ingest-training-datalake
Pass
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: SAFE
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill uses the 'uv' package manager to handle its environment. The
run.shscript suggests a standard installation command fromastral.sh, which is a well-known and trusted source for this tool. This is a common setup procedure for Python-based skills using modern tooling. - [COMMAND_EXECUTION]: The skill executes local scripts and calls other internal tools (like
fetcher,memory, andtaxonomy) via subprocesses. These calls are used for its primary purpose of data ingestion and are restricted to the local workspace environment. - [DATA_EXPOSURE]: The skill includes path-traversal guardrails in
datalake_utils.py(_validate_training_root). It ensures that file operations (reading and writing) are restricted to an approved training corpus root (defaulting to/mnt/storage12tb/extractor_corpus), preventing unauthorized access to sensitive system files. - [INDIRECT_PROMPT_INJECTION]: The skill processes external URLs and metadata from downloaded documents. While it ingests untrusted data, it treats this data as files for storage and does not execute them as instructions. It uses structured JSON for reporting and includes metadata tags to maintain context, which aligns with best practices for handling untrusted content.
Audit Metadata