ingest-training-datalake

Pass

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: SAFE
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill uses the 'uv' package manager to handle its environment. The run.sh script suggests a standard installation command from astral.sh, which is a well-known and trusted source for this tool. This is a common setup procedure for Python-based skills using modern tooling.
  • [COMMAND_EXECUTION]: The skill executes local scripts and calls other internal tools (like fetcher, memory, and taxonomy) via subprocesses. These calls are used for its primary purpose of data ingestion and are restricted to the local workspace environment.
  • [DATA_EXPOSURE]: The skill includes path-traversal guardrails in datalake_utils.py (_validate_training_root). It ensures that file operations (reading and writing) are restricted to an approved training corpus root (defaulting to /mnt/storage12tb/extractor_corpus), preventing unauthorized access to sensitive system files.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes external URLs and metadata from downloaded documents. While it ingests untrusted data, it treats this data as files for storage and does not execute them as instructions. It uses structured JSON for reporting and includes metadata tags to maintain context, which aligns with best practices for handling untrusted content.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 17, 2026, 06:35 AM
Security Audit — agent-trust-hub — ingest-training-datalake