data-engineering-medallion-pipeline
Fail
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: HIGHEXTERNAL_DOWNLOADSREMOTE_CODE_EXECUTIONCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill instructs the agent to clone a repository from an unverified GitHub user account (
LucasGoulartCouto/data-engineering-medallion.git). - [REMOTE_CODE_EXECUTION]: The skill requires the execution of a
make setupcommand immediately after cloning the external repository. This initiates the execution of arbitrary scripts and commands defined within the downloaded project's Makefile and associated scripts without prior verification. - [COMMAND_EXECUTION]: The skill makes heavy use of local shell operations including
docker-compose,make, andgit. These commands operate on content downloaded from the internet, which could lead to privilege escalation or system compromise if the external repository is malicious. - [INDIRECT_PROMPT_INJECTION]:
- Ingestion points: The pipeline is designed to ingest raw data from external sources into a MinIO bucket (
raw-data) which is then processed through multiple database layers. - Boundary markers: There are no clear boundary markers or instructions to the agent to treat the ingested data as untrusted, which could lead the agent to follow instructions embedded within the processed data.
- Capability inventory: The skill possesses significant capabilities including arbitrary shell execution via Airflow
BashOperator, Python execution viaPythonOperator, and full database read/write access. - Sanitization: Sanitization is limited to standard SQL type casting (e.g.,
::INTEGER) and basic string cleaning in the DBT models, which is insufficient to prevent complex injection attacks.
Recommendations
- AI detected serious security threats
Audit Metadata