spark-data-engineering-pipeline
Warn
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: MEDIUMEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONREMOTE_CODE_EXECUTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill instructions direct the user to clone a repository from an unverified GitHub user account (github.com/Giri-25/spark-data-engineering-pipeline.git). As this is not a trusted organization or well-known service, the integrity of the downloaded source code cannot be confirmed.
- [COMMAND_EXECUTION]: The installation and execution steps involve running shell commands (git clone, pip install, python main.py) that facilitate the execution of code from the external repository.
- [REMOTE_CODE_EXECUTION]: The skill instructs the user to download code from an external URL and then execute it locally (python main.py), which is a pattern for executing remote code from an untrusted source.
- [INDIRECT_PROMPT_INJECTION]: The skill facilitates the ingestion and processing of untrusted data with the following profile: 1. Ingestion points: Reads JSON data from AWS S3 buckets in src/extract/s3_extractor.py. 2. Boundary markers: The skill lacks explicit delimiters or instructions to ignore embedded commands within the processed data. 3. Capability inventory: The pipeline performs database writes to PostgreSQL via JDBC and file writes to S3 and local logs. 4. Sanitization: Includes schema validation in src/validation/schema_validator.py and data quality checks in src/validation/quality_checker.py, which verify data structure but do not sanitize against natural language injection.
Audit Metadata