ai-data-engineering-rag-pipeline
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill instructs the user to clone a repository from a third-party GitHub account (Nahid-mahmud555/ai-data-engineering-roadmap) and install dependencies from its requirements.txt files.
- [INDIRECT_PROMPT_INJECTION]: The skill implements a pipeline that ingests and processes data from external document files (corpus directory) and structured datasets (questions.jsonl).
- Ingestion points: The search implementation reads all text files within a 'corpus' directory and a JSONL file for evaluation queries.
- Boundary markers: The provided implementation does not include boundary markers or 'ignore instructions' delimiters when processing document content.
- Capability inventory: The code performs file system reads, building a BM25 index, semantic vector search using FAISS, and console logging of retrieval results.
- Sanitization: Content is only subjected to basic lowercase conversion and whitespace tokenization, without specific sanitization against embedded prompt instructions.
Audit Metadata