data-scraper-agent
Pass
Audited by Gen Agent Trust Hub on Apr 1, 2026
Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill is designed to ingest untrusted data from the public internet and process it using an LLM, creating a surface for indirect prompt injection.
- Ingestion points: Data is fetched from arbitrary URLs or APIs in the
scraper/sources/modules andscraper/main.py. - Boundary markers: The prompt construction in
ai/pipeline.pyuses markdown-style headers (e.g., '# Items', '# Instructions') to separate data from system prompts. However, it lacks specific instructions telling the LLM to ignore directives found within the scraped content. - Capability inventory: The agent can write to external databases (Notion) and has permissions to commit and push changes back to its GitHub repository.
- Sanitization: The skill uses
json.dumpsto serialize scraped data before interpolation, which provides structural delimitation but does not sanitize the semantic content for adversarial instructions. - [COMMAND_EXECUTION]: The GitHub Actions configuration in
.github/workflows/scraper.ymlexecutes shell commands for dependency installation (pip install), running the main script (python -m scraper.main), and committing feedback data to the repository via git CLI. These operations are standard for the skill's intended CI/CD workflow. - [EXTERNAL_DOWNLOADS]: The skill utilizes the
requestsandplaywrightlibraries to download content from arbitrary external web sources. It also makes network requests to the Google Gemini API for AI enrichment tasks.
Audit Metadata