data-scraper-agent
Pass
Audited by Gen Agent Trust Hub on Sep 12, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest data from untrusted external sources (websites, APIs) and include this content directly in prompts sent to the Gemini AI model.
- Ingestion points: Scraper source files (e.g., scraper/sources/source_name.py) fetch data from external endpoints; profile/context.md and data/feedback.json provide additional context.
- Boundary markers: The prompt in ai/pipeline.py uses simple Markdown headers (# Items, # User Context) and json.dumps() for content. These markers are insufficient to prevent an attacker from embedding instructions within the scraped text that the LLM might follow.
- Capability inventory: The resulting agent has network capabilities (requests, notion-client) and the ability to update its own repository state (git push in GitHub Actions).
- Sanitization: No explicit content filtering, escaping, or "ignore instructions" directives are included in the prompt construction logic.
- [EXTERNAL_DOWNLOADS]: The skill instructs the user to install several standard Python packages via a requirements.txt file.
- Packages include: requests, beautifulsoup4, lxml, python-dotenv, pyyaml, notion-client, and optionally playwright. These are well-known, industry-standard libraries.
- [COMMAND_EXECUTION]: The provided GitHub Actions workflow (.github/workflows/scraper.yml) executes shell commands to set up the environment, install dependencies, run the scraper script, and commit feedback data.
- The execution is transparent and scoped to the automation task defined in the skill.
Audit Metadata