data-scraper-agent

Pass

Audited by Gen Agent Trust Hub on Sep 12, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADSCOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill is designed to ingest data from untrusted external sources (websites, APIs) and include this content directly in prompts sent to the Gemini AI model.
  • Ingestion points: Scraper source files (e.g., scraper/sources/source_name.py) fetch data from external endpoints; profile/context.md and data/feedback.json provide additional context.
  • Boundary markers: The prompt in ai/pipeline.py uses simple Markdown headers (# Items, # User Context) and json.dumps() for content. These markers are insufficient to prevent an attacker from embedding instructions within the scraped text that the LLM might follow.
  • Capability inventory: The resulting agent has network capabilities (requests, notion-client) and the ability to update its own repository state (git push in GitHub Actions).
  • Sanitization: No explicit content filtering, escaping, or "ignore instructions" directives are included in the prompt construction logic.
  • [EXTERNAL_DOWNLOADS]: The skill instructs the user to install several standard Python packages via a requirements.txt file.
  • Packages include: requests, beautifulsoup4, lxml, python-dotenv, pyyaml, notion-client, and optionally playwright. These are well-known, industry-standard libraries.
  • [COMMAND_EXECUTION]: The provided GitHub Actions workflow (.github/workflows/scraper.yml) executes shell commands to set up the environment, install dependencies, run the scraper script, and commit feedback data.
  • The execution is transparent and scoped to the automation task defined in the skill.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 12, 2026, 03:40 PM
Security Audit — agent-trust-hub — data-scraper-agent