data-scraper-agent
Pass
Audited by Gen Agent Trust Hub on Sep 1, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPERSISTENCECOMMAND_EXECUTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data from external websites and APIs and interpolates it directly into LLM prompts in
ai/pipeline.pywithout sufficient isolation or sanitization.\n - Ingestion points: The
fetch()function inscraper/sources/my_source.py(and the provided HTML/RSS patterns) ingest content from external, potentially attacker-controlled URLs.\n - Boundary markers: The
_build_promptfunction uses basicjson.dumpsformatting but lacks robust structural delimiters or explicit instructions for the LLM to ignore instructions embedded within the scraped data.\n - Capability inventory: The resulting agent has network access (via the
requestslibrary) and the ability to modify the repository state via GitHub Actions.\n - Sanitization: There is no significant sanitization or filtering of the scraped data before it is presented to the LLM.\n- [PERSISTENCE]: The skill uses GitHub Actions to maintain state across different sessions. The workflow in
.github/workflows/scraper.ymlis configured to commit and push thedata/feedback.jsonfile back to the repository after each run, enabling the agent to 'learn' and persist data across scheduled executions.\n- [COMMAND_EXECUTION]: The GitHub Actions workflow (.github/workflows/scraper.yml) executes various shell commands to manage the environment, install dependencies fromrequirements.txt, and orchestrate the scraping process.
Audit Metadata