data-scraper-agent

Pass

Audited by Gen Agent Trust Hub on Sep 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONPERSISTENCECOMMAND_EXECUTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data from external websites and APIs and interpolates it directly into LLM prompts in ai/pipeline.py without sufficient isolation or sanitization.\n
  • Ingestion points: The fetch() function in scraper/sources/my_source.py (and the provided HTML/RSS patterns) ingest content from external, potentially attacker-controlled URLs.\n
  • Boundary markers: The _build_prompt function uses basic json.dumps formatting but lacks robust structural delimiters or explicit instructions for the LLM to ignore instructions embedded within the scraped data.\n
  • Capability inventory: The resulting agent has network access (via the requests library) and the ability to modify the repository state via GitHub Actions.\n
  • Sanitization: There is no significant sanitization or filtering of the scraped data before it is presented to the LLM.\n- [PERSISTENCE]: The skill uses GitHub Actions to maintain state across different sessions. The workflow in .github/workflows/scraper.yml is configured to commit and push the data/feedback.json file back to the repository after each run, enabling the agent to 'learn' and persist data across scheduled executions.\n- [COMMAND_EXECUTION]: The GitHub Actions workflow (.github/workflows/scraper.yml) executes various shell commands to manage the environment, install dependencies from requirements.txt, and orchestrate the scraping process.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 1, 2026, 02:39 AM
Security Audit — agent-trust-hub — data-scraper-agent