data-scraper-agent

Pass

Audited by Gen Agent Trust Hub on Apr 1, 2026

Risk Level: SAFEPROMPT_INJECTIONCOMMAND_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
  • [PROMPT_INJECTION]: The skill is designed to ingest untrusted data from the public internet and process it using an LLM, creating a surface for indirect prompt injection.
  • Ingestion points: Data is fetched from arbitrary URLs or APIs in the scraper/sources/ modules and scraper/main.py.
  • Boundary markers: The prompt construction in ai/pipeline.py uses markdown-style headers (e.g., '# Items', '# Instructions') to separate data from system prompts. However, it lacks specific instructions telling the LLM to ignore directives found within the scraped content.
  • Capability inventory: The agent can write to external databases (Notion) and has permissions to commit and push changes back to its GitHub repository.
  • Sanitization: The skill uses json.dumps to serialize scraped data before interpolation, which provides structural delimitation but does not sanitize the semantic content for adversarial instructions.
  • [COMMAND_EXECUTION]: The GitHub Actions configuration in .github/workflows/scraper.yml executes shell commands for dependency installation (pip install), running the main script (python -m scraper.main), and committing feedback data to the repository via git CLI. These operations are standard for the skill's intended CI/CD workflow.
  • [EXTERNAL_DOWNLOADS]: The skill utilizes the requests and playwright libraries to download content from arbitrary external web sources. It also makes network requests to the Google Gemini API for AI enrichment tasks.
Audit Metadata
Risk Level
SAFE
Analyzed
Apr 1, 2026, 11:09 AM
Security Audit — agent-trust-hub — data-scraper-agent