extract-html

Pass

Audited by Gen Agent Trust Hub on Mar 17, 2026

Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
  • [PROMPT_INJECTION]: The skill is vulnerable to Indirect Prompt Injection (Category 8) because it incorporates untrusted HTML content directly into LLM prompts to perform data extraction.
  • Ingestion points: Untrusted data enters the agent context via the raw_html parameter in extract_html/pipeline.py (read from the user-provided --html file).
  • Boundary markers: The prompt in extract_html/schematron.py uses simple markdown headers like ### HTML CONTENT to delimit the input, but these are insufficient to prevent a sophisticated model from obeying instructions embedded within the HTML text or comments.
  • Capability inventory: The skill has the capability to perform network POST requests to LLM/Vision providers (extract_html/schematron.py), fetch remote images via HTTP GET (extract_html/media.py), and write JSON/text results to the local filesystem (extract_html/cli.py).
  • Sanitization: While extract_html/cleaning.py removes some HTML tags (script, style, etc.), it does not sanitize or escape natural language instructions that may be present in the remaining visible text, which the LLM processes as context.
  • [EXTERNAL_DOWNLOADS]: The skill can fetch remote resources from arbitrary domains found within the input HTML.
  • Evidence: extract_html/media.py contains load_image_bytes, which uses httpx.Client to download images from http(s) URLs discovered in <img> tags if the --fetch-remote-media flag is enabled. This could be leveraged for SSRF (Server-Side Request Forgery) or tracking if the agent processes maliciously crafted HTML.
Audit Metadata
Risk Level
SAFE
Analyzed
Mar 17, 2026, 06:35 AM
Security Audit — agent-trust-hub — extract-html