extract-html
Pass
Audited by Gen Agent Trust Hub on Mar 17, 2026
Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill is vulnerable to Indirect Prompt Injection (Category 8) because it incorporates untrusted HTML content directly into LLM prompts to perform data extraction.
- Ingestion points: Untrusted data enters the agent context via the
raw_htmlparameter inextract_html/pipeline.py(read from the user-provided--htmlfile). - Boundary markers: The prompt in
extract_html/schematron.pyuses simple markdown headers like### HTML CONTENTto delimit the input, but these are insufficient to prevent a sophisticated model from obeying instructions embedded within the HTML text or comments. - Capability inventory: The skill has the capability to perform network POST requests to LLM/Vision providers (
extract_html/schematron.py), fetch remote images via HTTP GET (extract_html/media.py), and write JSON/text results to the local filesystem (extract_html/cli.py). - Sanitization: While
extract_html/cleaning.pyremoves some HTML tags (script, style, etc.), it does not sanitize or escape natural language instructions that may be present in the remaining visible text, which the LLM processes as context. - [EXTERNAL_DOWNLOADS]: The skill can fetch remote resources from arbitrary domains found within the input HTML.
- Evidence:
extract_html/media.pycontainsload_image_bytes, which useshttpx.Clientto download images fromhttp(s)URLs discovered in<img>tags if the--fetch-remote-mediaflag is enabled. This could be leveraged for SSRF (Server-Side Request Forgery) or tracking if the agent processes maliciously crafted HTML.
Audit Metadata