benchmark

Pass

Audited by Gen Agent Trust Hub on Sep 12, 2026

Risk Level: SAFEPROMPT_INJECTIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDYNAMIC_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill mandates a specific activation sentinel, [BENCHMARK-SKILL-ACTIVE v1 / skills/benchmark/SKILL.md], which must appear as the very first line of the agent's response after reading the file, overriding default response formatting.\n- [INDIRECT_PROMPT_INJECTION]:\n * Ingestion points: The skill ingests external data from https://realfood.gov using capture and extraction tools.\n * Boundary markers: The skill lacks explicit delimiters or warnings to treat the ingested web content as untrusted data.\n * Capability inventory: The skill utilizes bash for script execution, npm for project setup, and agent-browser eval for executing JavaScript.\n * Sanitization: Ingested content from the reference site is processed without explicit sanitization steps.\n- [COMMAND_EXECUTION]: The workflow relies on shell commands for environment setup and management, including symlink manipulation (ln -sfn) and process termination (pkill -f) of development servers.\n- [DYNAMIC_EXECUTION]: The skill executes dynamic JavaScript within a browser session via agent-browser eval and automates the creation and execution of a Next.js project using npx create-next-app and npm run dev.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 12, 2026, 12:02 PM
Security Audit — agent-trust-hub — benchmark