gan-style-harness
Pass
Audited by Gen Agent Trust Hub on Sep 1, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDYNAMIC_EXECUTIONMETADATA_POISONING
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill architecture creates a feedback loop where user prompts and agent-generated content are processed and executed.
- Ingestion points: User-provided one-line prompts, 'spec.md' (agent-generated), and 'feedback-NNN.md' (agent-generated).
- Boundary markers: No explicit delimiters or 'ignore embedded instructions' warnings are provided for these ingestion points to prevent malicious overrides.
- Capability inventory: The system uses 'Bash' for command execution, 'Write' and 'Edit' for file manipulation, and 'Playwright MCP' for interacting with live web applications.
- Sanitization: No sanitization or validation of the generated code or feedback is mentioned before it is executed or re-ingested.
- [METADATA_POISONING]: The skill provides deceptive references to gain trust.
- Evidence: It claims to be based on an 'Anthropic's March 2026 harness design paper' and provides a corresponding link. As of the current date, this is a non-existent future-dated resource.
- [COMMAND_EXECUTION]: The skill instructions direct the agent to perform sensitive shell operations.
- Evidence: Usage examples include starting development servers ('npm run dev'), managing version control ('git'), and executing shell scripts ('./scripts/gan-harness.sh').
- [DYNAMIC_EXECUTION]: The Evaluator agent is instructed to interact with live, dynamically generated application code.
- Evidence: It uses Playwright to 'click through features, fill forms, test API endpoints' on code just written by the Generator agent. This allows potentially malicious generated code to execute within the agent's browser context.
Audit Metadata