perf-theory-tester
Pass
Audited by Gen Agent Trust Hub on Sep 18, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill instructs the agent to process data from external benchmark commands and file modifications which could potentially contain malicious instructions.
- Ingestion points: The agent is instructed to read and record evidence from the output of the benchmark command and specified changed files.
- Boundary markers: The instructions lack explicit delimiting markers or warnings to the agent to disregard instructions potentially embedded in benchmark results.
- Capability inventory: The agent is expected to perform shell command execution for benchmarks and file modifications to apply and revert experiment changes.
- Sanitization: No instructions are included for sanitizing or validating external content before it is processed by the agent.
Audit Metadata