add-runner-eval
Pass
Audited by Gen Agent Trust Hub on Sep 17, 2026
Risk Level: SAFE
Full Analysis
- [COMMAND_EXECUTION]: The skill executes local Python scripts (
eval_program.py,run_live_roster.py, etc.) to validate evaluation definitions and rosters. These are internal project tools located within the repository structure. - [DATA_EXFILTRATION]: The skill contains strong proactive security guidelines, explicitly instructing the agent and user that 'credentials, secrets, private references, raw trusted artifacts, and hidden reasoning' must not be exposed or published in output reports.
- [REMOTE_CODE_EXECUTION]: There are no remote downloads or executions from untrusted sources. All execution is limited to local scripts within the Paperclip environment.
- [INDIRECT_PROMPT_INJECTION]: While the skill processes evaluation fixtures (JSON), it includes specific instructions for sanitization and uses a deterministic validation process, reducing the risk of malicious content affecting the agent's behavior during the evaluation lifecycle.
Audit Metadata