agentic-evals
Pass
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The skill's primary function is to execute shell commands specified in evaluation fixtures and regression proof commands. This is performed using
subprocess.Popeninsrc/runner.pyandsubprocess.runinsrc/regressions.pyandsrc/remediation.py. - [COMMAND_EXECUTION]: The remediation loop logic in
src/remediation.pyinvokes external tools such as the GitHub CLI (gh) and sibling skills (phart-dag-chart,ticket) to automate the creation and tracking of issue reports based on evaluation failures. - [DATA_EXPOSURE]: The runner implements a redaction mechanism in
src/runner.py(via_redactand_SECRET_RE) designed to scrub sensitive strings like tokens, API keys, and passwords from captured stdout/stderr before they are persisted in the evaluation reports. - [COMMAND_EXECUTION]: The skill manages process group isolation during command execution in
src/runner.py, ensuring that timed-out trials do not leave orphan background processes that could corrupt subsequent tests.
Audit Metadata