agentic-evals

Pass

Audited by Gen Agent Trust Hub on Aug 26, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill's primary function is to execute shell commands specified in evaluation fixtures and regression proof commands. This is performed using subprocess.Popen in src/runner.py and subprocess.run in src/regressions.py and src/remediation.py.
  • [COMMAND_EXECUTION]: The remediation loop logic in src/remediation.py invokes external tools such as the GitHub CLI (gh) and sibling skills (phart-dag-chart, ticket) to automate the creation and tracking of issue reports based on evaluation failures.
  • [DATA_EXPOSURE]: The runner implements a redaction mechanism in src/runner.py (via _redact and _SECRET_RE) designed to scrub sensitive strings like tokens, API keys, and passwords from captured stdout/stderr before they are persisted in the evaluation reports.
  • [COMMAND_EXECUTION]: The skill manages process group isolation during command execution in src/runner.py, ensuring that timed-out trials do not leave orphan background processes that could corrupt subsequent tests.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 26, 2026, 06:00 PM
Security Audit — agent-trust-hub — agentic-evals