task-completition-evaluation

Fail

Audited by Gen Agent Trust Hub on Jun 24, 2026

Risk Level: HIGHCREDENTIALS_UNSAFECOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
  • [CREDENTIALS_UNSAFE]: The skill requires reading and exporting environment variables from a hidden file in the user's home directory (~/.env.d/codex.env). Accessing such paths is a high-risk activity as they are commonly used to store sensitive authentication tokens and credentials.
  • [COMMAND_EXECUTION]: The skill performs automated evaluations by executing shell commands and local scripts, such as pytest and internal end-to-end evaluation scripts (eval_slack_e2e.py, eval_github_e2e.py).
  • [PROMPT_INJECTION]: The skill is vulnerable to indirect prompt injection by processing external communication data. * Ingestion points: The agent reads Slack thread transcripts, GitHub issue/PR content, and evaluation reports. * Boundary markers: No delimiters or specific instructions are provided to the agent to treat external content as untrusted. * Capability inventory: The agent has the ability to read sensitive system files and execute shell commands and Python scripts. * Sanitization: There is no evidence of sanitization or validation of the data retrieved from external sources before it is used to influence the agent's verdict.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Jun 24, 2026, 04:10 AM
Security Audit — agent-trust-hub — task-completition-evaluation