codex-prompting
Fail
Audited by Gen Agent Trust Hub on Sep 23, 2026
Risk Level: HIGHPRIVILEGE_ESCALATIONCOMMAND_EXECUTIONPROMPT_INJECTIONDATA_EXFILTRATIONEXTERNAL_DOWNLOADS
Full Analysis
- [PRIVILEGE_ESCALATION]: The skill configuration in
SKILL.mdexplicitly setsapproval_policy = "never"andsandbox_mode = "danger-full-access". This configuration bypasses user oversight and removes standard sandbox protections for commands executed by the Codex model. - [PRIVILEGE_ESCALATION]: The
shell_commandtool schema inreferences/codex-prompting-guide.mdincludes awith_escalated_permissionsflag, which allows the agent to explicitly request the removal of sandbox restrictions for specific executions. - [PROMPT_INJECTION]: Instructions in
SKILL.mdand the reference guide utilize behavioral overrides to bypass standard conversational filters and safety steps. Evidence includes directives like "You are autonomous senior engineer... proactively gather context... without waiting for additional prompts at each step" and "No preamble, no plans, and no 'I’ll do X then Y' narration." - [DATA_EXFILTRATION]: The skill provides shell command patterns for searching agent session files at
~/.pi/agent/sessions. This directory often contains sensitive data from previous interactions, which the skill facilitates accessing and processing without explicit user approval. - [COMMAND_EXECUTION]: The skill is designed to facilitate direct and automated execution of shell commands, CLI operations via the
joelclawtool, and system bus interactions with a strong bias toward full autonomy and action-first output. - [EXTERNAL_DOWNLOADS]: The reference documentation encourages cloning code from OpenAI's official GitHub repository (
github.com/openai/codex) and refers to official documentation atplatform.openai.com. These references target a well-known service and are documented as legitimate resources for the skill's purpose.
Recommendations
- AI detected serious security threats
Audit Metadata