agents-introspection

Warn

Audited by Gen Agent Trust Hub on Sep 4, 2026

Risk Level: MEDIUMPROMPT_INJECTIONDATA_EXFILTRATIONINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTION
Full Analysis
  • [PROMPT_INJECTION]: The skill includes a specific instruction to "skip the ai-coord gate," which is an attempt to override or bypass platform-level orchestration and safety controls for its operations.
  • [DATA_EXFILTRATION]: The skill is designed to access and read sensitive data from personal directories, specifically ~/.codex and ~/.claude, which contain interaction history. The instructions grant the agent authority to read transcripts from any materially relevant local project without asking for user permission for each access.
  • [COMMAND_EXECUTION]: The skill executes shell commands (pwd -P) and bundled Python scripts via uv run (scripts/transcript-miner.py and scripts/transcript-inspect.py) to search and digest local transcript files.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted historical data found in transcripts, creating a vulnerability surface where malicious instructions from previous sessions could influence the agent's current recommendations or actions.
  • Ingestion points: Historical transcript JSONL files located in ~/.codex/sessions, ~/.codex/archived_sessions, and ~/.claude/projects/.
  • Boundary markers: The transcript-inspect.py script uses per-line channel prefixes (e.g., user:, assistant:) to delimit historical content in the generated report.
  • Capability inventory: The skill has authority to write modifications to AGENTS.md and local skill files, and executes shell scripts via uv run as documented in SKILL.md and references/transcript-sources.md.
  • Sanitization: The skill implements robust regex-based redaction for emails, API keys, and blockchain addresses in scripts/transcript_common.py, but it does not filter or sanitize executable instructions or prompt injection patterns found in the mined content.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 4, 2026, 09:31 AM