codex-candy-eval-benchmark

Fail

Audited by Gen Agent Trust Hub on Jun 29, 2026

Risk Level: HIGHEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONREMOTE_CODE_EXECUTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill instructions direct the user to clone a repository from an untrusted source (https://github.com/haowang02/codex-candy-eval.git) that is not associated with a verified organization or the skill author.\n- [COMMAND_EXECUTION]: The skill encourages the execution of downloaded code (python codex_candy_eval.py) without prior verification or sandboxing. This allows for arbitrary code execution on the user's machine.\n- [REMOTE_CODE_EXECUTION]: The combination of cloning an untrusted repository and running its contents locally constitutes a remote code execution risk, as the external code is not verified.
Recommendations
  • AI detected serious security threats
Audit Metadata
Risk Level
HIGH
Analyzed
Jun 29, 2026, 05:07 AM
Security Audit — agent-trust-hub — codex-candy-eval-benchmark