paper-figure
Pass
Audited by Gen Agent Trust Hub on Sep 15, 2026
Risk Level: SAFECOMMAND_EXECUTIONDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [DYNAMIC_EXECUTION]: The skill generates Python scripts at runtime (e.g.,
paper_plot_style.pyandgen_fig*.py) and executes them to produce figures. This involves assembling executable code from internal templates and project data, which is a significant capability. - [COMMAND_EXECUTION]: The skill uses the
bashtool to execute the generated Python scripts in a loop (for script in gen_fig*.py; do python "$script"; done). - [INDIRECT_PROMPT_INJECTION]: The skill ingests data from the project environment which is used to influence code generation and external model queries.
- Ingestion points: The skill reads and parses
PAPER_PLAN.mdand experiment data files such asexp.jsonor CSV files. - Boundary markers: There are no explicit boundary markers or instructions to ignore embedded commands when the agent processes the contents of these files.
- Capability inventory: The skill can write files, execute arbitrary Python code via
bash, and call themcp__codex__codextool to interact with other LLM models. - Sanitization: The skill lacks sanitization or validation of the data ingested from
PAPER_PLAN.mdor experiment logs before it is interpolated into Python scripts or sent to theREVIEWER_MODELfor evaluation.
Audit Metadata