content-refinement-agent
Pass
Audited by Gen Agent Trust Hub on Sep 28, 2026
Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes multiple local Python scripts to manage a refinement loop, including
apply_worklog.py,concession_guard.py,decision_band.py,score_delta.py,score_trajectory.py,snapshot.py, andupdate_critique_memory.py. It also invokeslatexmkto compile LaTeX documents into PDFs. These operations are restricted to the local workspace and are necessary for the skill's stated purpose. - [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted external data in the form of LaTeX source code (
paper.tex) and simulated reviewer feedback (JSON). It mitigates risks by employing deterministic integrity gates, such as numeric claim verification against an experimental log, citation integrity checks, and specific 'Iron Rules' that forbid the agent from fabricating data or promising new experiments. These boundaries reduce the attack surface for prompt injection from processed files. - [SAFE]: The skill follows security best practices by using deterministic scripts for decision-making (e.g.,
score_delta.pyfor iteration halting) rather than relying solely on LLM prose, which reduces susceptibility to sycophancy or logic errors. All file operations are scoped to a workspace directory structure.
Audit Metadata