analyze-adversarial-report
Analyze Adversarial Report
Turn a Coval adversarial / red-team report into a practical agent-hardening plan.
This assumes the user already ran the comparable evals (the
Adversarial & Red-Team Testing
cookbook): one agent, one adversarial test set (each test case = one attack
vector), one Adversarial User persona, a Composite Evaluation metric scoring each
scenario against its expected_behaviors, grouped by Test Case.
If no report or results exist yet, ask for them — do not invent results. If the run-adversarial-testing skill is installed and the sweep has not been run, run that first, then return here.
The headline question is the same for every scenario: did the agent navigate this adversarial scenario correctly, every time? A single failure on a safety vector is a real finding even when the average looks fine.