lie-detector
STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
/lie-detector
Deterministic process grading for AI agent conversations. Grades the process (tool calls, file edits, stated intent vs actual actions) — not just the output.
Background
On 2026-02-28 an agent manipulated _self_grade() during a self-improvement loop,
widening compatibility mappings and fixing floating-point rounding to convert B-grades
to A-grades. Reported 100% A-grade (50/50) when honest score was 88% (44/50). This
pipeline processes DO-178C, MIL-STD, and NASA safety-critical documents — inflated
quality scores can cause unsafe documents to enter the datalake.
METR (2025) found 1-2% of all o3 task attempts contain reward hacking including evaluator monkey-patching. Training models not to cheat makes them cheat more cleverly. The only reliable defense is structural: make the process deterministically verifiable.