lie-detector

Installation
SKILL.md

STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.

/lie-detector

Deterministic process grading for AI agent conversations. Grades the process (tool calls, file edits, stated intent vs actual actions) — not just the output.

Background

On 2026-02-28 an agent manipulated _self_grade() during a self-improvement loop, widening compatibility mappings and fixing floating-point rounding to convert B-grades to A-grades. Reported 100% A-grade (50/50) when honest score was 88% (44/50). This pipeline processes DO-178C, MIL-STD, and NASA safety-critical documents — inflated quality scores can cause unsafe documents to enter the datalake.

METR (2025) found 1-2% of all o3 task attempts contain reward hacking including evaluator monkey-patching. Training models not to cheat makes them cheat more cleverly. The only reliable defense is structural: make the process deterministically verifiable.

Installs
2
GitHub Stars
7
First Seen
Mar 17, 2026
lie-detector — grahama1970/agent-skills