test-lab

Installation
SKILL.md

STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.

test-lab: Adversarial Blind Evaluation Skill

Enforces an information barrier between the agent writing code and the tests verifying it.

The Problem

Coding agents write code and tests in the same context, then optimize for passing their own tests rather than actual correctness. ImpossibleBench (arXiv:2510.20270) showed GPT-5 cheats 76% of the time when it can see tests, but near-zero when tests are hidden.

How It Works

  1. Generate: Parse a task file, identify applicable best-practices domains, generate hidden tests
  2. Guard: Extension blocks all agent reads of test dirs and generator source
  3. Run: Execute hidden tests in a separate subprocess
  4. Report: Return only "PASS" or "FAIL: {description}" — no test source, no assertion code
Installs
3
GitHub Stars
7
First Seen
Mar 17, 2026
test-lab — grahama1970/agent-skills