challenge
Installation
SKILL.md
Challenge — does the result survive the choices you didn't make?
A single specification is one draw from a distribution you never looked at.
Why this exists, measured rather than asserted. In a controlled study, 150 autonomous agents were given the same data and the same questions. Effect-size interquartile ranges reached ~10.7 %/yr, and the spread concentrated in discrete measure-choice forks — not in estimation noise. Within a measure family, agents agreed to ~0.25 %/yr. Two findings from that study shape this skill:
- AI peer review left the spread essentially unchanged. Review catches errors; it does not reduce analytical-choice variance. A clean referee report is not robustness.
- Exposure to exemplar papers collapsed the spread by 80–99 % — convergence by imitation, not by correctness. Herding is not agreement.
So the spread has to be measured, not reviewed away.