improvement-evaluator

Installation
SKILL.md

Improvement Evaluator

Measures whether a Skill actually makes AI perform better on real tasks, not just whether the SKILL.md document looks well-structured.

Why Execution Testing Matters

Structural scoring (word count, section presence, formatting) correlates poorly with actual AI task performance. Internal benchmarks showed R²=0.00 between document-structure scores and execution pass rates across 40+ skill evaluations. A perfectly formatted SKILL.md can still produce failing task outputs if the instructions mislead the model or omit critical constraints.

Installs
1
GitHub Stars
5
First Seen
Apr 8, 2026
improvement-evaluator — lanyasheng/auto-improvement-orchestrator-skill