sample-app-benchmark
Installation
SKILL.md
Sample App Benchmark
Use this skill when the question is whether a new system, workflow CLI, or instruction layer actually helps on real app work.
This skill standardizes the benchmark shape:
- one clean source repo
mainas the committed baseline- one committed challenger branch
- one top-level Codex session per arm
- one isolated benchmark
CODEX_HOMEper arm using thesample_app_evalprofile - one normalized post-run verifier for both arms
- no benchmark residue in the source repo
Use When
- benchmarking a new agent aid on a fresh sample app
- comparing docs-first guidance vs a challenger branch
- testing whether a workflow CLI improves discovery on real implementation work
- creating reusable benchmark evidence for multiple systems with the same method