ln-34-benchmark-comparator

Installation
SKILL.md

Benchmark Comparator

Goal: Compare alternatives under controlled, reproducible conditions. Correctness comes before speed, and measured data must remain separate from estimates, setup cost, and interpretation.

Execution contract: The ordered checkboxes are the Definition of Done. Track every item internally as PENDING, PROVEN with concrete evidence, CLEARED with evidence that its condition is absent, or UNPROVEN with a gap; reading, delegation, or tool failure is not proof. Reconcile items after each section. Before returning, resolve all PENDING and count only PROVEN and CLEARED; apply the skill's verdict and approval rules to every gap. Preserve user intent, scope, and existing authorization. Continue authorized work; ask only for consequential unresolved choices or required external approval. Scale depth to material risk without silently skipping checks. Preserve dependency and safety ordering; otherwise choose the verification method appropriate to each obligation.

Tool Routing

Need Preferred tool Use it when Fallback
Canonical workload and oracle Repository fixtures, tests, expected diffs, schemas, or independently specified outcomes Defining what success means before either candidate runs Create the smallest deterministic fixture that represents the decision
Isolation Clean Git worktrees, temporary directories, controlled environment, fixed seeds, and resettable caches Preventing one candidate or run from contaminating another Sequential clean-room setup with verified cleanup
Execution The same shell runner and wrapper for every candidate Capturing commands, exit status, stdout, stderr, timing, and artifacts consistently Manual execution with an explicit reproducibility limitation
Activation proof Logs, traces, command records, process metadata, or candidate-specific artifacts Verifying the intended alternative actually ran and did not fall back Treat the run as invalid when activation cannot be proven
Correctness grading Tests, output parser, diff, schema validation, or independent oracle Every scenario before cost comparison Manual blind grading against written expectations
Performance and cost Monotonic timer, resource metrics, token or usage telemetry, tool-call logs, and failure counts Metrics are observable through the same method for all candidates Label derived or estimated values and keep them out of measured aggregates
External semantics Official documentation and specifications Candidate configuration or claimed behavior needs current verification Primary-source web research; otherwise mark the claim UNVERIFIED
Installs
49
GitHub Stars
559
First Seen
Jul 13, 2026
ln-34-benchmark-comparator — levnikolaevich/claude-code-skills