benchmark-and-baseline-selector
Installation
SKILL.md
Benchmark and Baseline Selector
When the user proposes a method, claim, or result, return the smallest set of baselines and benchmarks that makes the claim credible, plus a complementary set that makes it strong.
The deliverable is always a two-tier comparison plan, not a single list:
- Minimal — necessary. If the method ties or loses here, the claim is dead. Omitting any of these invites a desk-reject or "missing baseline" review.
- Suggested — complementary. Adds external validity (scale, transfer, robustness), but is not decision-critical for the core claim.
A baseline/benchmark only earns a slot if it can change the conclusion. Decorative comparisons that the method obviously wins are noise.
Procedure
0. Context intake — gate before recommending
Do not produce a plan from an underspecified prompt. First check whether the conversation supplies the context below. Pull what is present; for anything required that is missing, ask the user — do not guess, and do not hedge about asking. A wrong incumbent or an unowned self-ablation makes the whole plan misleading, so it is always cheaper to ask first.