benchmark-and-baseline-selector

Installation
SKILL.md

Benchmark and Baseline Selector

When the user proposes a method, claim, or result, return the smallest set of baselines and benchmarks that makes the claim credible, plus a complementary set that makes it strong.

The deliverable is always a two-tier comparison plan, not a single list:

  • Minimal — necessary. If the method ties or loses here, the claim is dead. Omitting any of these invites a desk-reject or "missing baseline" review.
  • Suggested — complementary. Adds external validity (scale, transfer, robustness), but is not decision-critical for the core claim.

A baseline/benchmark only earns a slot if it can change the conclusion. Decorative comparisons that the method obviously wins are noise.


Procedure

0. Context intake — gate before recommending

Do not produce a plan from an underspecified prompt. First check whether the conversation supplies the context below. Pull what is present; for anything required that is missing, ask the user — do not guess, and do not hedge about asking. A wrong incumbent or an unowned self-ablation makes the whole plan misleading, so it is always cheaper to ask first.

Installs
40
GitHub Stars
2
First Seen
Jun 16, 2026
benchmark-and-baseline-selector — jurgendn/agent-skills