autoresearch-method

Installation
SKILL.md

Autoresearch Method

autoresearch adds a verifiable autonomous experiment loop to a git repository:

agent edits bounded scope -> verifier runs -> result is scored -> commit is kept or reverted

The loop runs unattended. Every candidate is committed, scored against a fixed compute budget, and kept only if it strictly beats the current best on a single primary metric. Failures and regressions are reset to the parent commit. The methodology generalizes the pattern from karpathy/autoresearch.

When this method fits

A repo is a good candidate when all of the following hold:

  • Single, well-defined objective. One numeric scalar the agent should optimize (loss, latency, throughput, accuracy, score, error rate). If the goal is "make it better" with no metric, the loop has nothing to rank candidates by.
  • Verifier exists or can be written. Tests, benchmarks, or scoring code that run end-to-end without human judgment in a fixed compute budget. If grading requires a human in the loop, autoresearch will not help.
  • Bounded mutable scope. A small set of files where experiments make sense (one model file, one algorithm module, one config). Wider scope = noisier loop.
  • Verifier can be locked. Tests, fixtures, datasets, and scoring code can be marked immutable. If the verifier and the target are entangled, the agent will weaken the verifier instead of solving the problem.
  • Cheap, repeatable evaluation. Each candidate must finish in a budget you're willing to pay tens or hundreds of times. If a single run takes 12 hours, the loop is impractical.
Installs
1
First Seen
May 10, 2026
autoresearch-method — will-wright-eng/skills