error-analysis

Installation
SKILL.md

Error Analysis

Evals-first failure analysis for LLM applications. The method (Hamel Husain, Shreya Shankar) is qualitative research applied to traces: a human open-codes real failures, Claude axial-codes the notes into a named taxonomy, and only the recurring named modes earn an automated eval. Error analysis decides what to measure; it never writes an eval for a mode that has no name and no count.

When to Use

  • An LLM feature (chatbot, RAG, agent, extractor) misbehaves and the team is guessing which evals to write
  • You have production traces in Langfuse (or a JSONL export) and need the top failure modes with counts
  • Eval scores exist but nobody trusts them because they were never grounded in real failures
  • A judge or metric exists and its agreement with human labels is unknown

Do NOT use for Claude Code session errors (errors skill), failing CI runs (ci-debug), or writing the evaluators themselves once modes are named (testing-llm, ork:eval-runner).

Task Management (CC 2.1.16)

Multi-phase workflow: create tasks before Phase 1 and keep status current.

Installs
1
GitHub Stars
284
First Seen
4 days ago
error-analysis — yonatangross/orchestkit