llm-evaluation

Installation
SKILL.md

LLM Evaluation

Workflow

  1. Define the behavior to measure and the decision the eval supports.
  2. Build a representative dataset with normal, edge, adversarial, and failure cases.
  3. Choose grading: exact match, schema validation, heuristic checks, human rubric, or model judge.
  4. Track model, prompt, retrieval settings, and tool versions.
  5. Compare against a baseline before accepting changes.

Rules

  • Keep eval cases independent of implementation details when possible.
  • Include negative cases and ambiguity cases.
  • Avoid overfitting prompts to tiny eval sets.

Verification

Installs
2
Repository
yiweiwan/skills
GitHub Stars
1
First Seen
May 13, 2026
llm-evaluation — yiweiwan/skills