llm-evaluation
Installation
SKILL.md
LLM Evaluation
Workflow
- Define the behavior to measure and the decision the eval supports.
- Build a representative dataset with normal, edge, adversarial, and failure cases.
- Choose grading: exact match, schema validation, heuristic checks, human rubric, or model judge.
- Track model, prompt, retrieval settings, and tool versions.
- Compare against a baseline before accepting changes.
Rules
- Keep eval cases independent of implementation details when possible.
- Include negative cases and ambiguity cases.
- Avoid overfitting prompts to tiny eval sets.