evaluating-skills-with-models
Installation
SKILL.md
Evaluating Skills with Models
Evaluate skills across multiple Claude models using sub-agents with quality-based scoring.
Requirement: Claude Code CLI only. Not available in Claude.ai.
Why Quality-Based Scoring
Binary pass/fail ("did it do X?") fails to differentiate models - all models can "do the steps." The difference is how well they do them. This skill uses weighted scoring to reveal capability differences.
Workflow
Step 1: Load Test Scenarios
Check for tests/scenarios.md in the target skill directory.
Default to difficult scenarios: When multiple scenarios exist, prioritize Hard or Medium difficulty scenarios for evaluation. Easy scenarios often don't show meaningful differences between models and aren't realistic for production use.
Required scenario format: