create-eval
Installation
SKILL.md
Creating HolmesGPT Eval Tests
This skill provides the complete workflow for creating LLM evaluation tests in the HolmesGPT project. Eval tests validate that Holmes can correctly answer questions by querying real infrastructure and services.
Test Structure
Each eval lives in its own directory under tests/llm/fixtures/test_ask_holmes/:
tests/llm/fixtures/test_ask_holmes/<NNN>_<descriptive_name>/
├── test_case.yaml # Required: test definition
├── toolsets.yaml # Optional: enable specific toolsets
├── manifest.yaml # Optional: Kubernetes manifests
├── generate_*.py # Optional: data generation scripts
└── other supporting files
Naming convention: <3-digit-number>_<snake_case_description> (e.g., 212_large_configmap_needle).