benchmark-models
Installation
SKILL.md
benchmark-models
Run standardized compliance QRA tests against candidate LLMs to evaluate accuracy, latency, and cost before deploying to the inference pipeline.
Usage
Run a single-model benchmark
./run.sh run --model deepseek-v3 --suite compliance-basic
Runs 20 gold-set compliance QRA questions against the specified model via /scillm. Outputs a results table with TEST_CASE, EXPECTED, ACTUAL, MATCH, and LATENCY_MS columns, plus summary metrics (accuracy%, latency_p50, latency_p95, estimated_token_cost).
Compare multiple models
./run.sh compare --models "deepseek-v3,llama-3.1-70b" --suite compliance-basic