evaluations

Installation
SKILL.md

Route an Evaluation Request

This is a compatibility skill. Do not build an experiment, monitor, or guardrail from this skill.

Classify the user's intent:

Intent Correct skill
Batch test a dataset, compare prompts or models, benchmark, create a CI quality gate experiments
Score live traces or threads, monitor production quality, create a guardrail online-evaluations

If the request remains ambiguous after inspecting context (a bare "make me an eval" that names neither a dataset nor live traffic), do not create anything yet. This choice picks what gets tested, so it is the user's to make, not a default's. Ask it as a question card and stop; the answer arrives as the next message.

Where langy-card blocks render, ask it as a choices block (the only sanctioned question format) last in the reply:

Installs
114
GitHub Stars
3
First Seen
Mar 18, 2026
evaluations — langwatch/skills