llm-testing-expert

Installation
SKILL.md

LLM Testing and Evaluation

Description

Provide LLM test strategy design, test dataset construction, automated evaluation, and security red teaming recommendations to ensure models meet production requirements for functionality, performance, robustness, and safety.

When to Use

  • User requests "design LLM testing strategy" or "how to evaluate model quality"
  • User needs to build test datasets, design test cases, or perform regression testing
  • User seeks prompt optimization, engineering, and version management recommendations
  • User needs to evaluate model performance metrics (accuracy, hallucination rate, safety, latency, cost)
  • User asks how to conduct A/B testing or comparative evaluation (multi-model/multi-version)
  • User needs automated testing workflows or CI/CD integration
  • User requests red team testing design (jailbreak, prompt injection, adversarial attacks)
  • User asks about specialized testing strategies for RAG systems or Agent applications
Installs
2
GitHub Stars
5
First Seen
Jun 19, 2026
llm-testing-expert — miaoge-ge/coding-agent-skills