agent-instructions-evaluator

Installation
SKILL.md

Agent Instructions Evaluator

Overview

Evaluate agent instructions or agent definitions for operational achievability in production settings. This skill focuses on runtime reliability and runtime efficiency rather than writing quality, identifying issues from hidden state, conflicting rules, vague scope, brittle exact phrasing, underspecified tool behavior, instruction overload, and performance friction caused by excessive or contradictory runtime reasoning.

Use this skill when you need:

  • A practical evaluation report focused on runtime reliability
  • Evidence-backed prompt review with concrete recommendations
  • A reusable report artifact that can be shared with reviewers
  • Per-dimension scoring that preserves nuance rather than averaging away critical issues

Core Principle

Score the artifact not by how much behavior it describes, but by how much behavior the agent can reliably execute. More rules do not automatically make a better prompt—more rules often lower achievability. An instruction set that is technically understandable but expensive to reconcile at runtime should still be considered lower-achievability, because slow, unstable, or tool-heavy execution reduces production reliability.

Evaluation Workflow

Installs
7
GitHub Stars
177
First Seen
Jul 29, 2026
agent-instructions-evaluator — ibm/ibm-watsonx-orchestrate-adk