scoring-agent-skills
Installation
SKILL.md
Scoring Agent Skills
Produce a comparable quality scorecard, not a vague impression or a runtime capability claim.
Scope and Evidence
- Resolve the requested skill set. For “every skill,” use the repository's active discovery tree unless the user explicitly includes deprecated or external skills.
- Establish the target-model baseline from the user's request, the repository's model-specific guide, or the execution environment, in that order. If none identifies one, use a “general capable model” baseline and disclose that low-confidence assumption.
- Inspect applicable repository instructions, the model-specific prompting guide, each
SKILL.md, catalog entry, and directly linked resources relevant to a score. - For the trusted current repository, run its established non-destructive validators when feasible. For an external or untrusted repository, default to source inspection or a trusted validator. Do not run repository-supplied code unless it is sandboxed and the user explicitly authorizes the exact command.
- Treat structural checks, link integrity, metadata presence, and test results as supporting evidence. They do not by themselves prove prompt quality, incremental value, or task success.
- Before scoring, read the rubric and assign an integer from 1 to 10 for each assessable dimension. Apply the same anchors to every skill and assess each dimension proportionately to what the skill can do.
- Support each score with direct evidence. For incremental knowledge value, label representative content as domain delta, justified activation, or redundancy candidate without inventing precise ratios. Revisit conspicuous outliers after the first pass so differences reflect the rubric rather than category or ordering bias.
Calculate and Report
Use equal weight for all assessed dimensions. When all six are assessed, compute the overall score as their arithmetic mean and show one decimal place. If inaccessible evidence prevents a defensible dimension score, mark that dimension unassessed, exclude it from the aggregate, and report score coverage and confidence; do not add hidden bonuses or penalties.
Return in the user's language unless requested otherwise: