eval-rubric-citation-quality
Installation
SKILL.md
Eval Rubric — Citation Quality (0–5)
When to use this
Apply this rubric whenever an AI legal output includes references to statutes, regulations, cases, or other legal instruments. It is the most important rubric for preventing harm: a practitioner who acts on a fabricated statute faces professional liability; a client who relies on an invented court ruling may lose their case.
Run in the [[eval-llm-as-judge-system-prompt]] ensemble on every research, advisory, and review output. Apply independently of [[eval-rubric-hallucination-detection]] — citation quality measures the quality of real citations, while hallucination detection is the binary gate.
Scoring (0–5)
| Score | Label | Criteria |
|---|---|---|
| 5 | Excellent | All citations are real and verified; accurately quoted or paraphrased; pin-cited where appropriate (article number, paragraph, section); formatted per applicable citation style (Bluebook for US, OSCOLA for English, style appropriate to jurisdiction); every legal proposition has a supporting authority or is clearly marked as general background |
| 4 | Good | All citations real; minor formatting issues (e.g., missing year or publisher in case citation); or missing pin-cites in 1–2 instances; no misleading attribution |
| 3 | Acceptable | All citations real and substantively correct; inconsistent style across citation types; missing pin-cites in several instances; not misleading but would require cleanup before professional use |
| 2 | Poor | Most citations real but at least one is questionable (source may exist but quoted proposition is materially different from what it says) or clearly wrong (wrong jurisdiction statute cited) |
| 1 | Very poor | Mix of real and fabricated citations; or a single clearly fabricated source |
| 0 | Fail | Significant fabrication — invented case names, invented statute numbers, invented article numbers, invented pin-cites; or a single egregious fabrication on a material point |