eval-dataset-competitor-comparison-set
Installation
SKILL.md
Eval Dataset — Competitor Comparison Set
Scope
A curated set of legal AI prompts run against this system and comparable legal AI tools (Harvey, Clio Duo, LexisNexis AI, generic Claude/GPT-4o) to produce a structured quality comparison. The primary hypothesis: this system outperforms general-purpose LLMs and US-centric legal AI tools on MENA-jurisdiction tasks (UAE, KSA, Lebanon, DIFC/ADGM) while being competitive on standard tasks.
Results feed [[eval-leaderboard-updater]] and the weekly AI quality trend report.
How to use this pack
- Select 20–30 prompts from the comparison set (see categories below).
- Run each prompt through:
- This system (production endpoint)
- The baseline tool(s) (API or UI, documented)
- Score each output with [[eval-llm-as-judge-system-prompt]] using [[eval-rubric-legal-soundness]], [[eval-rubric-jurisdiction-awareness]], and [[eval-rubric-hallucination-detection]].
- Record scores and qualitative notes in the comparison spreadsheet.
- Publish to the internal leaderboard monthly.
Never publish raw competitor output without legal review of their terms of service regarding benchmarking.