prompt-benchmark

Installation
SKILL.md

Prompt Benchmark Framework

Systematic evaluation of prompts against standard AI benchmarks.

Core Benchmarks

Benchmark Domain Metric Baseline Meta-Prompt Target
MATH Mathematics Accuracy ~50% ≥70%
GSM8K Grade School Math Accuracy ~80% ≥90%
Game of 24 Arithmetic Reasoning Success Rate ~30% ≥90%
HumanEval Code Generation pass@1 ~67% ≥80%
MMLU Knowledge Accuracy ~70% ≥80%

Benchmark Implementation

Game of 24

Installs
1
GitHub Stars
7
First Seen
3 days ago
prompt-benchmark — manutej/categorical-meta-prompting