leaderboard
Pass
Audited by Gen Agent Trust Hub on Aug 26, 2026
Risk Level: SAFECOMMAND_EXECUTIONPROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes multiple shell commands to manage the leaderboard workflow.
- Uses
git logandgit diffto identify new submissions based on the last generation of the benchmarks page. - Uses
uv sync --lockedto manage the Python virtual environment. - Runs local Python modules via
uv run python -m nurb_evals.reportanduv run python -m nurb_evals.siteto regenerate the site and reports. - Executes tests using
uv run pytest -qto verify the state of the repository before publishing. - [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted data from the
evals/submissions/directory, which contains content provided by external contributors. - Ingestion points: Reads
results.jsonl, trial transcripts, and part sources fromevals/submissions/(SKILL.md Step 1 and Step 2). - Boundary markers: None identified. The agent is explicitly instructed to "Read one transcript" per contributor to verify the session content (Step 2).
- Capability inventory: The skill has significant capabilities including file writing (
REPORT.md,benchmarks.html), local code execution viauv run, and the ability to open pull requests (Step 4). - Sanitization: The instructions include sanity checks for sensitive file paths (e.g.,
/Users/,/home/) and usernames, but do not provide specific sanitization or escaping for LLM instructions embedded within the transcripts. - [DYNAMIC_EXECUTION]: The skill involves rebuilding and grading projects based on submitted content.
- Step 2 instructs the agent to "rebuild the trial project" by materializing tasks and dropping in the submitted "part" (likely code), followed by running a grader. This involves executing logic that may be influenced or directly provided by the untrusted submission content.
Audit Metadata