project-evals

Pass

Audited by Gen Agent Trust Hub on Aug 22, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill documents the use of a local development tool via the command node ./bin/gd.ts dev <path-to-guide-directory>. This tool is used to generate test files and execute them as part of an evaluation pipeline. This is a functional component of the documented development process.
  • [DYNAMIC_EXECUTION]: The instructions describe a process where natural language expectations in expectations.md are used to dynamically generate grader.ts (Playwright test files) which are subsequently executed. This code generation and execution is the primary intended purpose of the evaluation tool.
  • [INDIRECT_PROMPT_INJECTION]: The skill defines a surface for processing untrusted data, as the eval harness ingests prompts from tasks/task.md and assertions from expectations.md.
  • Ingestion points: The tool reads content from expectations.md and tasks/task.md (located in the guide directory).
  • Boundary markers: None explicitly defined in the provided markdown instructions.
  • Capability inventory: The tool gd.ts has the capability to generate and execute JavaScript/TypeScript code (Playwright tests) based on these inputs.
  • Sanitization: The skill mentions that a human may manually edit the generated tests if needed, providing a manual review step.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 22, 2026, 05:58 PM
Security Audit — agent-trust-hub — project-evals