project-evals
Pass
Audited by Gen Agent Trust Hub on Aug 22, 2026
Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: The skill documents the use of a local development tool via the command
node ./bin/gd.ts dev <path-to-guide-directory>. This tool is used to generate test files and execute them as part of an evaluation pipeline. This is a functional component of the documented development process. - [DYNAMIC_EXECUTION]: The instructions describe a process where natural language expectations in
expectations.mdare used to dynamically generategrader.ts(Playwright test files) which are subsequently executed. This code generation and execution is the primary intended purpose of the evaluation tool. - [INDIRECT_PROMPT_INJECTION]: The skill defines a surface for processing untrusted data, as the eval harness ingests prompts from
tasks/task.mdand assertions fromexpectations.md. - Ingestion points: The tool reads content from
expectations.mdandtasks/task.md(located in the guide directory). - Boundary markers: None explicitly defined in the provided markdown instructions.
- Capability inventory: The tool
gd.tshas the capability to generate and execute JavaScript/TypeScript code (Playwright tests) based on these inputs. - Sanitization: The skill mentions that a human may manually edit the generated tests if needed, providing a manual review step.
Audit Metadata