nightly-eval-investigation
Pass
Audited by Gen Agent Trust Hub on Aug 22, 2026
Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes shell commands via Node.js
child_process.execSyncto facilitate its diagnostic workflow. - The commands utilized include
gcloud storage ls,gcloud storage rsync, and severalghCLI subcommands (issue create, project item-add, api graphql). - Execution is used exclusively for synchronizing evaluation artifacts and generating diagnostic reports.
- Inputs for these commands are primarily derived from system state or restricted through regular expression filtering, limiting the risk of arbitrary command injection.
- [EXTERNAL_DOWNLOADS]: The skill synchronizes diagnostic data from a remote Google Cloud Storage bucket.
- Evaluation runs are retrieved from
gs://guidance-evals/usinggcloud storage rsyncto the localharness/results/directory. - The downloaded content consists of JSON evaluation metadata used for local analysis scripts.
- [PROMPT_INJECTION]: The skill possesses a surface for indirect prompt injection as it processes untrusted markdown content from the
guides/directory. - Ingestion points: The
extractTaskPromptandextractGuideDescriptionfunctions ininvestigate.tsread raw markdown content from the filesystem. - Boundary markers: The
SKILL.mdinstructions define aCRITICAL CONSTRAINTand explicit workflow rules that prevent the agent from modifying the source evaluation files, enforcing a read-only diagnostic posture. - Capability inventory: The skill has access to the filesystem and CLI tools (
gcloud,gh) for report generation and synchronization. - Sanitization: Content is extracted and embedded into markdown reports for human review; the processed data is not executed as shell commands or code.
Audit Metadata