benchmark-models
Pass
Audited by Gen Agent Trust Hub on Sep 28, 2026
Risk Level: SAFECOMMAND_EXECUTIONDATA_EXFILTRATIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill invokes several local binaries to manage its workflow.
- Executes
gstack-skill-startin the preamble to initialize the session and sync artifacts. - Executes
gstack-model-benchmarkto perform the actual cross-model comparisons across providers. - Executes
gstack-learnings-logto record operational insights to the local filesystem. - Executes
gstack-skill-endto handle telemetry and cleanup at the end of the workflow. - [DATA_EXFILTRATION]: The skill accesses sensitive local credential files to determine feature availability.
- Evidence: In
Step 3, it searches for the string 'ANTHROPIC' within~/.claude/.credentials.jsonto check if the quality judge feature (which requires an API key) can be enabled. - [INDIRECT_PROMPT_INJECTION]: The skill reads and processes external data, which is a potential vector for indirect prompt injection attacks.
- Ingestion points: User input for inline prompts (
Step 1B), local skill documentation files (Step 1A), and specific prompt files on the filesystem (Step 1C). - Boundary markers: None identified in the prompt construction logic.
- Capability inventory: The skill performs network operations (via the benchmark binary) and file system writes.
- Sanitization: No explicit sanitization or validation of the external prompt content is implemented before it is passed to the models.
Audit Metadata