skills/google/capsem/dev-benchmark/Gen Agent Trust Hub

dev-benchmark

Pass

Audited by Gen Agent Trust Hub on Sep 26, 2026

Risk Level: SAFECOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • Local Command Execution: The skill interacts with the system by executing just tasks and a specialized Rust binary (capsem-bench-rs). These tools are used to orchestrate benchmark runs, verify machine fitness for measurement, and generate performance reports.
  • Subprocess Interaction: The skill utilizes "collectors," which are standalone executables designed to print JSON metrics to standard output. The benchmarking engine executes these subprocesses and captures their output to compute statistics.
  • File System Interaction: To measure storage and I/O performance, the skill accesses various system paths including /tmp, /var/log, and /run. It also manages a local SQLite database located at cache/target/tests/benchmarks/benchmarks.db to store and query historical evidence.
  • Data Ingestion Surface: The skill processes structured JSON data from external collectors and Criterion benchmark samples. While this represents a surface for indirect prompt injection, the skill employs protocol rules (such as line limits and schema enforcement) to manage how data is ingested for analysis.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 26, 2026, 06:44 PM
Security Audit — agent-trust-hub — dev-benchmark