dev-benchmark
Pass
Audited by Gen Agent Trust Hub on Sep 26, 2026
Risk Level: SAFECOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- Local Command Execution: The skill interacts with the system by executing
justtasks and a specialized Rust binary (capsem-bench-rs). These tools are used to orchestrate benchmark runs, verify machine fitness for measurement, and generate performance reports. - Subprocess Interaction: The skill utilizes "collectors," which are standalone executables designed to print JSON metrics to standard output. The benchmarking engine executes these subprocesses and captures their output to compute statistics.
- File System Interaction: To measure storage and I/O performance, the skill accesses various system paths including
/tmp,/var/log, and/run. It also manages a local SQLite database located atcache/target/tests/benchmarks/benchmarks.dbto store and query historical evidence. - Data Ingestion Surface: The skill processes structured JSON data from external collectors and Criterion benchmark samples. While this represents a surface for indirect prompt injection, the skill employs protocol rules (such as line limits and schema enforcement) to manage how data is ingested for analysis.
Audit Metadata