benchmark

Pass

Audited by Gen Agent Trust Hub on Jun 24, 2026

Risk Level: SAFECOMMAND_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill instructs the agent to execute various local shell commands to measure build performance, including cold builds, hot reload (HMR), test suites, TypeScript checks, and Docker builds.
  • [DATA_EXPOSURE_AND_EXFILTRATION]: The skill performs network operations by navigating to target URLs and hitting API endpoints (up to 100 times) to measure browser metrics and latency. While this involves network access, it is performed for the explicit purpose of benchmarking as described.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes external data from browser metrics, API responses, and build logs. This creates a potential surface where untrusted data could influence agent behavior, though the risk is low given the specific use case of performance measurement.
  • [DATA_EXPOSURE_AND_EXFILTRATION]: The skill writes benchmarking results to the local filesystem within the .ecc/benchmarks/ directory to maintain performance baselines.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 24, 2026, 12:57 PM
Security Audit — agent-trust-hub — benchmark