j-rig-eval

Pass

Audited by Gen Agent Trust Hub on Sep 9, 2026

Risk Level: SAFECOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill invokes the j-rig CLI tool via Bash to perform skill evaluations. It specifically allows for the execution of commands defined within a local evaluation specification (eval-spec.yaml) when the --run-self-test flag is enabled. The skill documentation provides explicit warnings that these commands are controlled by the target skill and should be reviewed for safety before execution.
  • [INDIRECT_PROMPT_INJECTION]: As an evaluation tool, the skill is designed to ingest and process untrusted data from other skill directories. This creates an indirect prompt injection surface where a malicious skill under test could attempt to subvert the evaluator's logic. However, the skill explicitly includes test criteria (e.g., no-prompt-leakage) to measure and mitigate these risks in the target content.
  • [PROMPT_INJECTION]: A static analysis flag detected prompt injection patterns in eval.yaml. Upon manual review, these patterns (e.g., "Ignore all previous instructions") are identified as benign adversarial test cases used to verify the security of the skill being evaluated, rather than attempts to hijack the agent's current session.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 9, 2026, 11:21 AM
Security Audit — agent-trust-hub — j-rig-eval