rules-distill

Pass

Audited by Gen Agent Trust Hub on Sep 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill scans the full content of other installed skills to propose new rules. If a skill contains malicious instructions, it could trick the LLM into suggesting unsafe rule changes. Ingestion points: scripts/scan-skills.sh extracts text from all SKILL.md files. Boundary markers: The LLM instructions use simple headers but lack strict separation to prevent the LLM from following instructions found inside the scanned skills. Capability inventory: The skill generates drafts that append, revise, or create new rule files in ~/.claude/rules/. Sanitization: There is no filtering of skill content prior to analysis. Safety relies on the requirement that principles appear in multiple skills and the final human approval step.
  • [COMMAND_EXECUTION]: The skill runs shell scripts to read file metadata and headings from the local filesystem. Evidence: SKILL.md calls scan-skills.sh and scan-rules.sh located in the skill directory.
  • [DATA_EXFILTRATION]: The skill reads the contents of the user's agent configuration and rule files located in ~/.claude/rules/ to perform its analysis.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 1, 2026, 02:39 AM
Security Audit — agent-trust-hub — rules-distill