rules-distill
Pass
Audited by Gen Agent Trust Hub on Sep 1, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONCOMMAND_EXECUTIONDATA_EXFILTRATION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill scans the full content of other installed skills to propose new rules. If a skill contains malicious instructions, it could trick the LLM into suggesting unsafe rule changes. Ingestion points:
scripts/scan-skills.shextracts text from allSKILL.mdfiles. Boundary markers: The LLM instructions use simple headers but lack strict separation to prevent the LLM from following instructions found inside the scanned skills. Capability inventory: The skill generates drafts that append, revise, or create new rule files in~/.claude/rules/. Sanitization: There is no filtering of skill content prior to analysis. Safety relies on the requirement that principles appear in multiple skills and the final human approval step. - [COMMAND_EXECUTION]: The skill runs shell scripts to read file metadata and headings from the local filesystem. Evidence:
SKILL.mdcallsscan-skills.shandscan-rules.shlocated in the skill directory. - [DATA_EXFILTRATION]: The skill reads the contents of the user's agent configuration and rule files located in
~/.claude/rules/to perform its analysis.
Audit Metadata