learning-hardened
Pass
Audited by Gen Agent Trust Hub on Apr 21, 2026
Risk Level: SAFE
Full Analysis
- [PROMPT_INJECTION]: The skill incorporates defensive instructions intended to protect against adversarial prompts, explicitly directing the agent to ignore attempts to override safety boundaries using authority claims or pseudo-system messages.
- [DATA_EXFILTRATION]: The skill implements a persistent behavioral profile by updating its own preference sections in SKILL.md. This is constrained by mandatory guardrails that prohibit the recording of sensitive personal data such as health diagnoses, disabilities, or demographic attributes, ensuring the data remains focused on learning style rather than personal identity.
- [INDIRECT_PROMPT_INJECTION]: The skill includes a potential surface for indirect injection as it evolves based on user interactions. Ingestion points: Behavioral patterns extracted from user chat sessions. Boundary markers: Explicit instructions in SKILL.md and criteria.md defining data exclusion and signal thresholds. Capability inventory: Permission for the agent to modify its own preference sections in SKILL.md. Sanitization: Strict guidelines to record only behavioral traits (e.g., 'visual learner') while rejecting sensitive personal disclosures.
Audit Metadata