sparse-autoencoder-training

Pass

Audited by Gen Agent Trust Hub on Sep 9, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill describes workflows that ingest external text data to analyze neural network activations. While this creates a surface for indirect prompt injection, it is the intended primary purpose of the skill and follows standard research practices.
  • Ingestion points: User-provided prompts and dataset paths in SKILL.md and references/tutorials.md.
  • Boundary markers: None explicitly demonstrated in code snippets.
  • Capability inventory: Documents model generation, local file saving, and optional HuggingFace uploads.
  • Sanitization: Not applicable for activation analysis workflows.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 9, 2026, 07:07 PM
Security Audit — agent-trust-hub — sparse-autoencoder-training