sparse-autoencoder-training
Pass
Audited by Gen Agent Trust Hub on Sep 9, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill describes workflows that ingest external text data to analyze neural network activations. While this creates a surface for indirect prompt injection, it is the intended primary purpose of the skill and follows standard research practices.
- Ingestion points: User-provided prompts and dataset paths in SKILL.md and references/tutorials.md.
- Boundary markers: None explicitly demonstrated in code snippets.
- Capability inventory: Documents model generation, local file saving, and optional HuggingFace uploads.
- Sanitization: Not applicable for activation analysis workflows.
Audit Metadata