sparse-autoencoder-training
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill facilitates the download of the
sae-lenslibrary and its dependencies via standard package managers. It also includes functions to fetch pre-trained Sparse Autoencoder weights and configuration files from HuggingFace repositories. - [COMMAND_EXECUTION]: The documentation provides Python code snippets for training models, analyzing activations, and performing feature steering, as well as shell commands for library installation.
- [INDIRECT_PROMPT_INJECTION]: The skill processes arbitrary text inputs to analyze model activations. As is standard for mechanistic interpretability research, this creates an ingestion surface where external data is converted into model tokens. Mandatory evidence:
- Ingestion points: Text prompts are passed to
model.to_tokens()inSKILL.mdandreferences/tutorials.md. - Boundary markers: None present; prompts are processed directly as is typical for activation analysis.
- Capability inventory: Includes local file writing (
sae.save_model) and network operations (Weights & Biases logging, HuggingFace uploads). - Sanitization: No specific sanitization or filtering of input text is described.
- [DATA_EXFILTRATION]: The skill describes integrated support for Weights & Biases (
wandb) for experiment tracking and provides methods for uploading models to HuggingFace. These are standard features for collaborative ML development and research.
Audit Metadata