sparse-autoencoder-training

Pass

Audited by Gen Agent Trust Hub on Sep 17, 2026

Risk Level: SAFEMETADATA_POISONINGINDIRECT_PROMPT_INJECTION
Full Analysis
  • [SAFE]: The skill documentation and associated Python scripts follow standard practices for mechanistic interpretability research. All dependencies (sae-lens, transformer-lens, torch) and external resources (HuggingFace, GitHub, W&B) are appropriate and correctly referenced. No malicious patterns or behaviors were detected in the instructions or scripts.\n- [METADATA_POISONING]: There is a discrepancy between the platform-reported author ('firecrawl') and the author stated in the skill's YAML frontmatter ('Orchestra Research'). This mismatch is documented here for transparency but does not appear to indicate malicious intent or deceptive practices regarding the skill's capabilities.\n- [INDIRECT_PROMPT_INJECTION]: The skill involves processing external text datasets and user-provided prompts during analysis and feature steering. While this is an inherent surface for untrusted data, the capabilities are limited to model training and inference within a research context.\n
  • Ingestion points: Dataset paths in SKILL.md and interactive prompts in references/tutorials.md.\n
  • Boundary markers: None explicitly used for data interpolation.\n
  • Capability inventory: Limited to tensor operations and model generation using established ML libraries.\n
  • Sanitization: Standard tokenization is used; no specialized sanitization for embedded instructions is present.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 17, 2026, 07:53 PM
Security Audit — agent-trust-hub — sparse-autoencoder-training