transformer-lens-interpretability
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFE
Full Analysis
- [SAFE]: The skill provides legitimate guidance and documentation for mechanistic interpretability research using the TransformerLens library. All analyzed components align with standard research and development practices.- [EXTERNAL_DOWNLOADS]: The skill provides instructions for installing the
transformer-lenslibrary via pip from PyPI and the official GitHub repository (TransformerLensOrg/TransformerLens). These are well-known and reputable sources for this research tool.- [CREDENTIALS_UNSAFE]: The documentation includes the use ofos.environ["HF_TOKEN"] = "your_token". This uses a clear placeholder string and is standard practice for authenticating with Hugging Face to access gated models.- [INDIRECT_PROMPT_INJECTION]: The skill's primary function involves processing text prompts and model activations for analysis. While this represents a data ingestion surface, it is a core requirement of the library's research purpose and does not pose an atypical security risk.- [DYNAMIC_EXECUTION]: The skill utilizes HookPoints to execute Python functions during model inference. The provided examples (ablation, patching, and steering) demonstrate legitimate and safe use of this feature for research analysis.
Audit Metadata