transformer-lens-interpretability

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFE
Full Analysis
  • [SAFE]: The skill provides legitimate guidance and documentation for mechanistic interpretability research using the TransformerLens library. All analyzed components align with standard research and development practices.- [EXTERNAL_DOWNLOADS]: The skill provides instructions for installing the transformer-lens library via pip from PyPI and the official GitHub repository (TransformerLensOrg/TransformerLens). These are well-known and reputable sources for this research tool.- [CREDENTIALS_UNSAFE]: The documentation includes the use of os.environ["HF_TOKEN"] = "your_token". This uses a clear placeholder string and is standard practice for authenticating with Hugging Face to access gated models.- [INDIRECT_PROMPT_INJECTION]: The skill's primary function involves processing text prompts and model activations for analysis. While this represents a data ingestion surface, it is a core requirement of the library's research purpose and does not pose an atypical security risk.- [DYNAMIC_EXECUTION]: The skill utilizes HookPoints to execute Python functions during model inference. The provided examples (ablation, patching, and steering) demonstrate legitimate and safe use of this feature for research analysis.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — transformer-lens-interpretability