pyvene-interventions
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill facilitates the installation of the
pyvenelibrary from PyPI and demonstrates loading intervention models from the HuggingFace Hub (e.g.,zhengxuanzenwu/intervenable_honest_llama2_chat_7B). These actions involve downloads from well-known services essential for the skill's stated purpose. - [INDIRECT_PROMPT_INJECTION]: The tutorials and code examples involve processing user-supplied text strings (e.g.,
clean_prompt,corrupted_prompt) to analyze model activations. This represents a data ingestion surface common in model interpretability tasks. - [DYNAMIC_EXECUTION]: The skill utilizes
pv.IntervenableModel.load()to import pre-trained intervention parameters from external sources. This involves deserializing model weights and configurations, which is a standard component of PyTorch-based machine learning workflows.
Audit Metadata