pyvene-interventions

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEEXTERNAL_DOWNLOADSINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill facilitates the installation of the pyvene library from PyPI and demonstrates loading intervention models from the HuggingFace Hub (e.g., zhengxuanzenwu/intervenable_honest_llama2_chat_7B). These actions involve downloads from well-known services essential for the skill's stated purpose.
  • [INDIRECT_PROMPT_INJECTION]: The tutorials and code examples involve processing user-supplied text strings (e.g., clean_prompt, corrupted_prompt) to analyze model activations. This represents a data ingestion surface common in model interpretability tasks.
  • [DYNAMIC_EXECUTION]: The skill utilizes pv.IntervenableModel.load() to import pre-trained intervention parameters from external sources. This involves deserializing model weights and configurations, which is a standard component of PyTorch-based machine learning workflows.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — pyvene-interventions