pyvene-interventions

Pass

Audited by Gen Agent Trust Hub on Sep 9, 2026

Risk Level: SAFE
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill fetches the pyvene library from PyPI and demonstrates loading pre-trained interventions from the HuggingFace Hub. These are standard operations within the AI research ecosystem using well-known, reputable platforms.
  • [REMOTE_CODE_EXECUTION]: Documents the use of the IntervenableModel.load() method to retrieve and apply intervention configurations from HuggingFace. This is a core feature of the pyvene library designed for experiment reproducibility and sharing steering directions.
  • [INDIRECT_PROMPT_INJECTION]: The skill processes text prompts to conduct causal analysis. While these inputs are technically untrusted data, the skill follows standard interpretability practices and does not expose dangerous capabilities or execution environments that could be exploited via injection.
Audit Metadata
Risk Level
SAFE
Analyzed
Sep 9, 2026, 07:07 PM
Security Audit — agent-trust-hub — pyvene-interventions