blip-2-vision-language

Pass

Audited by Gen Agent Trust Hub on Oct 1, 2026

Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
  • [INDIRECT_PROMPT_INJECTION]: The skill provides instructions and code for processing untrusted external data (images and text questions) through a large language model, which introduces a surface for indirect prompt injection.
  • Ingestion points: Untrusted data is ingested into the agent context through local file reads (e.g., Image.open("photo.jpg") in SKILL.md), URL-based image loading, and API endpoints defined in the FastAPI server example (e.g., UploadFile in references/advanced-usage.md).
  • Boundary markers: The implementation uses basic string formatting to delimit inputs (e.g., "Question: {question} Answer:" in SKILL.md), which provides minimal isolation against adversarial instructions embedded in the data.
  • Capability inventory: The skill's capabilities are primarily restricted to model inference and feature extraction; it does not contain high-risk functions for arbitrary file system modifications or exfiltration of system credentials.
  • Sanitization: There is no implementation of input sanitization, text filtering, or image validation logic to mitigate potential prompt injection attacks before data reaches the LLM.
Audit Metadata
Risk Level
SAFE
Analyzed
Oct 1, 2026, 07:50 AM
Security Audit — agent-trust-hub — blip-2-vision-language