multimodal-ai

Warn

Audited by Gen Agent Trust Hub on Sep 24, 2026

Risk Level: MEDIUMCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
  • [COMMAND_EXECUTION]: Several code snippets in references/sharp_edges.md (e.g., convertToSupportedFormat, normalizeAudioForTranscription, and extractVideoFrames) use child_process.exec to run shell commands like ffmpeg and mkdir. The variables inputPath, outputPath, and outputDir are interpolated directly into the command string. If these paths are derived from user-supplied filenames, an attacker could execute arbitrary commands by including shell metacharacters (e.g., ;, &, |, or backticks) in the filename.
  • [INDIRECT_PROMPT_INJECTION]: The skill exhibits multiple vulnerabilities to indirect prompt injection across its core patterns in references/patterns.md. The UnifiedMultimodalPipeline.process method, extractStructuredData function, and extractDocument function all interpolate untrusted external data (such as instruction, schema, and fieldNames) directly into the prompts sent to OpenAI and Anthropic models.
  • Ingestion points: Inputs to UnifiedMultimodalPipeline.process, extractStructuredData, and extractDocument in references/patterns.md.
  • Boundary markers: None. No delimiters or 'ignore' instructions are used to separate user data from system instructions.
  • Capability inventory: The skill includes file system read/write access (fs) and shell execution (child_process.exec).
  • Sanitization: Absent. Data is interpolated using standard template literals without escaping or validation.
  • [DYNAMIC_EXECUTION]: The skill uses dynamic require calls within methods (e.g., UnifiedMultimodalPipeline.transcribeAudio in references/patterns.md and validateAudioForWhisper in references/sharp_edges.md). While loading standard modules like fs, this pattern of deferred, dynamic loading is a hallmark of techniques used to hide dependencies or perform runtime code assembly.
Audit Metadata
Risk Level
MEDIUM
Analyzed
Sep 24, 2026, 08:22 AM
Security Audit — agent-trust-hub — multimodal-ai