multimodal-ai
Warn
Audited by Gen Agent Trust Hub on Sep 24, 2026
Risk Level: MEDIUMCOMMAND_EXECUTIONINDIRECT_PROMPT_INJECTIONDYNAMIC_EXECUTION
Full Analysis
- [COMMAND_EXECUTION]: Several code snippets in
references/sharp_edges.md(e.g.,convertToSupportedFormat,normalizeAudioForTranscription, andextractVideoFrames) usechild_process.execto run shell commands likeffmpegandmkdir. The variablesinputPath,outputPath, andoutputDirare interpolated directly into the command string. If these paths are derived from user-supplied filenames, an attacker could execute arbitrary commands by including shell metacharacters (e.g.,;,&,|, or backticks) in the filename. - [INDIRECT_PROMPT_INJECTION]: The skill exhibits multiple vulnerabilities to indirect prompt injection across its core patterns in
references/patterns.md. TheUnifiedMultimodalPipeline.processmethod,extractStructuredDatafunction, andextractDocumentfunction all interpolate untrusted external data (such asinstruction,schema, andfieldNames) directly into the prompts sent to OpenAI and Anthropic models. - Ingestion points: Inputs to
UnifiedMultimodalPipeline.process,extractStructuredData, andextractDocumentinreferences/patterns.md. - Boundary markers: None. No delimiters or 'ignore' instructions are used to separate user data from system instructions.
- Capability inventory: The skill includes file system read/write access (
fs) and shell execution (child_process.exec). - Sanitization: Absent. Data is interpolated using standard template literals without escaping or validation.
- [DYNAMIC_EXECUTION]: The skill uses dynamic
requirecalls within methods (e.g.,UnifiedMultimodalPipeline.transcribeAudioinreferences/patterns.mdandvalidateAudioForWhisperinreferences/sharp_edges.md). While loading standard modules likefs, this pattern of deferred, dynamic loading is a hallmark of techniques used to hide dependencies or perform runtime code assembly.
Audit Metadata