transcribe-video
Pass
Audited by Gen Agent Trust Hub on Jun 18, 2026
Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
- [COMMAND_EXECUTION]: The skill executes local system commands using Node.js spawnSync and Python subprocess calls to manage Docker containers, extract audio with ffmpeg, and run the transcription binary.
- These commands are implemented using argument lists rather than shell strings, which provides protection against standard command injection.
- The Docker container is configured with the --restart unless-stopped flag, which acts as a persistence mechanism by ensuring the service continues running across system reboots.
- [EXTERNAL_DOWNLOADS]: The skill fetches necessary software and data components from well-known technology platforms.
- Clones the whisper.cpp source code from GitHub during the initial Docker image build.
- Downloads pre-trained machine learning models from HuggingFace's official repository.
- [PROMPT_INJECTION]: The skill has a surface for indirect prompt injection because it converts untrusted audio/video content into text that is then processed by the agent.
- Ingestion points: Reads user-supplied video and audio files from the local filesystem via scripts/transcribe.mjs.
- Boundary markers: The output files (transcript.txt and transcript.srt) are provided to the agent as plain text without delimiters or instructions to ignore embedded commands.
- Capability inventory: The skill can execute local shell commands (docker, ffmpeg) and perform file system operations.
- Sanitization: While the script performs basic filename sanitization to prevent path traversal, it does not validate or filter the spoken content for malicious instructions directed at the agent.
Audit Metadata