transcribe-video

Pass

Audited by Gen Agent Trust Hub on Jun 18, 2026

Risk Level: SAFECOMMAND_EXECUTIONEXTERNAL_DOWNLOADSPROMPT_INJECTION
Full Analysis
  • [COMMAND_EXECUTION]: The skill executes local system commands using Node.js spawnSync and Python subprocess calls to manage Docker containers, extract audio with ffmpeg, and run the transcription binary.
  • These commands are implemented using argument lists rather than shell strings, which provides protection against standard command injection.
  • The Docker container is configured with the --restart unless-stopped flag, which acts as a persistence mechanism by ensuring the service continues running across system reboots.
  • [EXTERNAL_DOWNLOADS]: The skill fetches necessary software and data components from well-known technology platforms.
  • Clones the whisper.cpp source code from GitHub during the initial Docker image build.
  • Downloads pre-trained machine learning models from HuggingFace's official repository.
  • [PROMPT_INJECTION]: The skill has a surface for indirect prompt injection because it converts untrusted audio/video content into text that is then processed by the agent.
  • Ingestion points: Reads user-supplied video and audio files from the local filesystem via scripts/transcribe.mjs.
  • Boundary markers: The output files (transcript.txt and transcript.srt) are provided to the agent as plain text without delimiters or instructions to ignore embedded commands.
  • Capability inventory: The skill can execute local shell commands (docker, ffmpeg) and perform file system operations.
  • Sanitization: While the script performs basic filename sanitization to prevent path traversal, it does not validate or filter the spoken content for malicious instructions directed at the agent.
Audit Metadata
Risk Level
SAFE
Analyzed
Jun 18, 2026, 10:11 PM
Security Audit — agent-trust-hub — transcribe-video