speech-to-text
Pass
Audited by Gen Agent Trust Hub on Sep 30, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONREMOTE_CODE_EXECUTIONEXTERNAL_DOWNLOADS
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted audio and video data from various sources, including file uploads and remote URLs. This input serves as a potential vector for indirect prompt injection, where malicious instructions spoken in the media could affect the agent's subsequent actions.
- Ingestion points: Transcription methods accepting
file,source_url,cloud_storage_url, and theaudio_base_64parameter in real-time streaming (SKILL.md, references/transcription-options.md). - Boundary markers: The provided code examples do not demonstrate the use of delimiters or specific instructions to the agent to ignore commands within the transcribed text.
- Capability inventory: While the skill primarily outputs text, the agent consuming this text may have extensive file system, network, or tool access.
- Sanitization: No sanitization or filtering of transcribed text is implemented in the reference code.
- [REMOTE_CODE_EXECUTION]: The documentation includes a command to install the vendor's CLI via a shell script downloaded from GitHub and piped directly to the shell.
- Evidence:
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/elevenlabs/cli/releases/latest/download/elevenlabs-cli-installer.sh | sh(references/installation.md). - Context: The resource is hosted on the official GitHub repository of the skill's author, which is a standard distribution method for this vendor.
- [EXTERNAL_DOWNLOADS]: The skill documentation guides users to install various software packages and utilities from external registries and repositories.
- Evidence: References to official NPM packages (
@elevenlabs/elevenlabs-js), PyPI packages (elevenlabs), and platform-specific package managers like Homebrew and Scoop (references/installation.md). - Context: All external dependencies are official vendor-maintained resources.
Audit Metadata