speech-to-text

Pass

Audited by Gen Agent Trust Hub on Aug 8, 2026

Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
  • [EXTERNAL_DOWNLOADS]: The skill refers to external resources for installation and functionality, specifically directing users to download the belt CLI tool and associated skills from the official inference.sh GitHub organization and service domains.
  • [PROMPT_INJECTION]: The skill includes functionality that processes content from untrusted external sources, specifically audio and video URLs, creating a surface for indirect prompt injection.
  • Ingestion points: Media URLs provided to the audio_url and video_url parameters within belt app run command examples in SKILL.md.
  • Boundary markers: None present; the skill does not instruct the agent to distinguish between transcribed content and system instructions.
  • Capability inventory: Transcribes audio/video to text, translates speech, and generates video captions via the belt CLI.
  • Sanitization: The skill does not implement validation or filtering of the resulting transcription before it is ingested into the agent context.
Audit Metadata
Risk Level
SAFE
Analyzed
Aug 8, 2026, 08:29 PM
Security Audit — agent-trust-hub — speech-to-text