speech-to-text
Pass
Audited by Gen Agent Trust Hub on Jun 19, 2026
Risk Level: SAFEPROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [PROMPT_INJECTION]: The skill processes audio from external URLs and utilizes the resulting transcripts in downstream tasks like video captioning. This creates an indirect prompt injection surface where malicious instructions embedded in audio could influence the agent's behavior in subsequent steps.
- Ingestion points:
audio_urlandvideo_urlparameters in SKILL.md. - Boundary markers: Absent.
- Capability inventory:
belt app run(execution of remote inference applications). - Sanitization: None documented for processed audio content.
- [EXTERNAL_DOWNLOADS]: The documentation provides links to fetch configuration and installation instructions from the inference-sh official GitHub repository.
Audit Metadata