speech-to-text
Pass
Audited by Gen Agent Trust Hub on Sep 17, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTION
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted audio files to produce text transcripts, creating a surface where malicious instructions embedded in audio could influence the agent's behavior.
- Ingestion points: Audio data is read from the filesystem in
scripts/stt.pyand uploaded to the transcription API. - Boundary markers: None; the skill returns the raw transcript text or JSON structured data without explicit delimiters to the agent.
- Capability inventory: The skill utilizes
networkaccess to communicate with the transcription service andfilesystemaccess to read source audio, write output transcripts, and manage its configuration file. - Sanitization: No sanitization is performed on the audio binary or the resulting transcript string.
- [SAFE]: All network operations are directed to the vendor's official API endpoint (
noiz.ai). The skill follows security best practices for local credential storage, using a configuration file with restrictive0600permissions to prevent unauthorized access by other local users.
Audit Metadata