voice-ai-integration
Pass
Audited by Gen Agent Trust Hub on Sep 16, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill architecture converts audio input to text and passes it to a Large Language Model (LLM) for processing. This creates a vulnerability where spoken instructions in the audio stream could maliciously influence the AI's behavior.
- Ingestion points:
VoiceAssistant.process_voice_inputinexamples/voice_assistant.pytranscribes audio files or streams into text strings. - Boundary markers: The skill lacks delimiters or "ignore instructions" warnings when interpolating the transcribed user speech into the
conversation_historyused for LLM response generation. - Capability inventory: The skill manages conversation state and interacts with multiple external APIs, which could be exploited through successful prompt injection.
- Sanitization: No filtering or validation is performed on the transcribed text before it is processed by the AI.
- [EXTERNAL_DOWNLOADS]: The skill communicates with external APIs to process audio data, including AssemblyAI (
api.assemblyai.com) for transcription and ElevenLabs (api.elevenlabs.io) for speech synthesis. These are established services for voice-related tasks.
Audit Metadata