voice-ai-engine-development
Pass
Audited by Gen Agent Trust Hub on Sep 14, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill implements a pipeline that transcribes user audio and passes it directly to an LLM agent, creating a surface for indirect prompt injection. 1. Ingestion points: User audio is received via WebSocket in examples/complete_voice_engine.py and processed by the DeepgramTranscriber. 2. Boundary markers: The GeminiAgent uses a System Instruction prefix in its conversation history construction, but the implementation lacks explicit delimiters or warnings to ignore embedded instructions in the transcribed user text. 3. Capability inventory: The agent has the ability to generate audio responses and transmit them back to the user via WebSocket connections. 4. Sanitization: There is no evidence of text-based filtering or validation applied to the transcription before it is ingested by the LLM.
- [EXTERNAL_DOWNLOADS]: The skill documentation and implementation templates reference several well-known and trusted service providers for transcription, LLM, and TTS services, including Google, OpenAI, Deepgram, and ElevenLabs. These references are essential for the skill's functionality and target official API endpoints.
Audit Metadata