voice-audio-engineer
Pass
Audited by Gen Agent Trust Hub on Sep 17, 2026
Risk Level: SAFEINDIRECT_PROMPT_INJECTIONEXTERNAL_DOWNLOADS
Full Analysis
- [INDIRECT_PROMPT_INJECTION]: The skill processes untrusted external audio and text data through ElevenLabs and Firecrawl integrations, which creates a potential surface for indirect prompt injection attacks. Any malicious instructions embedded in audio transcripts or text inputs could influence the agent's behavior during processing.
- Ingestion points: Input audio for voice cloning, speech-to-speech transformation, and transcription, as well as text data for speech synthesis.
- Boundary markers: The skill does not specify the use of clear delimiters or instructions to ignore instructions embedded within the processed data.
- Capability inventory: The skill includes file system access (Read, Write, Edit), network operations (WebFetch, Firecrawl), and extensive control over voice generation via the ElevenLabs API.
- Sanitization: No explicit validation or filtering logic for ingested content or transcriptions is mentioned in the skill instructions.
- [EXTERNAL_DOWNLOADS]: The skill utilizes standard, well-known libraries for digital signal processing and machine learning, including numpy, scipy, librosa, and transformers. It also references established academic datasets for speech disfluency research, such as FluencyBank, UCLASS, and SEP-28k. These are reputable resources commonly used in the audio engineering and AI research community.
Audit Metadata