d-id-avatar-video
D-ID avatar video production
Use this skill when the selected provider is D-ID or when the user asks for a D-ID talking avatar, photo-to-speaking-video, presenter clip, expressive avatar, interactive visual agent, real-time avatar stream, or D-ID-powered localization/support/training workflow.
Do not treat D-ID as a general cinematic video generator. It is primarily for human-presenter video: a face or presenter speaks text or audio with lip sync. For scenes that need object interaction, product handling, full-body action choreography, multi-shot narrative motion, or cinematic camera movement, use a broader video-generation or composition pipeline and use D-ID only for the presenter segments.
Facts in this skill were verified against official D-ID documentation and policies on 2026-07-10. Re-check live docs before quoting prices, enabled models, supported languages, moderation categories, endpoint schemas, or plan limits in user-facing commitments.
Choose the D-ID route before scripting
Pick the route from the deliverable, not from habit: