multimodal-voice-and-haptics
Multimodal: Voice, Sound & Haptics
Screens are one channel. Real products speak, vibrate, chime, and listen — in cars, on wrists, across rooms, hands-free. This skill covers designing for every output and input channel a human has, and how to combine them without overwhelming the user. The governing principle: pick the modality that matches the user's context and the task's nature, then layer modalities so each does what it's best at.
Modality strengths — pick the right channel
Every modality has a job it does well and jobs it does badly. Design by matching, not defaulting.
| Modality | Best for | Bad at | Context fit |
|---|---|---|---|
| Visual | Browsing, comparison, precision, dense info, scanning, persistent reference | Hands-busy, eyes-busy, no-screen | Desk, phone-in-hand, TV |
| Voice input | Fast input, hands-busy, eyes-busy, long dictation, search by name, "fire-and-forget" commands | Privacy, noisy rooms, precise editing, browsing options | Car, kitchen, accessibility, walking |
| Voice/audio output | Eyes-free confirmation, short answers, ambient notification, alerts | Long lists, comparison, precise data, privacy | Car, smart speaker, screen reader |
| Haptic | Confirmation, alert, texture/boundary feedback, silent notification | Conveying content, complex info | Wearable, phone-in-pocket, controller |
| Touch/gesture | Direct manipulation, precision, spatial control | Eyes-free, hands-busy, distance | Phone, tablet, touchscreen kiosk |
Rule: combine modalities so each does its strength. A smart-display timer shows remaining minutes (visual: precise), confirms "Timer set" (voice: eyes-free), and chimes when done (audio: ambient). Don't make one channel do everything.
Eyes-free vs hands-free are different constraints. Driving is eyes-busy (can glance briefly, can't read). Cooking is hands-busy (can look, can't touch). Walking-with-coffee is both. Identify which the user has before choosing.