Adding a voice assistant to your website or app
A voice assistant does not have to live on a phone line. The same engine can sit behind a microphone button on your website or inside your product, answering questions, guiding someone through a task or completing it for them while they talk. It removes the telephone from the picture, which makes it cheaper per conversation and lets the assistant use the screen as well as speech. It also raises questions a phone agent never faces: whether people will grant microphone access, what happens to what they say, and whether voice is actually better than typing for the job you have in mind.
How does a voice assistant on a website work?
The browser captures audio from the microphone after the user grants permission and streams it, usually over WebRTC, to the voice service. Speech recognition, the language model and text-to-speech run as they would on a phone call, and the reply streams back and plays in the page. Managed platforms such as Vapi and Retell offer web SDKs for this, and speech-to-speech models such as OpenAI's Realtime API can be reached from the browser directly.
In a mobile app the path is similar, with the app's own microphone permission instead of the browser's. Either way there is no phone number and no telephony charge, which is one reason in-product voice can be cheaper per conversation than a phone line.
The screen changes what the assistant can do. It can show a live transcript so the user sees what it heard, display options instead of reading them out, and open the page or fill the form it is talking about. Voice plus a screen is often better than either alone.
Push-to-talk or always listening?
Push-to-talk for almost every website and most apps. The user presses a button, speaks, and the assistant responds. It makes consent obvious, avoids picking up conversations in the room, and removes the problem of the assistant interrupting when the user pauses to think.
Continuous conversation, where the microphone stays open for the session and the assistant detects when the user has finished speaking, feels more natural and suits hands-free use: someone cooking, driving, working with their hands, or using the product through accessibility tools. It needs good turn detection and a clear visual indicator that the microphone is live.
Always-on listening for a wake word, the smart-speaker model, is rarely right for a business website. It needs persistent permission most users will not grant and raises privacy questions that outweigh the convenience.
| Mode | Best for | Watch out for |
|---|---|---|
| Push-to-talk | Websites, most apps, noisy environments | Feels less conversational |
| Open session with turn detection | Hands-free tasks, accessibility, longer help | Cutting in during pauses; needs a live indicator |
| Wake-word, always listening | Dedicated devices, not websites | Permission, privacy and battery |
What about privacy and the microphone?
Ask for the microphone only when the user presses the voice button, never on page load. A permission prompt that appears before someone has chosen to use voice gets refused, and a refused permission is hard to recover because the browser remembers it.
Say what happens to the audio in a sentence next to the button: whether it is recorded or only transcribed, how long transcripts are kept, and which providers process it. Where you process personal data, the usual data protection rules apply, and audio that could reveal health or other sensitive information deserves particular care.
Offer text as an equal alternative, not a fallback. Some users cannot or will not speak, and a voice-only assistant excludes them.
When is text chat the better choice?
When the answer is something the user needs to read, copy or keep: a tracking number, an address, a list of steps, a price comparison. Speaking these aloud is slower and harder to follow than showing them.
When the user is in a quiet or public place, which is where most people browse. Many people who happily talk to a phone agent will not speak to a website in an open office.
And when the job is mostly typing anyway. Voice wins where hands are busy, where the user struggles with typing or reading, where the conversation is exploratory, or where the product itself is used by voice. If none of those apply, a good text assistant may serve your users better, and the same knowledge base can power both.
What makes in-app voice feel fast?
The same things that make a phone agent feel fast, with one advantage: you control more of the path. Keep the response under a second where you can, stream the reply so speech starts before the whole answer is generated, and show a visible "thinking" state so a short pause reads as work rather than failure.
Let the user interrupt. If they start speaking while the assistant is talking, it should stop and listen, as a person would. Assistants that talk over their users are the most common reason people abandon voice features.
And keep answers short. A spoken answer longer than a few sentences loses people; offer to show the detail on screen instead.
Common questions
- Can I add an AI voice assistant to my website?
- Yes. A microphone button captures the visitor's speech after they grant permission and streams it, usually over WebRTC, to a voice service that transcribes it, generates an answer and speaks it back in the page. Platforms such as Vapi and Retell provide web SDKs, and the assistant can also show a transcript, options or the page it is describing.
- Is a website voice assistant better than a chatbot?
- Sometimes. Voice wins for hands-free tasks, users who find typing or reading hard, exploratory conversations and products used by voice. Text wins when the answer needs to be read or kept, such as a tracking number, and when users are in quiet or public places. Offering both, from the same knowledge base, serves the most people.
- Should a website voice assistant always be listening?
- Almost never. Push-to-talk suits most websites: the microphone opens only when the user presses a button, which makes consent obvious and avoids picking up the room. An open session with turn detection suits hands-free use. Always-on wake-word listening needs persistent permission and is rarely appropriate for a business site.
- When should a website ask for microphone permission?
- Only when the user presses the voice button, never on page load. A prompt shown before someone chooses voice is usually refused, and browsers remember the refusal. Explain beside the button whether audio is recorded or only transcribed, how long it is kept and who processes it.
- How much does an in-app voice assistant cost to run?
- Usually less per conversation than a phone agent, because there is no telephony charge; you pay for speech recognition, the language model and text-to-speech, typically per minute or per token. The larger costs are building it well: turn-taking, interruptions, accessibility and the integrations that let it act on what the user asks.