Vapi, Retell or a custom pipeline: choosing a voice stack
Every voice agent is built from the same parts: a phone line or browser connection, speech recognition, a language model, a synthetic voice, and something that orchestrates them fast enough to hold a conversation. The choice is how much of that you assemble yourself. Managed platforms such as Vapi and Retell AI do the orchestration for you and get an agent live in days. A custom pipeline, typically Twilio for telephony with OpenAI's models and specialist providers such as Deepgram and ElevenLabs, gives you more control over latency, cost and data at the price of more engineering. Neither is better in general. The right answer depends on your volume, your integrations and how much control your use case demands.
What are the parts of a voice agent stack?
Telephony connects the call: a phone number and the audio stream, from a carrier or a provider such as Twilio, or a WebRTC connection for a website. Speech-to-text turns audio into words; Deepgram is a common choice for its speed and accent handling. The language model decides what to say and which tools to call. Text-to-speech turns the reply into audio; ElevenLabs is known for natural voices.
Orchestration is the part people underestimate. It decides when the caller has finished speaking, handles interruptions, streams partial results so speech starts before the full reply exists, manages tool calls to your calendar or CRM, and keeps the round trip short enough to feel like a conversation. This is most of what a managed platform sells.
Speech-to-speech models, such as OpenAI's Realtime API, collapse recognition, reasoning and voice into one model. They can reduce latency and handle tone better, and they trade away some of the control you get from choosing each component separately.
| Layer | What it does | Examples |
|---|---|---|
| Telephony or connection | Phone numbers, call audio, transfers, web audio | Twilio, SIP trunks, WebRTC |
| Speech-to-text | Turns the caller's audio into text | Deepgram, OpenAI |
| Language model | Decides the reply and calls tools | OpenAI, Anthropic, Google |
| Text-to-speech | Speaks the reply | ElevenLabs, OpenAI, Deepgram |
| Speech-to-speech (alternative) | One model hears, reasons and speaks | OpenAI Realtime API |
| Orchestration | Turn-taking, interruptions, streaming, tool calls | Vapi, Retell AI, or custom code |
When is a managed platform the right choice?
When speed to launch matters and your needs are within what the platform supports, which covers most receptionist, reminder and lead-qualification agents. Vapi and Retell AI handle turn-taking, interruptions and streaming out of the box, provide phone numbers and transfers, and connect to calendars and CRMs through tool calls and webhooks. A competent agent can be live in days.
They differ in approach. Vapi is developer-oriented and lets you choose and configure each provider, often with your own API keys, so you keep flexibility over models and voices. Retell AI packages more of the stack into an agent builder with telephony included. Both charge per minute of conversation, on top of or bundling the underlying model and voice costs, so read how each counts the components before comparing.
The trade-off is dependency. Your call flows, prompts and integrations live in the platform's format, so moving later takes work, and you inherit its limits on latency, regions and data handling. Keeping prompts and business logic in your own systems, with the platform as the voice layer, makes that dependency manageable.
When does a custom pipeline make sense?
When you need control the platforms do not give you: very low latency tuned end to end, data that must stay in particular regions or providers, unusual call flows, deep integration with your own systems, or volumes high enough that platform margins become a significant cost. A custom build on Twilio, streaming audio to your own service that runs Deepgram, an OpenAI model and ElevenLabs, gives you every lever.
It also gives you every problem. Turn detection, barge-in, streaming, retries when a provider is slow, fallbacks when one is down, call recording and observability all become your code to write and maintain. That is weeks of engineering the platforms have already done, and it only pays back when the control is genuinely needed.
A common middle path is to launch on a managed platform, learn what the calls actually need, and move to a custom pipeline only for the agents where volume, latency or data rules justify it.
How do the costs compare?
Every option ultimately pays for the same components by usage: telephony per minute, speech recognition per minute of audio, the language model per token, and text-to-speech per character or minute. Managed platforms add a per-minute fee for orchestration, or bundle everything into one per-minute price. A custom pipeline pays the component costs directly and replaces the platform fee with your own engineering and hosting.
At low and moderate volumes the platform fee is usually cheaper than building and running your own orchestration. At high, steady volumes the balance can tip, but only after counting the engineering time honestly. Model choice matters more than people expect: a smaller, faster model often answers routine calls just as well at a fraction of the cost of the largest one.
Compare on the cost per resolved call rather than per minute. An agent that is cheaper per minute but transfers twice as many calls to people is the expensive one.
What should you check before committing?
Latency on your own calls, not on a demo. Ring a test agent from a mobile in a noisy place, interrupt it, and listen for the pauses. A voice agent that hesitates for two seconds sounds broken whatever its other qualities.
Data handling: where audio and transcripts are stored and for how long, which regions they are processed in, whether they are used for training, and, for healthcare or other regulated data, whether the providers will sign the agreements your rules require, such as a business associate agreement for HIPAA in the US.
And portability: whether you can export your prompts, call logs and configuration, and whether your business logic can live outside the platform. The easiest time to keep an exit open is at the start.
Common questions
- What is the difference between Vapi and Retell AI?
- Both are managed platforms that orchestrate speech recognition, a language model and a voice into a real-time phone or web agent. Vapi is developer-oriented and lets you choose and configure each provider, often with your own keys. Retell AI packages more of the stack into an agent builder with telephony included. Both charge per minute; compare how each counts the underlying model and voice costs.
- Should I use a voice AI platform or build a custom pipeline?
- Use a platform when speed to launch matters and your needs fit what it supports, which covers most receptionist, reminder and lead-qualification agents. Build a custom pipeline, for example on Twilio with OpenAI, Deepgram and ElevenLabs, when you need tight latency control, specific data residency, unusual call flows or very high volumes. Many businesses launch on a platform and move specific agents later.
- What is a speech-to-speech model?
- A model such as OpenAI's Realtime API that hears audio, reasons and speaks in one step, instead of chaining separate speech-to-text, language model and text-to-speech services. It can reduce latency and carry tone more naturally, at the cost of some control over each component, such as choosing a specialist voice or recogniser.
- How is an AI voice agent priced?
- By usage of its components: telephony per minute, speech recognition per minute, the language model per token and text-to-speech per character or minute. Managed platforms add or bundle a per-minute orchestration fee. The fairest comparison is cost per resolved call, because a cheaper agent that transfers more calls to people costs more overall.
- What should I test before choosing a voice agent platform?
- Latency on real calls from a mobile in a noisy place, including interruptions; how it handles accents and background noise; where audio and transcripts are stored and whether they are used for training; whether providers sign the agreements regulated data requires, such as a HIPAA business associate agreement; and whether you can export your prompts, logs and configuration.