Vocalis

Every model worth running, on your keys

Twenty-seven models across four stages of a call, from six vendors, and a catalogue that tells you what each choice costs you in speed and money before you commit. Swap a model without rewriting an agent.

Engine
VVocalis Realtime

Our speech-native stack, managed end to end.

Sub-second responsesNatural turn-taking
Custom Pipeline

Assemble your own transcription, language and speech models.

Speech-to-Speech

A third-party realtime model, with its own built-in voice.

Custom pipeline · modelsYour keys
Transcription7 modelsReasoning11 modelsSpeech4 modelsSpeech-to-speech5 models
GPT-5.4 MiniRecommended
OpenAI

Fast and cheap enough to answer mid-sentence.

SpeedCost
Claude Sonnet 5New
Anthropic

The strongest reasoning in the catalogue.

SpeedCost
Gemini 3.5 Flash Lite
Google

Lowest latency to first token.

SpeedCost
Claude Haiku 4.5
Anthropic

Small, quick, and careful with instructions.

SpeedCost
Speech modelsYour keys
TranscriptionReasoningSpeechSpeech-to-speech
Cartesia Sonic 3.6Recommended
Cartesia
SpeedCost
ElevenLabs Flash v2.5
ElevenLabs
SpeedCost
ElevenLabs Turbo v2.5
ElevenLabs
SpeedCost

Three engines

The engine decides how the call is assembled, and it is the only choice that changes the shape of everything after it.

Vocalis Realtime
Our speech-native stack, managed end to end. Sub-second responses, natural turn-taking, nothing to assemble. This is the product; the other two are for teams with a reason.
Custom pipeline
Transcription, reasoning and speech chosen separately. Three models, three vendors if you want, and full control over the tradeoff at each hop.
Speech-to-speech
One third-party realtime model that listens and speaks directly, with its own built-in voices. No transcription hop, which removes latency and removes a text transcript to steer with.

Four stages

A pipeline agent hears, thinks and speaks as three separate steps, and each step is a model you choose. A speech-to-speech agent collapses all three into one.

Transcription
Turns the caller's audio into text. This is where phone audio is won or lost: narrowband, background noise, a name spelled out.
Reasoning
Decides what to say and when to call a tool. The stage with the widest range of cost and the most room to over-buy.
Speech
Turns the answer back into audio. Judged on prosody, language coverage and time to first sound, in that order.
Speech-to-speech
One model doing all three. Fastest, and the least steerable, because there is no text in the middle to inspect.

The catalogue

Every model here is one we hold a key for and one the worker knows how to build a client for. Nothing in the picker is a name passed straight through in the hope that it resolves, because a model that cannot be reached does not fail when you pick it. It fails mid-call, as dead air.

Transcription7 models
Deepgram Nova-3Deepgram FluxDeepgram Flux MultilingualDeepgram Nova-3 MedicalDeepgram Nova-2 Phone CallCartesia Ink-2Cartesia Ink Whisper
Reasoning11 models
GPT-5.6 LunaGPT-5.5GPT-5.4 MiniGPT-5.4 NanoGPT-4.1 MiniGPT-4o MiniClaude Sonnet 5Claude Sonnet 4.6Claude Haiku 4.5Gemini 3.7 FlashGemini 3.5 Flash Lite
Speech4 models
Cartesia Sonic 3.6Cartesia Sonic 3.5ElevenLabs Flash v2.5ElevenLabs Turbo v2.5
Speech-to-speech5 models
GPT Realtime 2.1GPT Realtime 2.1 MiniGPT Realtime 2GPT Realtime 1.5Gemini Live 2.5 Flash

Who you can run

OpenAI
Reasoning and realtime
Anthropic
Reasoning
Google
Reasoning and realtime
Deepgram
Transcription
Cartesia
Transcription and speech
ElevenLabs
Speech

Speed and cost, in front of you

Each card carries a three-step meter for speed and one for cost, read left to right. Deliberately not a number: when you are picking, what matters is whether this one is faster or dearer than the one beside it, and a meter answers that without arithmetic.

Cards are badged too. Recommended is what to take if you have no opinion. New is recent enough to be worth trying and not yet what most lines run. Proven is older and still what most production voice stacks are built on. Retiring carries the date it goes away, so you find out before your calls do.

Transcription modelsYour keys
Transcription7 modelsReasoning11 modelsSpeech4 modelsSpeech-to-speech5 models
Deepgram Nova-3Recommended
Deepgram

The broad default. 48 languages, strong on phone audio.

SpeedCost
Deepgram FluxNew
Deepgram

Detects the end of a turn as it hears it. English only.

SpeedCost
Cartesia Ink-2New
Cartesia

Built for interruptions. English only.

SpeedCost
Nova-3 Medical
Deepgram

Tuned for clinical vocabulary and drug names.

SpeedCost
The transcription stage. Featured models come first; the rest sit behind an expander.

Language coverage is checked before you pick

A speech model takes a fixed list of languages and rejects anything outside it when the session starts. On a phone call that failure has no error message and no stack trace. The line simply produces no sound, and there is nothing for anyone to point at.

So the language picker only offers what the models you chose can actually serve, and choosing a language narrows the models in the other direction. A language that would fail is a language that was never offered.

LanguageFiltered by your models
English (US)
Spanish
Vietnamese
Haitian CreoleNot served by Cartesia Sonic 3.6
HmongNot served by Deepgram Nova-3
Coverage differs by model in ways that are easy to miss. Deepgram Nova-3 serves 48 languages while Flux is English only. Cartesia Sonic 3.5 dropped French and Italian and 3.6 brought them back. The picker knows; you do not have to.

Voices

Pipeline agents take their voice from the speech model. Speech-to-speech agents take one of the voices built into the realtime model, each with a short note on how it lands, and any of them can be heard before you save.

VoiceEnglish · US
Marin
Warm, unhurried
Cedar
Low, steady
Sage
Calm, reassuring
Kore
Clear, professional

Swapping a model is not a rebuild

Models are a property of the agent, not something threaded through its prompt, so changing one is a version like any other. Pick the new model, save, test by voice, and roll back if the call sounds worse. The prompt, the fields and the tools are untouched.

Developer documentation: the API, rate limits, webhooks and error handling.

Make the front door intelligent