Every model worth running, on your keys
Twenty-seven models across four stages of a call, from six vendors, and a catalogue that tells you what each choice costs you in speed and money before you commit. Swap a model without rewriting an agent.
Our speech-native stack, managed end to end.
Assemble your own transcription, language and speech models.
A third-party realtime model, with its own built-in voice.
Fast and cheap enough to answer mid-sentence.
The strongest reasoning in the catalogue.
Lowest latency to first token.
Small, quick, and careful with instructions.
Three engines
The engine decides how the call is assembled, and it is the only choice that changes the shape of everything after it.
Four stages
A pipeline agent hears, thinks and speaks as three separate steps, and each step is a model you choose. A speech-to-speech agent collapses all three into one.
The catalogue
Every model here is one we hold a key for and one the worker knows how to build a client for. Nothing in the picker is a name passed straight through in the hope that it resolves, because a model that cannot be reached does not fail when you pick it. It fails mid-call, as dead air.
Who you can run
Speed and cost, in front of you
Each card carries a three-step meter for speed and one for cost, read left to right. Deliberately not a number: when you are picking, what matters is whether this one is faster or dearer than the one beside it, and a meter answers that without arithmetic.
Cards are badged too. Recommended is what to take if you have no opinion. New is recent enough to be worth trying and not yet what most lines run. Proven is older and still what most production voice stacks are built on. Retiring carries the date it goes away, so you find out before your calls do.
The broad default. 48 languages, strong on phone audio.
Detects the end of a turn as it hears it. English only.
Built for interruptions. English only.
Tuned for clinical vocabulary and drug names.
Language coverage is checked before you pick
A speech model takes a fixed list of languages and rejects anything outside it when the session starts. On a phone call that failure has no error message and no stack trace. The line simply produces no sound, and there is nothing for anyone to point at.
So the language picker only offers what the models you chose can actually serve, and choosing a language narrows the models in the other direction. A language that would fail is a language that was never offered.
Voices
Pipeline agents take their voice from the speech model. Speech-to-speech agents take one of the voices built into the realtime model, each with a short note on how it lands, and any of them can be heard before you save.
Swapping a model is not a rebuild
Models are a property of the agent, not something threaded through its prompt, so changing one is a version like any other. Pick the new model, save, test by voice, and roll back if the call sounds worse. The prompt, the fields and the tools are untouched.