Built for the Real World · Essay · Conversational AI

Voice is the bottleneck: no language model recovers a wrong product name

On 22 July OpenAI launched Presence, a platform for voice and chat agents delivered only by its own engineers, with no self-serve access and no published price. The delivery model admits that deployment is the hard part. In India the hard part comes earlier, at the speech stage, where no recogniser stays under 20 percent error on telephone audio in every language.

Abstract illustration for “Voice is the bottleneck: no language model recovers a wrong product name”

On 22 July OpenAI launched Presence, an enterprise product for voice and chat agents that answer questions, use company systems, take approved actions and escalate to people. It bundles scoped system access, policies and standard operating procedures, guardrails, simulations, evaluation graders and an improvement loop in which Codex proposes changes that a human approves. You cannot sign up for it. Deployments are led by OpenAI’s forward-deployed engineers and a small set of global systems integrators, access is a limited general availability programme gated on workflow fit, implementation readiness and OpenAI’s delivery capacity, and no price has been published. The proof point was OpenAI’s own English-language phone support line, which the company says now resolves 75 percent of inbound issues without a human and cut hand-offs by 15 percentage points in ten days. Three design partners were named: BBVA for voice banking in Mexico, SoftBank for Japanese-language conversations, and IAG for support during peak demand.

I have built voice systems for Indian callers twice. At Reliance Jio Haptik the platform moved from chat to voice, and the components that made it work in twelve Indic languages were language identification and transliteration, not the intent model. At Salesken the product listened to live sales calls and showed the representative cues in real time, and the hardest programme was transcribing Indian English and code-mixed speech. The lesson both times was the same, and Presence does not change it. Voice agents are being sold as a model problem. In India, at least, they are a speech problem. The error budget of a voice agent is set by its first stage, and no language model recovers a wrong product name, a wrong amount or a wrong date once the transcript has committed to it. This post is about that stage, with the numbers that exist as of August 2026 and the ones that do not.

What Presence is, and what its delivery model says

Read the component list again and notice what is not on it: a model. Presence sits on top of OpenAI’s models, and the Realtime API they are served through remains self-serve for anyone with a card. What is new is the layer above the model, delivered as a six-stage engagement from scoping to post-launch iteration, led by OpenAI’s own engineers throughout.

The delivery model is the honest part. If the lab with the strongest speech-to-speech models will not let an enterprise deploy a voice agent without its own people in the room, the difficulty is not in the model. It is in the scoping, the policy, the evaluation and the escalation path, which is where it has been since the intent era, as I argued in June. Presence productises the work OpenAI’s forward-deployed engineers were already doing one customer at a time, and the absence of a price says the work does not yet have a repeatable cost.

The second thing to notice is the languages. The phone line is English. SoftBank is testing Japanese. BBVA is in Mexico. Nothing in the announcement mentions an Indian language, and nothing in the component list addresses the stage that decides whether an Indian deployment works at all.

The anatomy of a voice turn

A turn of a voice agent passes through six stages, and it helps to name them because vendors sell the fourth and fifth as if they were the whole.

Endpointing. Deciding that the caller has finished. Too eager and the agent interrupts a caller who paused to find a number; too patient and every turn carries a second of dead air. Barge-in, the caller talking over the agent, is the same problem in reverse.

Speech recognition. Turning audio into text, incrementally, while the caller is still speaking. This is where the error budget is set: everything downstream reads the transcript, not the audio.

Language identification. Which language, which script, and whether the utterance switches mid-sentence. At Haptik we ran it on every message, because Hinglish and English went to different pipelines and the identifier’s mistakes were cheap to count.

Understanding. Intent and entities, now usually a language model reading the transcript with the conversation so far.

Decision. Answer, ask, confirm, call a tool, hand off. This is the stage Presence is built around.

Synthesis. Turning the reply into speech, streamed so the first syllable starts before the sentence is finished, in the caller’s language, with the product names pronounced correctly.

Two things accumulate across those stages, differently. Latency is serial: the caller hears nothing until endpointing has fired, recognition has finalised, the model has produced its first tokens and the synthesiser its first audio. At Salesken we learned this the hard way, because a cue that reached the representative after the customer had moved on was worse than no cue. Errors are conditional: each stage conditions on the output of the one before, so a wrong word at stage two becomes a wrong entity at stage four and a wrong action at stage five, with nothing in between that can hear the audio and object. Figure 1 sets out the stages with the budgets I use when I design one. They are targets, not measurements.

The six stages of a voice turn with latency and error budgetsSix boxes in a row, endpointing, speech recognition, language identification, understanding, decision and synthesis, joined by arrows, with a latency target beneath each stage and the kind of error each stage adds, and a bar at the bottom showing the mouth-to-ear budget of about 1.2 seconds and the entity error floor set at the speech recognition stage. ONE TURN, SIX STAGES Budgets are design targets, not measurements 1 Endpointinghas the caller stopped? 2 Speech recognitionaudio to text, streaming 3 Language IDlanguage, script, switch 4 Understandingintent and entities 5 Decisionanswer, confirm, act, hand off 6 Synthesistext to speech, streaming LATENCY TARGET 200 to 300 msafter last speech under 300 msto final transcript under 50 msinside the recogniser 300 to 400 msshared with decision to first tokenone model call for both under 200 msto first audio ERROR IT ADDS cut-off callers,dead air wrong names, numbers,dates: the floor wrong pipeline,wrong reply script right intent on awrong transcript acting on anunconfirmed value mispronounced names,wrong language MOUTH TO EAR endpointrecogniseunderstand and decidespeak about 1.2 s total ERROR FLOOR An entity the recogniser gets wrong stays wrong through stages 3 to 6. Nothing downstream can hear the audio. So the entity error rate at stage 2 is the lower bound on the error rate of the whole agent.
Figure 1. A voice turn and its budgets. Latency is serial and should be budgeted by stage; errors are conditional and are floored by the recogniser. The millisecond figures are design targets for a conversational feel on a telephone line, not measurements of any product.

Word error rate is the wrong number

Every speech vendor reports word error rate: substitutions, deletions and insertions divided by the words in the reference. It is the right number for comparing recognisers in a paper and the wrong one for a voice agent, because it weights every word equally. In a collections call, “haan”, “the” and “four four seven one” each count as one word. The agent survives a wrong “the”. It does not survive a wrong “four four seven one”.

The authors of Voice of India, a benchmark from IIT Madras and Josh Talks published on arXiv in April, make a related point about the metric: a single reference transcript, scored strictly, penalises the spelling variation that is normal in Indian languages, especially for English-origin words in a native script. Their answer was a lattice of valid spellings per utterance. The problem for an agent is different: what matters is not the fraction of words that are wrong but the fraction of the words that carry the decision.

So measure those directly, on your own calls.

Entity error rate. Take every span in the reference that is an entity the agent will act on, normalise it to a value, and score it right or wrong after the same normalisation of the hypothesis. Report it by entity type. A vendor post from Caller Digital on 10 July suggests that for banking and insurance the rate on identifiers, amounts and date phrases should be under 2 percent; a vendor’s rule of thumb, but the right order of magnitude, and a number most vendors do not report unless asked.

Number and name error rates. Digit strings and proper nouns are the entities most likely to be wrong and most expensive when they are. A number read in Hindi with an English unit, a date given as “teesri October”, a scheme name, the name of your own product: none of these are in the recogniser’s training data in the form your callers say them.

Code-switch errors. An English word inside a Hindi sentence transcribed as a Hindi near-homophone, or a whole segment emitted in the wrong script. Score these separately, because they are the errors a language model is worst at repairing: the transcript is fluent and wrong.

Then slice, because the same recogniser is a different product under different conditions. Utterances under two seconds were the worst case for every system: Amazon Transcribe went from 10.45 percent on utterances over five seconds to 18.74 under two. The lowest audio-quality quartile added ten points to the best commercial systems, Gemini 3 Pro going from 13.4 to 23.4 percent. Male speakers were 3.1 to 4.3 points worse than female. Five hundred calls per language, labelled for entities and scored by condition, is a week of work, and it is the same bench discipline I described for text models in July, applied to audio.

The Indian specifics

Four things make Indian speech recognition harder than the demo, and each now has a number attached.

Twenty-two languages, measured on clean speech. Sarvam’s Saaras V3, released on 11 February, covers all 22 scheduled languages and English in one model, trained on more than a million hours of audio, and reports a word error rate of about 19 percent on IndicVoices, down from about 22 for its predecessor. IndicVoices is AI4Bharat’s 2024 corpus: 7,348 hours from 16,237 speakers in 145 districts, 74 percent extempore and 17 percent conversational. It is a serious result, and an average over languages and conditions, and the conditions are not a telephone line.

Telephony audio. Voice of India was built from unscripted telephone conversations: 536 hours, 306,230 utterances, 36,691 speakers, 15 languages, 139 regional clusters sampled in proportion to population. Fourteen systems were evaluated. Most exceeded 20 percent word error, the threshold the authors associate with practical usability, and none stayed under it in every language. Sarvam’s production recogniser, which the paper calls Sarvam Audio, had the lowest error in 13 of 15 languages: 5.0 percent on Hindi, 14.2 on Tamil, 18.2 on Telugu, 16.3 on Kannada, and above 20 on Bhojpuri and Maithili. Gemini 3 Pro was 6.0 on Hindi and best on Bhojpuri and Chhattisgarhi. OpenAI’s GPT-4o Transcribe was 33.9 on Hindi and 84.2 on Kannada, and the mini variant reached 295.9 on Gujarati, which the authors attribute to failed language detection and hallucination. The cloud catalogues agree: Google’s Speech-to-Text lists 13 Indian-language locales on Chirp 3 but telephony-tuned models only for Hindi and Indian English, and Azure’s 13 Indian locales offer custom phrase lists, the mechanism for teaching a recogniser your product names, only for Hindi and Indian English.

Dialect and region. District-level error in Voice of India ran from about 4 percent in Nainital to 44 in Mannarakkat. The Hindi belt clustered below 10 percent; Kerala, interior Karnataka and north Bihar were worst. Bhojpuri and Maithili carried four to five times Hindi’s error on the best systems, and Chhattisgarhi speakers recorded in Tamil Nadu saw 55 to 65 percent. A national average says nothing about the branch network you are deploying to.

Code-switching and product names. No public recogniser has heard your product name, and most callers will say it inside a sentence in another language. At Haptik the hard cases were never the clean ones; they were “mera card block ho gaya” with the identifier deciding on “card” and “block” that the message was English. The vendor ranges in the Caller Digital post, described as typical of the firm’s bake-offs, put Indian-trained recognisers with telephony specialisation at 6 to 10 percent on clean Hindi telephony and 7 to 12 with accent and code-switching, Tamil at 9 to 14 and 11 to 17, Telugu at 10 to 15 and 12 to 18, and global recognisers at 22 to 32 on the same Hindi audio. They are a vendor’s view, consistent with the independent numbers above.

The public programme aimed at this layer is Gnani. On 30 May 2025 the IndiaAI Mission selected it to build a 14-billion-parameter voice foundation model. In December it released Vachana, a speech-to-text model for eleven Indian languages trained on more than a million hours of call audio, with lower error rates claimed and none published. In February, at the India AI Impact Summit, it showed Inya VoiceOS, a five-billion-parameter speech-to-speech research preview, with the 14-billion version to come. Table 1 sets the options side by side with what each number was measured on.

SystemIndian languagesReported error rateMeasured onTelephony and vocabulary
Saaras V3 / Sarvam Audio
Sarvam · 11 Feb 2026
22 scheduled languages plus English, one modelWER about 19% (vendor). VoI: Hindi 5.0, Tamil 14.2, Telugu 18.2, Kannada 16.3IndicVoices, mixed read and extempore (vendor); Voice of India telephony, independent, best in 13 of 15 languagesStreaming, diarisation, language detection; sub-150 ms first token is a stated target
Gemini 3 Pro / Chirp 3
Google
13 Indian-language locales on Chirp 3VoI: Hindi 6.0, Tamil 15.7, Telugu 21.9, Kannada 19.9Voice of India telephony, independent; no Indic WER published by GoogleTelephony models for Hindi and Indian English only
Azure Speech
Microsoft
13 Indian localesVoI: Hindi 11.4, Tamil 28.0, Malayalam 40.9Voice of India telephony, independent; no Indic WER published by MicrosoftCustom phrase lists for Hindi and Indian English only
Amazon Transcribe
AWS
12 of the 15 VoI languages evaluatedVoI: Hindi 6.8, Tamil 19.3, Telugu 19.7, Kannada 18.6Voice of India telephony, independentWorst case on short utterances: 10.45% over 5 s, 18.74% under 2 s
GPT-4o Transcribe
OpenAI
Prompt-conditioned; all 15 evaluatedVoI: Hindi 33.9, Tamil 64.2, Kannada 84.2Voice of India telephony, independent; mini variant 295.9 on GujaratiFailures attributed to language detection and hallucination
IndicConformer
AI4Bharat · open weights
All 15 evaluatedVoI: Hindi 8.2, Tamil 19.9, Telugu 23.7, Kannada 21.4Voice of India telephony, independent; Tier I in the authors’ groupingSelf-hosted; you own the vocabulary and the audio path
Haptik voice
Reliance Jio Haptik
Platform extended to 12 Indic languages by 2021; “100 plus regional dialects” claimed in May 2026No error rate publishedOutcome metrics only: up to 40% lower handle time claimedCode-switching and transliteration layer; recogniser not disclosed
Vachana STT / Inya VoiceOS
Gnani · Dec 2025, Feb 2026
Vachana 11 languages; Inya “15 plus”“Lower error rates” claimed; no figureNot published; 14B model funded by IndiaAI Mission from 30 May 2025Trained on over a million hours of call audio; Inya is speech-to-speech, 5B preview, sub-second latency claimed
Table 1. Speech systems for Indian languages as of August 2026. VoI is Voice of India (arXiv 2604.19151, IIT Madras and Josh Talks), word error rate in percent on unscripted telephone speech with a spelling lattice; vendor figures are from the vendors’ own pages. Locale counts are from the Google and Microsoft documentation pages.

Design patterns that survive bad audio

A 5 percent word error rate on Hindi, the best independent number on the table, still means a wrong entity on a meaningful share of calls once numbers and names are counted properly. The design question is not how to get the recogniser to zero but how to build an agent whose outcomes do not depend on it being right every time.

Confirm entities against the record, not the transcript. The customer’s own data is a small candidate set. “The loan ending 4471, is that right?” is a yes-or-no question with a known answer, and a yes-or-no answer is the one thing every recogniser on Table 1 gets right. Read identifiers back digit by digit, in the caller’s language, and never act on a value heard only once.

Bound the decisions. Give the agent a closed set of actions, each with its own confirmation requirement, and let the policy layer rather than the prompt decide which needs a read-back. This is what Presence’s approved actions are, and what a 2017 dialogue manager was.

Hand off with everything. When the agent escalates, the human should receive the transcript, the audio, the recogniser’s alternatives for the disputed span and the agent’s confidence, so the review shows which stage failed.

Put a decision model in front of the language model. A small, calibrated classifier that takes the recogniser’s confidence, whether the entity validated against the record, the language identifier’s confidence and the turn count, and outputs proceed, confirm or hand off. Trained on your own calls, it gives the agent the threshold a language model does not have. It is the grounding gate I described for retrieval, one stage earlier: a bounded decision before the unbounded one.

Bias the recogniser. Every recogniser worth deploying accepts a vocabulary: product names, scheme names, branch names. Where the cloud vendors restrict that to Hindi and English, the Indian vendors and the open models do not.

Reply in the caller’s language, every turn. Identify the language per utterance, not per call, because switching mid-call is normal, and a Kannada caller answered in Hindi has been failed by the agent whatever the transcript said.

Ask for answers the recogniser can hear. Short utterances are the worst case in every system measured, so design prompts that elicit short but predictable answers: a yes, a day of the week, a choice between two amounts.

Build or buy

Presence is the clearest statement yet of the managed option: the vendor’s engineers scope the workflow, write the policy, run the simulations and own the improvement loop. For an English phone line that is a reasonable trade, and OpenAI’s own line is the evidence. For an Indian deployment the stage that decides the outcome is the one the managed offer says least about, and the questions to ask any vendor are about that stage.

Which recogniser, for which languages, measured on what: a vendor benchmark, or telephone audio with your entities in it. Can the recogniser be swapped without redoing the policy and the evaluation, because the best Indian system in August 2026 will not be the best in August 2027. Where does the audio go, and where does it stay. Who owns the evaluation set and the simulated callers, and do you keep them if you leave. What does the improvement loop see, and who approves its changes. What is the mouth-to-ear latency on a real telephone leg in Bengaluru, not a browser in San Francisco. What is the price per minute of call, not per token. And what happens the day the contract ends.

My rule: if the language is English and the workflow is narrow, buy the managed deployment and spend your own effort on the evaluation set. If the language is not English, own the speech layer, whichever model sits behind it, because that is the layer where the vendors’ numbers do not yet exist and yours will have to.

A worked example: loan collections in Hindi and Kannada

Take a common outbound workflow in Indian retail lending: a reminder call to a borrower whose instalment is overdue, with the goal of a promise to pay on a date. Two languages, Hindi and Kannada, with English product names and amounts in both. The entities the agent must get right are four: the borrower’s identity, the overdue amount, the promised date and the promised amount. Everything else is conversation.

Stage by stage: where a figure is a measurement it is from Voice of India; where it is a target it is mine, and labelled so.

Recognition. The best independent word error rates on telephone speech are 5.0 percent for Hindi and 16.3 for Kannada, with the next best cloud option at 6.0 and 19.9. Those are measurements, and they set the floor. The targets I would write into the contract are an entity error rate under 2 percent on amounts and dates in both languages after read-back, and under 1 percent on the four-digit loan suffix, measured on five hundred of your own recorded calls per language. Kannada will not meet the Hindi floor on raw transcription; it has to meet the entity target through confirmation.

Language identification. Target: a wrong-language reply on fewer than 1 percent of turns, with Hindi-Kannada and either-English switches in the test set.

Understanding and decision. The language model reads the confirmed entities, not the raw transcript, and the decision model gates three actions: record a promise, reschedule the call, hand off to a collections officer. Target: no promise recorded without a confirmed date and amount, a rule rather than a rate, and a hand-off rate set from the first month’s data rather than the vendor’s deck.

Latency. Target: 1.2 seconds mouth to ear at the median on a real telephone leg, with the per-stage budgets in Figure 1.

Outcome. The business reads the promise-to-pay rate and the share of promises honoured. The platform team must read, beside it, the share of promises recorded with a wrong amount or date, found by a human sample a week later: the entity error rate arriving on the balance sheet. If it is above 2 percent, the fix is at stage two, not in the prompt.

None of these targets is exotic. What is unusual is writing them down per stage before the vendor is chosen, so the vendor is chosen against them.

Recommendations

  1. Budget the first stage first. Before you pick a language model, pick a recogniser against telephone audio in each language you serve, and write its entity error rate into the contract.
  2. Stop reporting word error rate alone. Report entity, number, name and code-switch error rates, by language, audio quality and utterance length, on your own calls.
  3. Build the five-hundred-call set. Per language, per channel, labelled for entities, refreshed quarterly. It is the only benchmark that predicts your deployment.
  4. Confirm against the record. Every entity that drives an action is validated against the customer’s data and read back before the action fires.
  5. Put a calibrated decision model before the language model. Proceed, confirm or hand off, from recogniser confidence, entity validation and language identification, trained on your calls.
  6. Keep the recogniser swappable. Table 1 will look different in a year. Policy, evaluation and vocabulary should survive the swap.
  7. Measure latency mouth to ear on a real line. Per stage, at the median and the ninety-fifth percentile, in the city you deploy to.
  8. Buy managed for English and a narrow workflow; own the speech layer for everything else. Presence shows, by what it leaves out, where an Indian deployment must do the work itself.

OpenAI’s decision to deliver Presence only through its own engineers is a confession that the deployment is the hard part. In India the deployment has a harder part inside it, and it sits before the model: a telephone line, twenty-two languages, a caller switching between two of them, and a product name the recogniser has never heard. Get that stage right and a modest model will do. Get it wrong and no model can.

Sources

  1. OpenAI. Introducing OpenAI Presence. 22 July 2026.
  2. AI News. Dashveenjit Kaur. OpenAI Presence: enterprise AI agents. 24 July 2026.
  3. dev.to. Luke Ocodes. OpenAI Presence: voice agents you can’t self-serve. 28 July 2026.
  4. Sarvam AI. Saaras V3 speech recognition for 22 Indian languages and English. 11 February 2026.
  5. arXiv. Bhogale, Dhir, Walecha et al., IIT Madras and Josh Talks. Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India. 21 April 2026, revised 3 July 2026.
  6. arXiv. Javed et al., AI4Bharat. IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages. 4 March 2024.
  7. Caller Digital. Kanan Richhariya. Voice AI WER benchmarks for Indian languages, 2026. 10 July 2026. Vendor figures.
  8. Inc42. Gnani.ai Launches Indic Speech-To-Text Model Under IndiaAI Mission. 19 December 2025.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn