On 27 July 2026 a single-author paper on arXiv, “The Tokenizer Tax”, measured what Indian languages cost a language model before it does any work. On the same parallel sentences, ten Indian languages averaged eight times as many tokens as English under cl100k_base, the tokeniser behind GPT-3.5 and GPT-4, and Malayalam thirteen times. OpenAI’s newer o200k_base cut the average to 2.1 times. It is a useful measurement, and the easiest step in any Indic LLM evaluation.
The same paper shows why it is not enough. On a reading-comprehension benchmark, once the author controlled for how much data each language has, fertility barely predicted accuracy. Token counts tell you what a model will cost in Hindi or Tamil. They do not tell you whether it will answer correctly, in the customer’s script, without promising a refund your policy does not allow.
In June I argued that public Indic benchmarks do not let a buyer compare models. This is the practical sequel: a 30-day plan for an enterprise about to put a model in front of Hindi, Tamil and Hinglish users, from the test set to the rule that turns numbers into a decision.
Most of it I learnt at Haptik, where we extended Jio Haptik’s conversational AI platform to twelve Indic languages through transliteration, language identification and multilingual entailment, and where automated evaluation against labelled transcripts was part of the platform work. The models have changed beyond recognition since. The inputs have not: Hindi typed in Latin script with English nouns, Tamil with the verb in English, three spellings of one word in one conversation, and a speech recogniser that has never heard the product name.
What an Indic language model benchmark can lend you
Public resources are components, not verdicts. Three kinds are worth borrowing; safety and speech suites come later.
Screens. IndicXTREME, from AI4Bharat in December 2022, is nine understanding tasks in 20 languages, including intent classification and slot filling from the MASSIVE dataset in seven Indian languages. IndicGenBench, from Google researchers in April 2024, tests generation in 29 Indic languages and found a significant gap to English in every one. MILU, from November 2024, draws on regional and state examinations in eleven languages, and models do worst on its culturally specific subjects. Use these to drop a candidate that cannot handle one of your languages at all. A pass tells you little more.
Tools for building your own set. Bhasha-Abhijnaanam, from May 2023, is a language-identification test set for all 22 scheduled languages in native and romanised script, published with IndicLID, by its authors’ account the first identifier for romanised Indian languages. Aksharantar, from May 2022, is 26 million transliteration word pairs in 21 languages. For code-mixing, GLUECoS, from April 2020, covers Hindi-English tasks, and DravidianCodeMix, from June 2021, holds about 44,000 Tamil-English, 20,000 Malayalam-English and 7,000 Kannada-English comments in Latin, native and mixed script.
Human preference. AI4Bharat’s Indic LLM-Arena, launched on 10 November 2025, ranks models from blind pairwise votes on prompts typed, transliterated or spoken in Indian languages, Hinglish and Tanglish included. It is the closest public signal to what Indian users prefer, averaged over strangers’ prompts rather than yours.
Indian language LLM comparison: the shortlist in September 2026
Hold three kinds of candidate and no more than four models; each costs native-speaker hours in week three.
Indian models. Sarvam open-sourced Sarvam-30B and Sarvam-105B on 6 March 2026 under Apache 2.0, with a tokeniser covering all 22 scheduled languages in 12 scripts. BharatGen’s Param-2, launched at the India AI Impact Summit in February, is a 17-billion-parameter mixture-of-experts model with 2.4 billion active, 23 languages including Bodo, Dogri and Santali, and a non-commercial licence. Krutrim-2, from early 2025, is a 12-billion-parameter model on Mistral NeMo under the Krutrim Community License. Behind them sits the IndiaAI Mission: a ministry reply reported on 4 April named twelve funded organisations, with BharatGen’s ₹1,058.52 crore the largest award, Sarvam’s and BharatGen’s models on the AIKosh platform, and others, including Gnani’s 14-billion-parameter voice model, listed as planned.
Hosted frontier models. The frontier labs’ multilingual numbers often come from MMMLU, OpenAI’s professionally translated MMLU in 14 languages, of which Hindi and Bengali are the only Indian ones. It says nothing about romanised Hindi or a telephone transcript, and frontier versions change every few months, so the evaluation has to be rerunnable.
Open-weight generalists. Global models you can host in India under your own controls. Their Indian-language quality is the least documented, which is why one belongs on the list.
In ten Indian languages, a model judge agreed with native speakers at a kappa of 0.31 on direct scores
Build the test set from your own traffic
The test set is the one part of the evaluation no vendor has tuned for.
Sample, then de-identify in kind. Draw thirty days of chat logs and call recordings. Replace names, phone numbers, addresses and order numbers with surrogates of the same kind: a Tamil name for a Tamil name, Devanagari numerals where the customer used them. A surrogate in the wrong script changes how the text tokenises and transliterates, and the test stops measuring production.
Tag language and script, then check the tags. Run an identifier that handles romanised text, such as IndicLID, over every message, and have a native speaker correct a sample. Identifiers fail on short messages, and short messages are most of a support queue.
Cover five forms per language. Native script; romanised text; code-mixed text, Hinglish or Tanglish, where English words carry the request; transliteration variants, the nahi, nahin and nhi of one word; and regional vocabulary, including your product and plan names as customers say them. A Hinglish LLM evaluation built on clean Devanagari tests a language few of your customers type. Where logs are thin, generate variants with a transliteration model and have a native speaker vet each; the reference does not change, so variant consistency is a free metric.
Write a reference for every item. The intent, the normalised slots, the facts the answer must contain, the action (answer, ask, act or hand off), and the language and script the reply must use. That last field is a policy decision. Most brands will answer romanised Hindi in romanised Hindi, and a Devanagari reply then fails the turn however fluent it is.
Freeze and version the set, and keep a held-out split that nobody tuning prompts can see, for the reasons I set out in the evaluation bench.
Stratify by language form and task, not by traffic share
A proportional sample reproduces production, which is the wrong goal. If most of your traffic is Hinglish, a thousand proportional items leave Tanglish with a few dozen. At 50 items and an accuracy near 90 percent, the 95 percent interval is about plus or minus eight points; at 400 items it is about three.
So stratify: language form crossed with task type, each cell large enough to support a decision on its own. Report every cell and decide on the worst. An average across languages is how a model that fails Tamil users gets deployed to Tamil users.
Measure tokens per word before you price anything
Fertility, the number of tokens per word, multiplies everything priced or timed in tokens. Aleksandar Petrov and colleagues showed at NeurIPS 2023 that the same text can differ in tokenised length by up to 15 times between languages. Neither paper measured your text. Run each candidate’s tokeniser over a thousand real messages per form and count tokens per word, separately for customer messages and for reference replies, because output costs more and is generated one token at a time.
Romanised text is made of Latin fragments every tokeniser knows; native scripts, the Dravidian ones most of all, are where vocabularies differ, so the cheapest model in Hinglish can be the dearest in Tamil.
Then put fertility where the July paper says it belongs: in cost and latency, not quality. Price it per conversation, in rupees, on a dated rate card. What finishes a task is tokens to done, and a model that needs an extra clarifying turn in Tamil pays for the whole turn.
Task accuracy: how to evaluate a Hindi or Tamil LLM against reference answers
Most of what a support assistant must get right is checkable. It calls tools, to look up an order or book a technician, and their arguments can be matched exactly against the reference.
Normalise before matching. An amount arrives as 2500, as “do hazaar paanch sau”, in Devanagari numerals, or as “twenty-five hundred” inside a Hindi sentence. “Parso” means the day after tomorrow or the day before yesterday, depending on the tense around it. Normalise reference and output to one canonical value, and test the normaliser first, because its bugs look exactly like model errors.
Score the reply’s language and script mechanically. Unicode blocks identify native scripts; a romanised-text identifier separates romanised Hindi from romanised Tamil and from English. It is cheap and deterministic, and catches what clean benchmarks cannot.
Include entailment traps. At Haptik we decided intent by asking whether an utterance entailed an intent description, and the same lens suits testing, because negation and contrast in code-mixed text are where keyword logic breaks. “Refund nahi chahiye, replacement bhej do” mentions a refund and asks for a replacement. Seed every stratum with items like it, and with requests the policy forbids, where the right action is to decline one option politely and offer the other.
Human rating: native speakers, a written rubric and agreement
Some properties only a native speaker can judge: whether a reply sounds like a support agent in that language, whether the form of address is right, whether an English idiom has been translated literally.
Recruit for the variety, not the language. Hire at least three raters per language who speak your customers’ variety, ideally from your own support team, and have them rate blind to the model.
Write binary criteria. Six checks cover most of it: the reply answers what was asked; its facts match the policy; it uses the customer’s language and script; its register is right, aap rather than tum in Hindi and neenga rather than nee in Tamil; it reads naturally to a native speaker; and it promises nothing the policy does not allow. Write each in English and the target language, with a worked pass and fail.
Measure agreement before you read scores. Have all raters score a shared subset and compute Krippendorff’s alpha per criterion and language; it handles several raters and missing ratings, and the DravidianCodeMix authors used it for code-mixed text. Low alpha means an ambiguous criterion: rewrite it and re-rate, rather than averaging disagreement into a score. In PARIKSHA, a 2024 study across ten Indian languages, native-speaker annotators agreed with each other at a Fleiss’ kappa of only 0.54 on pairwise comparisons and 0.49 on direct scores.
LLM-as-judge for Indic languages, and where it fails
PARIKSHA is the best evidence on how far to trust a model judge in Indian languages. Ishaan Watts and colleagues collected about 90,000 human and 30,000 model judgments on 30 models across ten Indian languages, with prompts written by native speakers rather than translated.
In pairwise comparisons the judge agreed with humans at a kappa of 0.49, close to the human 0.54. In direct scoring it fell to 0.31, and to 0.24 on culturally nuanced prompts. It was over-optimistic and missed hallucinations: when both responses were hallucinated it still picked a winner 87 percent of the time, against 53 percent for humans. The GPT-4 judge lifted GPT-4 by 1.4 places on average. Agreement was lowest for Marathi, Bengali and Punjabi in pairwise mode, and for Bengali and Odia in direct scoring.
The weak points follow: register; script, where a fluent Devanagari reply to a romanised question reads as a good answer; code-mixed input; and cultural content, where agreement was lowest. So use a judge only pairwise, in both orders, calibrated per language against your panel with Cohen’s kappa. Never let a candidate judge itself, and leave script and register to deterministic checks and people. Where native references are scarce, test a cross-lingual evaluator such as Hercule, from Sumanth Doddapaneni and colleagues in October 2024, which scores Indian-language answers against English references and on their RECON set aligned more closely with human judgments than proprietary models did.
Safety and refusal behaviour in Indian languages
Safety training does not transfer evenly across languages, and Indian measurements now show it. IndicSafe, posted on 18 March 2026 by Priyaranjan Pattnayak and Sanchari Chowdhuri, put 6,000 culturally grounded prompts on caste, religion, gender, health and politics to ten models in twelve Indian languages. Cross-language agreement was 12.8 percent, safe-response rates varied by more than 17 percent across languages, and some models over-refused benign prompts in low-resource scripts. The same authors’ IndicJR, accepted to the EACL 2026 industry track, found that jailbreaks written in English transfer strongly to Indian languages.
In October 2023 Yue Deng and colleagues had found low-resource languages, Bengali among them, about three times as likely as high-resource ones to draw harmful content from ChatGPT and GPT-4. In 2025 Darpan Aswal and Siddharth Jaiswal showed that phonetic misspellings in code-mixed prompts split safety-critical words into harmless-looking sub-word pieces while the model still understands the request.
So build two sets per language form, written by native speakers rather than machine-translated. One is adversarial: abuse, attempts to obtain another customer’s details, requests to skip verification, slurs, disclosures of self-harm. Any critical unsafe completion fails the stratum. The other is benign but sensitive: a festival gift order that names a religion, a warranty claim for a medical device. Score refusals. A model that is safe in Hindi and over-cautious in Tamil fails Tamil customers as surely as an unsafe one.
Voice: score the pipeline, not the model
On a voice channel the model reads a transcript, and the errors that matter most, names and numbers, are made before it is called. IndicVoices, from March 2024, is 7,348 hours from 16,237 speakers in 145 districts. Voice of India, posted on 21 April 2026 and accepted at Interspeech 2026, is a closed benchmark of 536 hours of unscripted telephone speech in 15 languages, scored so that valid spelling variants are not errors. Neither contains your product names.
Record a few hundred consented calls per language, transcribe them by hand, and run each candidate on the human transcript and on the recogniser’s output. The difference is the speech penalty, and it belongs to the pipeline, not the model. Score entity errors on order numbers, amounts and product names after the confirmation turn, and check that the assistant reads back every identifier before acting on it. On a voice line, a model that confirms beats one that is slightly more accurate on clean text.
Latency and cost per resolved conversation
Measure latency per stratum, median and 95th percentile, to first and last token, on your deployment infrastructure and from Indian network locations.
Then compute the number finance will ask for. Cost per resolved conversation is model cost across every turn, plus speech and infrastructure, plus human handling for the conversations that escalated, divided by the conversations resolved without a human and not reopened within a set window. It needs live traffic, which is why the plan ends with a pilot. A cheaper model that escalates Tamil conversations twice as often is not cheaper, and the per-token price will not show it.
A worked example: thirty days for a Hindi, Tamil and Hinglish support assistant
Take a consumer-appliance brand putting a support assistant on WhatsApp and a phone line, for users who write Hindi in Devanagari, Hinglish, Tamil in Tamil script and Tanglish, with English as a control. The shortlist is an Indian open-weight model served from Indian data centres, a hosted frontier model, that vendor’s cheaper tier, and an open-weight generalist. I will not name them; rankings move faster than methods. Every number below is a design choice of the plan, not a measured result.
Test set. 2,200 items: 400 Hindi in Devanagari, 500 Hinglish, 400 Tamil in Tamil script, 400 Tanglish, 200 English, and 300 recorded calls split between Hindi and Tamil. In each text stratum, 40 percent are requests that need a tool call, 30 percent are policy questions with reference answers, 20 percent are replayed multi-turn conversations and 10 percent should end in a hand-off. A tenth of the romanised items get two spelling variants each. Separately, there are 600 safety prompts: 100 adversarial and 50 benign but sensitive for each non-English form.
Week one builds and freezes the set, writes the rubric, recruits three raters each for Hindi and Tamil and measures fertility. Week two runs every candidate on every item and removes a candidate from any stratum where it fails a gate. Week three puts 150 items per stratum before the raters, with all three on 50 of them, calibrates the judge and runs the safety sets. Week four runs the voice slice, then a pilot on 5 percent of Hindi and Tamil chat traffic for the two finalists, with an agent watching every conversation and a 72-hour reopen window.

