Built for the Real World · Essay · Sovereign AI

Indic models: the gap between the benchmark and the bazaar

Sarvam became a unicorn on Monday, three months after open-sourcing two models trained on IndiaAI compute, and six weeks after Krutrim paused its own foundation-model work. India has a sovereign stack. What it lacks is a way to tell which model reads a romanised Kannada message correctly, or what an Odia reply costs. Those numbers decide deployments.

Abstract illustration for “Indic models: the gap between the benchmark and the bazaar”

On Monday 15 June, Sarvam closed the first $234 million of a $300 million Series B led by HCLTech, at a $1.5 billion valuation, and became India’s newest AI unicorn. Fourteen weeks earlier, on 6 March, it had published the weights of Sarvam-30B and Sarvam-105B under Apache 2.0, trained from scratch on IndiaAI Mission compute. Six weeks before the funding, on 5 May, Krutrim, where I ran applied AI from 2023 to 2025, confirmed to TechCrunch that it had paused chip design and foundation-model work, pulled its Kruti assistant from the app stores, and would sell cloud and GPU capacity instead, a business that produced ₹3 billion of revenue and a first annual profit in FY26. Two Indian frontier programmes, two opposite answers, six weeks apart.

I have spent most of the last decade on language technology for Indian languages: twelve Indic languages on Haptik’s platform at Reliance Jio, the Krutrim foundation models at Ola, and now an enterprise platform at Tata Group. From that vantage the sovereign-model debate is conducted at the wrong altitude: parameter counts, and who trained what from scratch. The costs that decide whether an Indian enterprise can deploy any of it sit lower in the stack: how many tokens a Telugu sentence costs, whether a benchmark score survives a transliterated WhatsApp message, and whether there is a speech model in front of the text model at all. This post is about those costs, with numbers.

The tokeniser tax

A language model sees text as tokens from a fixed vocabulary learned from whatever text the tokeniser was trained on. A tokeniser trained mostly on English and code learns long, efficient pieces for English and falls back to short fragments, sometimes single bytes, for scripts it rarely saw. The measure is fertility: tokens per word. English on a frontier tokeniser is a little over one token per word. Indian scripts are not.

In July 2024 my colleagues at Krutrim published the measurement I still use in every procurement conversation (arXiv 2407.12481). They trained a byte-pair-encoding tokeniser with SentencePiece on twelve languages, English plus eleven Indic, with a vocabulary of 100,000 entries, and compared its token-to-word ratio against the tokeniser behind GPT-4o. On English the frontier tokeniser wins, 1.33 against 1.48, as you would expect from a vocabulary built for English. On every Indian language it loses. Hindi is 1.62 against 1.39, a modest gap. Marathi is 2.53 against 1.80. The Dravidian languages are worse: Telugu 3.15 against 2.29, Tamil 3.16 against 2.30, Kannada 3.03 against 2.31, Malayalam 3.37 against 2.71. Punjabi is 2.68 against 1.56, Assamese 2.70 against 1.99, Gujarati 2.28 against 1.91. And Odia, one of the 22 scheduled languages, is 6.39 against 1.80: the frontier tokeniser spends three and a half times as many tokens on an Odia sentence as a tokeniser that was simply shown enough Odia.

The paper’s second finding matters as much. The tokeniser did not need petabytes: one trained on a 225-million sample of the corpus reached the same ratios as one trained on a sample more than fifty times larger. A good Indic tokeniser is a few GPU-hours and some care about sampling: the cheapest component in the stack, and the one buyers never ask about.

Indo-Aryan scripts

tokens per word · one scale across panels, Odia on GPT-4o = 100%

Hindi · GPT-4o
1.62
Hindi · Krutrim
1.39
Marathi · GPT-4o
2.53
Marathi · Krutrim
1.80
Gujarati · GPT-4o
2.28
Gujarati · Krutrim
1.91
Punjabi · GPT-4o
2.68
Punjabi · Krutrim
1.56

Dravidian scripts

tokens per word · Krutrim paper, Table 4, July 2024

Telugu · GPT-4o
3.15
Telugu · Krutrim
2.29
Tamil · GPT-4o
3.16
Tamil · Krutrim
2.30
Kannada · GPT-4o
3.03
Kannada · Krutrim
2.31
Malayalam · GPT-4o
3.37
Malayalam · Krutrim
2.71

Eastern scripts and English

tokens per word · lower is cheaper

Odia · GPT-4o
6.39
Odia · Krutrim
1.80
Assamese · GPT-4o
2.70
Assamese · Krutrim
1.99
English · GPT-4o
1.33
English · Krutrim
1.48
Figure 1. Token-to-word ratios for the GPT-4o tokeniser against Krutrim’s 100,000-entry Indic tokeniser, from Table 4 of the Krutrim tokenisation paper (arXiv 2407.12481, July 2024). The frontier tokeniser is cheaper only on English. On Odia it is three and a half times more expensive.

What fertility does to the bill, the window and the clock

An enterprise that serves Indian languages through a frontier API pays for the tokeniser three times.

Cost. API prices are per token, so fertility is a direct multiplier on the bill. GPT-4o’s list price is $2.50 per million input tokens and $10 per million output tokens. By the ratios above, a Hindi reply costs about 1.2 times the same reply in English, a Tamil reply about 2.4 times and an Odia reply 4.8 times. The output side is where it bites: output tokens cost four times input tokens, and a model answering in Odia produces Odia tokens at the Odia rate.

Context. A 128,000-token window is quoted as if it were a quantity of text. It is a quantity of tokens. At 1.33 tokens per word it holds about 96,000 English words; at 1.62 about 79,000 Hindi words; at 3.16 about 40,000 Tamil words; at 6.39 about 20,000 Odia words. A retrieval pipeline tuned on English documents overflows on Odia at a fifth of the length. Through the Krutrim tokeniser the same window holds about 71,000 Odia words. The sovereign model may be smaller, but for some languages it reads more.

Latency. Decoding is sequential, one token per step. An Odia reply that is three and a half times as many tokens takes roughly three and a half times as long to stream at the same tokens per second, and a voice agent that must wait for the text before it can speak inherits the whole delay. In a chat window that is an annoyance. On a voice line, where every second of silence is a dropped call, it is the product.

What training for a billion people taught us

In February 2025 we published the technical report for the first Krutrim model (arXiv 2502.09642); my contribution was the fine-tuning and alignment phase. The facts are modest by 2026 standards: seven billion parameters, two trillion tokens of pretraining, a 4,096-token context, a tokeniser built from scratch, supervised fine-tuning followed by direct preference optimisation. On sixteen standard English tasks the model averaged 0.569 against LLaMA-2’s 0.552 and won ten of them, and we reported the Indic side on IndicCOPA, IndicQA, IndicSentiment, IndicTranslation and IndicXParaphrase. What it does not say is what the work taught us.

Data. The scarce resource was never compute. It was Indic text written by a person rather than translated by a machine. Most of what a crawl returns in Hindi or Kannada is news, government notices and machine-translated commercial pages, and a model trained on it learns a register nobody speaks. The share of Indic text in the mix against English and code was the most consequential decision in the programme, because at seven billion parameters every point of Indic share cost something in English reasoning.

Alignment. Instruction and preference data in Indian languages did not exist off the shelf, and translating English instruction sets produced a model that answered Hindi questions in the shape of English answers. We built the safety preference set ourselves, around 20,000 instances focused on safety topics, and learned that what counts as a harmful or an impolite reply differs between Hindi and English, and again in Tamil. Alignment is cultural work before it is statistical work, and it has to be done per language by people who speak it.

Evaluation. The Indic benchmarks available to us were, almost without exception, translations of English benchmarks. A model that did well on them had learned to answer translated English. The score said nothing about whether it could handle a customer writing Hindi in Latin script with the verbs in English, which is how most of the country types. The 2026 models are larger, and that gap is unchanged.

The 2026 sovereign stack

The IndiaAI Mission was approved by the Union Cabinet on 7 March 2024 with ₹10,372 crore over five years and a target of more than 10,000 GPUs. By 30 May 2025 the common compute pool had crossed 34,000 GPUs; by October 2025 the ministry was quoting 38,000 at ₹65 per GPU-hour. The foundation-model pillar selected Sarvam on 26 April 2025, with 4,096 H100s through Yotta and a subsidy of nearly ₹99 crore; Soket, Gnani and Gan on 30 May 2025, for a 120-billion-parameter open model, a 14-billion-parameter voice model and a 70-billion-parameter multilingual model; and eight more on 18 September 2025, bringing the cohort to twelve. BharatGen, the IIT Bombay-led consortium of nine institutions, received ₹988.6 crore of a ₹1,500 crore allocation, the largest single award.

What has shipped is in Table 1. Sarvam-30B and Sarvam-105B are the first Indian models trained from scratch at a scale global buyers notice: mixture-of-experts transformers with 2.4 billion and 10.3 billion active parameters, 16 trillion and 12 trillion pretraining tokens, 32,000 and 128,000 tokens of context, a custom tokeniser covering the 22 scheduled languages across 12 scripts, and Apache 2.0. They were announced at the India AI Impact Summit in New Delhi on 18 February and open-sourced on 6 March. Sarvam’s co-founder Pratyush Kumar said at the launch that the company did not want to “do the scaling mindlessly”, and the active parameter counts bear that out: these are models sized to be served on Indian infrastructure, not to top a leaderboard. Artificial Analysis scores the 105B at 18 on its Intelligence Index, well behind the frontier on general tasks, which is beside the point.

BharatGen’s Param-2, launched at the same summit, is a 17-billion-parameter mixture-of-experts model with 2.4 billion active, 23 languages, a 32,768-token context and about 22.5 trillion training tokens, under a non-commercial licence. Its predecessor Param-1, 2.9 billion parameters trained on five trillion tokens of Hindi and English, came out in 2025 under CC BY 4.0 with a tokeniser that brought Hindi fertility to 1.43 against LLaMA’s 2.65. Krutrim-1 and Krutrim-2 were released in early 2025 under a community licence. Sarvam-M, a 24-billion-parameter post-training of Mistral Small from May 2025, first showed an 86 percent gain on romanised Indian-language arithmetic simply by training on romanised text.

ModelSizeLicenceLanguagesContextIndian-language benchmarks on the card
Krutrim-1
Krutrim · early 2025 (report Feb 2025)
7B dense · 2T tokensKrutrim Community LicenceEnglish + Indic (card lists 8 Indic)4KIndicCOPA, IndicQA, IndicSentiment, IndicTranslation, IndicXParaphrase
Krutrim-2
Krutrim · Feb 2025
12B dense on Mistral NeMoKrutrim Community LicenceEnglish + 12 Indic (incl. Sanskrit, Urdu)128KBharatBench (7 tasks), IndicXNLI, IndicQA, IN22, FLORES-IN, CrossSum-IN
Param-1
BharatGen · 2025 (paper Jul 2025)
2.9B dense · 5T tokensCC BY 4.0Hindi, English2KMILU, SANSKRITI, MMLU-Hindi, HellaSwag-Hindi
Sarvam-M
Sarvam · 23 May 2025
24B dense on Mistral Small 3.1Apache 2.0English + 10 Indicnot statedMILU, IndicGenBench, romanised MMLU / GSM-8K / ARC / TriviaQA, Indic Vibe Check
Param-2
BharatGen · Feb 2026
17B MoE · 2.4B active · ~22.5T tokensBharatGen non-commercial23 (English, Hindi + 21)32KSanskriti, Indic ARC, Indic TriviaQA, Indic BoolQ, HellaSwag-Hi, MMLU-Hi
Sarvam-30B
Sarvam · 6 Mar 2026
32B MoE · 2.4B active · 16T tokensApache 2.022 scheduled + English32KMILU (76.8), IndiVibe (89% pairwise wins)
Sarvam-105B
Sarvam · 6 Mar 2026
105B MoE · 10.3B active · 12T tokensApache 2.022 scheduled + English128KIndiVibe (90% pairwise wins); MILU not on the card
Table 1. Seven Indian models from the model cards and papers. MILU is the only Indian-language benchmark that appears on more than two cards; no benchmark appears on all seven, and the two largest models lead with suites their own labs built.

The benchmark problem

Look at the last column of Table 1. Seven Indian models, and no two report the same set of Indian-language benchmarks. Krutrim-2 reports BharatBench, which Krutrim built. Sarvam-105B reports IndiVibe, which Sarvam built. Param-2 reports Sanskriti and Hindi translations of English suites. MILU, the only benchmark on more than two cards, is dropped from the Sarvam-105B card. A buyer who wants to know whether the 105B is better than Param-2 at Marathi customer service has no published number that compares them.

The benchmarks themselves are good work. MILU, from AI4Bharat and IBM Research India in November 2024, is about 85,000 multiple-choice questions from Indian competitive examinations across eleven languages and 42 subjects; the best of 45 models, GPT-4o, scored 72 percent. IndicGenBench, from Google Research in April 2024, covers 29 languages and 13 scripts with summarisation, translation and cross-lingual question answering. IndicParam, posted in November 2025, is more than 13,000 questions in eleven low-resource languages, among them Dogri, Maithili, Bodo and Santali; the best of nineteen models, GPT-5, managed 45 percent. BharatBench is 300 native-speaker-rated examples across eight languages and five generation tasks, plus classification and entity recognition. IndiVibe is 110 English prompts translated into 22 languages in native and romanised script, judged pairwise by Gemini 3 on fluency, script correctness, usefulness and verbosity.

Three things follow. First, the two suites the biggest models lead with were built by the labs that sell the models, and both are scored by a model rather than a person. Sarvam-105B wins 90 percent of IndiVibe comparisons; Krutrim-2 is within a point of GPT-4o on BharatBench cultural context. I was on the team that shipped one of those models and I believe the numbers. I would still not buy on them. Second, every one of them is multiple-choice or judged generation on clean text in native script. None is a transcript. Third, the low-resource languages the sovereign argument is really about are where every model including the frontier scores lowest and where the fewest benchmarks exist.

At Haptik we extended a conversational platform from English to twelve Indic languages, and the hard part was never the clean case. The input was Hindi typed in Latin script with English nouns, Tamil with three spellings of one word in a session, Kannada voice transcribed by a speech model that had never heard the product name. Three unglamorous components made it work: a transliteration layer that normalised romanised text before anything else saw it, language identification that could tell Hinglish from English on a six-word message, and a multilingual entailment model that decided intent without a labelled dataset per language. None of them is measured by MILU or IndiVibe. A model that scores 76.8 on MILU can still answer a Latin-script Hindi message in Devanagari, which the customer cannot read, and a model that scores 72 percent can still route “mera card block ho gaya” to the English flow because the identifier saw “card” and “block”.

One query, two tokenisers, two models

Here is the arithmetic for a single turn, measured with the public tokenisers: the o200k encoding behind GPT-4o, the 262,144-entry tokeniser Sarvam ships with its 2026 models, and the 131,072-entry vocabulary Krutrim-2 shares with its Mistral NeMo base. I ran them this week; the counts are reproducible from the model cards.

The turn is a support exchange. A 92-word English system prompt sets the policy: answer in the customer’s language and script, offer a refund in five to seven working days or re-delivery, never promise same-day delivery, escalate damage or double payment to a human. The customer writes 60 words of Latin-script Hindi: an order placed on 14 June, the app says out for delivery, money already taken by UPI, either deliver today or refund in full, and how many days will the refund take. The model replies in 54 words of the same romanised Hindi.

On GPT-4o’s tokeniser that is 111 tokens of system prompt, 92 of query and 102 of reply: 203 in, 102 out. At list price the turn costs $0.0015, about $1,530 per million turns. Sarvam’s tokeniser makes it 112, 84 and 95, about five percent fewer, because romanised Hindi is Latin-script fragments both vocabularies know. Krutrim-2’s is 113, 102 and 105. For Hinglish the tokeniser tax is small, and the choice of model should be made on behaviour, not tokens.

Now move the same customer to Odia. The 14-word Odia query is 97 tokens on GPT-4o’s tokeniser, 26 on Sarvam’s and 250 on Krutrim-2’s. A 42-word Odia reply is 304, 80 and 819. The GPT-4o turn is now 208 tokens in and 304 out, $0.0036, about $3,560 per million turns: 2.3 times the Hinglish turn for the same content, almost all of it on the output side, and the reply takes three times as long to stream. Through Sarvam’s tokeniser the Odia turn is 138 in and 80 out, fewer tokens than the Hinglish turn on GPT-4o. And Krutrim-2, a model marketed as Indic, spends 819 tokens on 42 words because Odia is not one of its twelve languages and the text is broken down almost to the byte. Sovereign is a property of each language, not of the flag on the model card, and it has to be checked per language.

The failure modes are not in the token counts. The first is script: the customer wrote in Latin script, and a model that learned Hindi from Devanagari will often answer in Devanagari, a failed turn for a customer who cannot read it. The second is language identification: “Hello, maine 14 June ko ek order kiya tha” begins with two English words, and a classifier trained on clean text calls it English. The third is transliteration variance: nahi, nahin and nhi are the same word, and a retrieval step keyed on the surface form misses two of them. The fourth is entailment: the customer has asked for one of two outcomes, the policy forbids one, and the model has to recognise that “aaj hi delivery” is a same-day request and decline it politely while offering the other. We saw every one of these at Haptik; I see them on the platform today. No benchmark in Table 1 measures any of them.

What a sovereign programme should fund

The IndiaAI Mission’s money has gone overwhelmingly to pretraining. Pretraining is necessary; Sarvam’s two models and Param-2 exist because of it. But the mission’s own numbers show where the marginal rupee should go next, and it is not another base model.

Tokenisers. The cheapest fix in the stack and the one with the largest cost effect for low-resource languages. A public, Apache-licensed Indic tokeniser with a published fertility table across all 22 scheduled languages and their romanised forms would let any team, including teams adapting foreign open-weight models, start from a vocabulary that does not charge Odia six times the English rate. Sarvam says its 2026 tokeniser is significantly more efficient on Odia, Santali and Manipuri; the numbers should be published so buyers can compare.

Evaluation on field data. A national benchmark built from the text people actually send: romanised, code-mixed, voice-transcribed, with spelling variance, in all 22 languages, scored by native speakers, maintained by an institution that does not sell a model. AI4Bharat and the IndicParam team are the right institutions; the missing piece is the dirty-input track, and a requirement that any model trained on mission compute report on it.

Speech. Most Indian users who will ever talk to an enterprise AI system will do it with their voice, and a speech model has to recognise the language before any text model sees a token. Gnani’s 14-billion-parameter voice model is the one first-cohort selection aimed at this; it should not be the only one. Speech recognition at call-centre quality, with code-switching and product names, has been the bottleneck in every deployment I have seen since 2019.

Distribution. A model on AI Kosh and Hugging Face is a file. Enterprises consume endpoints with uptime, data residency, rate limits and an invoice, and most will not run a 105-billion-parameter model themselves. Funding should reach the serving layer: in-country endpoints for mission-funded models, at published prices, from more than one provider.

Buying Indic capability: eight checks

  1. Measure fertility on your own text first. Take a thousand real messages per language from your logs, run the candidate tokenisers over them and compute tokens per word. That ratio multiplies every price in the quote.
  2. Price output tokens, not blended tokens. Replies in Indian scripts are where fertility compounds against the higher output price. Model the bill per turn, per language, at your real reply length.
  3. Check context in words, not tokens. Divide the window by the fertility ratio for each language and compare it with your longest document and conversation.
  4. Test per language, not per model. An Indic label does not tell you which languages the model tokenises and which it spells out in bytes. Krutrim-2 covers twelve; Sarvam’s 2026 models claim 22; Param-2 claims 23. Verify the ones you serve.
  5. Build your own benchmark from your own traffic. Sample romanised, code-mixed and transcribed messages, label the intent and the correct reply script, and run every candidate on it. Treat MILU, IndiVibe and BharatBench as a screen, not a decision.
  6. Score script choice, language identification and entailment separately. These are the failures a clean-text benchmark cannot see and that produce most bad turns in production. Keep a transliteration normaliser and a language identifier in front of whichever model you pick.
  7. Put speech in the evaluation. If any channel is voice, evaluate the speech model and the text model together on recorded calls with your product names in them.
  8. Read the licence against the language. Apache 2.0 for Sarvam’s models and CC BY 4.0 for Param-1 clear a legal review; Param-2’s non-commercial licence and Krutrim’s community licence need reading. A model you cannot deploy commercially is not a sovereign option, however good its Odia.

The sovereign stack India has built since March 2024 is real: tens of thousands of GPUs, twelve funded teams, two Apache-licensed models trained from scratch, a 17-billion-parameter model from a public consortium, and a unicorn to carry the commercial side. What it does not yet have is a way for a buyer to tell which model will answer a romanised Kannada voice transcript correctly, and at what price. That is a smaller problem than pretraining and a cheaper one, and it is the one that decides whether any of this reaches the bazaar.

Sources

  1. arXiv. Krutrim LLM: A Novel Tokenization Strategy for Multilingual Indic Languages with Petabyte-Scale Data Processing. July 2024.
  2. arXiv. Krutrim LLM: Multilingual Foundational Model for over a Billion People. 10 February 2025.
  3. Sarvam. Open-Sourcing Sarvam 30B and 105B. 6 March 2026. Model cards: sarvam-105b, sarvam-30b; Sarvam-M, 23 May 2025. Artificial Analysis, Sarvam 105B and 30B.
  4. arXiv. PARAM-1 BharatGen 2.9B Model. 16 July 2025. Hugging Face, Param2-17B-A2.4B-Thinking model card. BharatGen, ₹988.6 crore under the IndiaAI Mission, 18 September 2025.
  5. arXiv. MILU: A Multi-task Indic Language Understanding Benchmark. 4 November 2024. IndicGenBench, 25 April 2024. IndicParam, 29 November 2025.
  6. Krutrim AI Labs. BharatBench: Comprehensive Multilingual Multimodal Evaluations of Foundation AI Models for Indian Languages. 4 February 2025. Model cards: Krutrim-2-instruct, Krutrim-1-instruct.
  7. DD News. IndiaAI Mission gets boost as compute capacity tops 34,000 GPUs. 30 May 2025. The Week, Cabinet approves ₹10,372 crore IndiaAI Mission, 7 March 2024. Analytics India Magazine, the eight phase-2 firms, 18 September 2025. Indian Masterminds, 38,000 GPUs at ₹65 an hour, 11 October 2025.
  8. TechCrunch. Sarvam becomes India’s newest AI unicorn with $234 million round led by HCLTech. 15 June 2026. Krutrim shifts to cloud services, 5 May 2026. Sarvam’s new models, 18 February 2026. The Hans India, Sarvam launches 30B and 105B, 18 February 2026. OpenAI, API pricing.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn