Ten years ago I joined a startup that wanted to let people talk to their bank. The tools were a sentence classifier, a slot filler, a dialogue state machine and a great deal of patience. This spring the tools are different. OpenAI released GPT-5.5 on 23 April and described it as its “smartest and most intuitive to use model” yet. Anthropic shipped Claude Opus 4.7 on 16 April and Opus 4.8 on 28 May. Google announced Gemini 3.5 Flash at I/O on 19 May and framed the next wave of its products around agents rather than chatbots. Yesterday, 3 June, Google DeepMind released Gemma 4 12B, an open model that takes text, images and audio in one decoder and runs on a laptop with 16 GB of memory. A bank that wants a talking machine in June 2026 can have a convincing one by Friday.
What follows is a history of four attempts to build that machine: at Active.Ai for banks, at Reliance Jio Haptik for Indian consumers in a dozen languages, at Salesken for live sales calls, and at Ola Krutrim for a foundation model and the agents on top of it. It is also an argument. The intent-and-entity era that the industry now treats as a false start got five things right, and the agent stacks being deployed this month have dropped all five. They can be put back, and the teams that put them back will ship agents that survive contact with a contact centre. The teams that do not will relearn 2017 at 2026 prices.
Chapter one: Active.Ai, 2016 to 2019
Active.Ai was a Singapore-headquartered startup with its engineering in Bengaluru, founded in 2016 to sell conversational banking to banks and fintechs. I set up its AI research and development function and architected Triniti, the platform underneath. The first version was honest about what it was: a natural-language understanding layer that mapped an utterance to one of a few hundred intents, an entity extractor that pulled out amounts, dates, account types and payee names, and a dialogue manager that held the state of a multi-turn conversation and decided what to ask next. “Send five thousand to Priya” became the intent fund_transfer with amount and payee filled and source_account empty, and the dialogue manager’s only job was to fill the empty slot, confirm, and call the bank’s API.
Three things made it harder than it sounds. Banking utterances are short, ambiguous and full of context from earlier turns: “the other one” after a list of cards, “same as last month” after a bill payment. We built multi-turn context into the dialogue state rather than into the classifier, so the system always knew which slot it was waiting on and could resolve a reference against a small, explicit set of candidates. Then there was the long tail of enquiries that no intent covered, the questions that live in a bank’s documents rather than its APIs. For those we built machine comprehension: given the bank’s own text, find the span that answers the question. In December 2018 we put a single-model entry built on BERT on the Stanford SQuAD 1.1 leaderboard, where it placed ninth at the time. And the whole stack had to run inside a bank’s perimeter, with latency low enough for voice, so every model had a size budget before it had an accuracy target.
The results took years to show. When Gupshup acquired Active.Ai in April 2022, the release said the platform had enabled more than 300 million user interactions across voice, video and messaging, served banks and fintechs in 43 countries, managed over 30 million service requests and fulfilled more than 50 million enquiries with 95 percent accuracy. Ninety-five percent accuracy on enquiries was a figure the company could state because the outputs were bounded and the evaluation was defined: each enquiry resolved to a label, each label was right or wrong, and someone counted. You cannot report an accuracy figure for a system that can say anything.
Chapter two: Reliance Jio Haptik, 2019 to 2021
Haptik had a different problem. Its customers were Indian consumer businesses, and its users typed the way Indians type: Hindi in Latin script, English with Marathi nouns, Tamil with the verb in English. The team had already seen bots break when users typed in Hindi. The platform I joined, as AVP leading conversational AI and NLU research, was English-first, and the job was to make it work in twelve Indic languages without multiplying the training data twelve times.
The approach was layered. A language identifier ran on every utterance first, because Hinglish and English routed to different pipelines and the identifier’s mistakes were cheap to measure. Romanised text was transliterated into native script, so that one Hindi model could serve users whether they typed in Devanagari or Latin characters, and the transliterator became a product surface tuned on real traffic. For intent recognition across languages we trained multilingual entailment models: rather than asking which of three hundred intents an utterance belonged to, the model asked whether the utterance entailed a given intent description, which let a new intent in a new language be added with a sentence rather than a dataset. Around that we built the tooling that mattered most. Intent discovery clustered unhandled utterances into candidate intents ranked by volume, so the fallback bucket became a backlog. Automated evaluation re-ran every bot against its own labelled transcripts on every model change, and the diff was the release gate. Training-data generation produced paraphrases that a human accepted or rejected. The platform also moved from chat to voice and to question answering over documents.
Haptik kept growing after I left. In February 2024 it reported passing 15 billion two-way conversations and 10 billion one-way notifications since its founding in 2013, with non-English conversations at 23 percent of the total. Nearly a quarter of a very large volume is in languages most model evaluations do not cover, and the only reason anyone could state the figure is that the pipeline identified the language of every message before it did anything else.
Chapter three: Salesken, 2021 to 2023
At Salesken, as VP Data Science, the conversation was a sales call and the user was the person selling, not the customer. The product listened to a live call, transcribed it, and showed the representative cues in real time: the objection just raised, the discovery question not yet asked, the moment the customer’s tone changed. This is a harder latency problem than chat, because a cue that arrives after the customer has moved on is worse than no cue, and a harder modelling problem than banking, because there is no API to call and no slot to fill.
Three choices mattered. We trained domain-aligned language models on sales conversations rather than relying on general ones, because the vocabulary of an insurance renewal call is not the vocabulary of the web. We quantised them, because a team of 35 serving hundreds of customers could not afford a GPU per call, and the accuracy we gave up was measured on our transcripts. And we ran research to product on six-week cycles, which forced every idea to meet a transcript before it met a slide. Speech transcription for Indian English and code-mixed speech was its own programme, and it taught me that the error budget of a conversational system is set by its first stage: a transcript with a wrong product name poisons every model downstream, and no amount of language-model cleverness recovers it.
Chapter four: Ola Krutrim, 2023 to 2025
Krutrim was the first time the model itself was the product. As Director of Applied AI I led fine-tuning and alignment of the Krutrim language models, including Krutrim Spectre, and the applied layer on top: the Krutrim chat assistant, an agent-builder platform and contact-centre AI for enterprises. The technical report we published in February 2025 describes a model trained on two trillion tokens with the largest Indic dataset we knew of, and the motivation in one line: Indic languages are about one percent of Common Crawl while India is 18 percent of the world’s population. On the report’s own benchmarks the model matched or exceeded Llama 2 on 10 of 16 tasks.
What I learned at Krutrim is the hinge of this essay. Fine-tuning and alignment are where you decide what a model will refuse, how it will hedge, and whether it will say “I do not know”. The agent builder and the contact-centre product were where those decisions met real customers. And the questions customers asked were the ones my 2017 self would have answered with a configuration file: what happens when the model is not sure, how do I stop it answering outside its remit, how do I know it worked last week and still works today. In the intent era those had mechanical answers. In the agent era they had a prompt.
What the intent era enforced
Here is what a 2017 platform gave you whether you wanted it or not.
Bounded outputs. The system could only do what an intent existed for, and could only say what a template allowed. This was the chief complaint against it, and it was also its safety case: a compliance officer could read every possible response in an afternoon.
Calibrated confidence with thresholds. Every classification came with a score, the score was calibrated against held-out transcripts, and the threshold below which the system declined to act was a number a product manager owned. Thresholds were tuned per intent, because the cost of a wrong fund_transfer is not the cost of a wrong balance_enquiry.
Explicit fallback and escalation. Below threshold, the system asked a clarifying question with a fixed set of options. Below a second threshold, or after two failed clarifications, it handed off to a human with the transcript and the top candidate intents attached. Escalation was a designed path with its own metrics, not a failure.
Evaluation on real transcripts. The test set was last month’s traffic, labelled. Every model change ran against it before release, and accuracy, fallback rate and escalation rate were compared with the previous build. Public benchmarks told us whether a technique was promising; transcripts told us whether to ship.
Slot-level grounding. An action fired only when every required slot held a value validated against the bank’s own data: an account the customer owned, a payee on their list, an amount inside their limit. The system could not act on an unconfirmed value.
None of these was a research result. They were engineering constraints that fell out of the architecture, and the architecture is what the industry threw away once a large model could understand the utterance without any of it.
The agent stacks of June 2026
The models released this spring are extraordinary. GPT-5.5 arrived on 23 April with OpenAI emphasising agentic coding and computer work. Claude Opus 4.7 on 16 April added an “xhigh” effort level between high and max; Opus 4.8 on 28 May scored 84 percent on Online-Mind2Web, a browser-agent benchmark, and Anthropic says it is about four times less likely than its predecessor to let a flaw in its own code pass unremarked, at an unchanged $5 and $25 per million tokens. Gemini 3.5 Flash was announced at I/O on 19 May and is generally available at $1.50 and $9 per million tokens with a context window of just over a million tokens; DeepMind’s chief technologist said it outperforms Gemini 3.1 Pro on nearly all of Google’s benchmarks. Gemma 4 12B, released yesterday, handles text, images and audio in one encoder-free decoder and runs on a 16 GB laptop.
The release notes are about effort levels, honesty, browser benchmarks, modalities and price. An enterprise agent built on these models typically has none of the five controls above. The output is unbounded by construction; the prompt asks the model to stay on topic and it usually does. There is no confidence score on a tool call, so there is no threshold, so there is no principled point at which the agent stops and asks. Fallback is a sentence in the prompt that says “if unsure, ask the user”, which loses to forty other sentences as the context fills. Evaluation is the vendor’s benchmark plus a handful of hand-written cases, because nobody built the harness that replays last month’s transcripts. Grounding is a tool call the model may make, may make with the wrong arguments, or may describe having made without making. Cost is proportional to tokens rather than fixed per turn, so a careful agent is an expensive one and a verbose customer is a cost centre. And multilingual handling is assumed because the model is multilingual, and never measured per language.
The labs are moving on some of this. Opus 4.8’s honesty work is calibration under another name, and both Anthropic and OpenAI now expose effort as a dial, a cost control the intent era never needed because inference was free. But these are properties of the model. The controls I am describing were properties of the platform, which is what a bank’s own team builds, and a bank will swap models three times in the next year and cannot re-derive its safety case each time.
A worked example: blocking a card
Take the most common urgent intent in retail banking: a customer has lost a card and wants it blocked.
The 2017 way. The customer types “lost my card block it now”. The NLU returns card_block at 0.94 with no card entity. The dialogue manager sees two cards on the profile and asks a bounded question: “Which card: the debit card ending 4421 or the credit card ending 0187?” The customer answers “debit”. The slot fills with a value validated against the customer’s own card list, the system reads back the action and asks for a yes, the API call fires, and the template responds with the confirmation and the reissue timeline. Total model inference: two classifier passes. Total possible responses: those in the template library. The failure modes are real. If the customer writes “my card got stuck in the ATM”, the classifier may return atm_dispute at 0.61, below threshold, and the customer gets a clarification menu that does not include the thing they mean. If they write “block my card and also tell me my balance”, the system handles one intent and drops the other. A new card product means new paraphrases, new labels and a retrain. The system is dumb in exactly the ways that are visible, predictable and cheap to fix.
The 2026 way. The same message goes to an agent built on one of this spring’s models with three tools: list_cards, block_card and raise_dispute. The model reads the message, calls list_cards, sees two cards, and here the path forks. A well-prompted agent asks which card. A slightly less well-prompted one reasons that “lost my card” most often means the debit card, blocks it, and tells the customer it has done so, and the customer who meant the credit card finds out at the checkout. The same agent handles “stuck in the ATM” perfectly, handles the balance request in the same turn, and copes with Hinglish without anyone having trained it. Its failure modes are the mirror of 2017’s. There is no score below which it will not act, so whether it asks or guesses depends on the model’s disposition and where the instruction sits in a long context. If block_card returns an error, the agent may report the error, retry, or describe a successful block that did not happen. If a merchant name in the transaction history reads like an instruction, the model may follow it, a well-documented class of failure that a 2017 slot validator could not suffer because it never read free text from a record. And the bank cannot state an accuracy figure for the agent, because its output is not a label and nobody has defined what correct means for a paragraph.
The right system in 2026 is the second one with the first one’s controls. The model understands the utterance; the platform decides whether to act. The model proposes block_card with the debit card as its argument; a policy layer checks that a confidence signal exists and clears the threshold for an irreversible action, that the card is on the customer’s own list, that the customer has confirmed, and that a human can reverse the block within the hour. The agent asks because the platform makes it ask, not because the prompt suggested it. It handles the ATM case, the two-intent case and the Hinglish case, costs more per turn, and can still report an accuracy figure, because what is scored is the proposed action, not the prose around it.
Principles a platform team should carry over
- Bound the action space, not the language. Let the model say what it likes; let it do only what a declared tool with a declared schema permits, and keep the tool list short enough for a compliance officer to read in an afternoon.
- Manufacture a confidence signal and threshold it. If the model will not give you a calibrated score, build one: sample the proposed action several times and measure agreement, or run a small verifier over the proposal, and calibrate it on your own transcripts. Set thresholds per action, because a wrong block is not a wrong balance.
- Design fallback and escalation as paths, with metrics. The clarifying question, the hand-off to a human with the transcript attached, and the rate at which each fires are product features. A sentence in a prompt is not a fallback.
- Evaluate on last month’s transcripts before every change. Build the replay harness first. Benchmarks such as Online-Mind2Web tell you which model to try; your transcripts tell you whether to ship it.
- Ground every argument against a record before acting. A card, a payee, an amount and a date are validated against the customer’s own data by code, not by the model, after the model proposes and before the tool fires.
- Treat cost per turn as a design input. A classifier was free; a frontier model at $5 per million input tokens reading a 50,000-token context on every turn is not. Route bounded decisions to small models, including open ones such as Gemma 4 12B on hardware you own, and reserve the frontier for the turns that need it.
- Identify the language first and measure per language. A multilingual model is a capability, not a result. Run language identification on every message, report accuracy, fallback and escalation by language, and expect the long tail to be where the system fails quietly.
I spent the first half of these ten years building systems that were too rigid to be loved and the second half building systems that were too fluent to be trusted. The platform that holds both is not a research problem. It is a set of controls we already had, applied to a model that no longer needs them to understand us and needs them more than ever to act for us.