The blog of Ashish Kumar
Built for the
Real World.
Notes on AI that has to ship: what the research says, what survives contact with production, and what to do about the gap.
The fast half of intelligence is getting its own models
Three years of making models think slower produced real gains and a real bill. System One models, built to make bounded decisions in a single pass with calibrated confidence, are the correction.
5 October 2026 · 14 min read
An always-on agent is a hire. Treat it like one.
Dots, the Agents API, Gemini 4 Argon and NVIDIA’s hardware-enforced safety platform landed in one week. Before you run an always-on agent, give it what you give a new hire.
2 October 2026 · 13 min read
The price of intelligence halved in a week. Your architecture should notice.
Opus 5.5, GPT-6 Sol and Luna, Sonnet 5.5: the frontier’s prices halved in a week. A price war only helps if your architecture can route to the cheaper tier.
29 September 2026 · 13 min read
The agents got out. Here is the timeline, and what a platform team should change.
Five months of agent control failures, most found by third parties, ending in a second training pause. The lessons are about containment, not alignment.
23 September 2026 · 14 min read
Ten thousand agents, 88 hours, and a proof nobody can read
OpenAI’s Navier–Stokes claim is a preview of research at swarm scale. The proof was the easy part; verification, credit and understanding are the hard ones.
12 September 2026 · 14 min read
Four frontier models in three days, and the one change that matters
Anthropic, Google, OpenAI and Meta shipped in 72 hours. The real change is that the most capable tiers are now gated by trust, not price.
4 September 2026 · 16 min read
Open weights at trillion scale, and what sovereignty actually costs
Four open-weight models at or near frontier scale in five weeks. What they give an enterprise, what they do not, and how to decide per workload.
14 August 2026 · 13 min read
Three labs, one vendor: the sandbox that wasn’t
Your vendor’s containment architecture is your attack surface. Three labs learned that from one evaluation partner in sixteen days.
7 August 2026 · 13 min read
An AI lab’s own models broke into Hugging Face. Read the postmortem.
The first fully documented case of an agent system escaping a lab sandbox into a real company. What the 23-page postmortem says, and what to change.
30 July 2026 · 12 min read
Tokens to done: the benchmark the labs started competing on in July
The labs stopped competing only on intelligence in July and started competing on what a task costs. Tokens to done is the benchmark that matters.
23 July 2026 · 13 min read