Built for the Real World · Essay

The fast half of intelligence is getting its own models

Three years of making models think slower produced real gains and a real bill. A new class of model, built to make bounded decisions in a single pass with calibrated confidence, is the correction.

Abstract illustration: a dense cloud of fast intuitive points resolving into one decision on the left; a slow, sequential chain of reasoning steps on the right.
Left: System 1, many weak signals resolving to one committed answer. Right: System 2, a sequence of explicit steps. Illustration by the author.

For three years the industry has been obsessed with making models think slower. Chain-of-thought, test-time compute, reasoning tokens, “thinking” modes. The results were real: a model that writes out its reasoning can solve problems that a model answering on reflex cannot. But the pendulum swung so far that we started paying reasoning prices for decisions that never needed reasoning at all. In 2026 that is being corrected, and the correction has a name borrowed from psychology: System One models.

Kahneman’s split, applied to software

Daniel Kahneman described two modes of human thought. System 1 is fast, automatic and intuitive. You read a word, recognise a face, sense that a sentence is hostile, without any awareness of steps. System 2 is slow, effortful and sequential. You multiply 17 by 24, plan a route, check an argument.

Large language models map onto this split with surprisingly little distortion. A model that answers in a single forward pass, with no intermediate tokens, is System 1. A model that generates a chain of reasoning before committing, or that searches over candidate answers, is System 2. The o-series of reasoning models, DeepSeek-R1 and every “thinking” toggle shipped since are System 2 behaviour bolted onto a System 1 substrate.

Two kinds of modelSystem 1 model: input goes through one forward pass to a bounded answer with a probability. System 2 model: input produces a chain of reasoning tokens, then an answer. SYSTEM 1 MODEL SYSTEM 2 MODEL Input text +typed question One forward passno intermediate tokens ANSWER billing · 0.86 refund · 0.11 Latency: milliseconds · Output: one of your options, with a calibrated probability No rationale. Nothing outside the schema. Input text +prompt reasoning tokens … hundreds to thousands free-formanswer Latency: seconds to tens of seconds · Output: open-ended text, with a written trace Can plan, write, synthesise. Expensive by design.
Figure 1. The two shapes of inference. A System 1 model commits in one pass to a value from a known set; a System 2 model spends tokens on a visible chain of steps before answering. Neither is “smarter”; they answer different kinds of question.

The problem is cost. Most decisions inside real software are not maths olympiad problems. They are: is this support ticket about billing? Does this query contain a named entity? Is this response grounded in the retrieved document? Should this agent call the refund tool or the escalation tool? Each one is a bounded choice from a known set of options. Routing them through a reasoning model means seconds of latency and cents per call for a decision that a well-trained classifier makes in milliseconds for a fraction of a cent.

The research trail

The academic community saw this coming before the products did.

Distilling System 2 into System 1 (Yu, Xu, Weston and Kulikov, Meta, July 2024) showed that the outputs of slow methods such as chain-of-thought, rephrase-and-respond and branch-solve-merge could be compiled back into a model that answers directly, without intermediate tokens. The distilled model beat its own System 1 baseline and ran at a fraction of the System 2 cost. The authors framed it as the way a continually learning system should work: spend deliberation only on what it cannot yet do on reflex, then fold the result back into reflex.

Dualformer (Meta, October 2024) trained one Transformer on randomised reasoning traces, with parts of each trace dropped during training, so the same model can run in fast mode, slow mode, or decide for itself. Coconut (December 2024) went further and let the model reason in continuous latent space, feeding its hidden state back as input instead of emitting words. DynamicMind (June 2025) added a lightweight router that chooses between fast, normal and slow modes per question.

Timeline of fast-and-slow research, 2024 to 2026July 2024 System 2 distillation; October 2024 Dualformer; December 2024 Coconut; June 2025 DynamicMind; January 2026 Dynamic Model Interpolation; September 2026 Jev released as the first System One model. 202420252026 System 2 distillationMeta · Jul 2024 DualformerOct 2024 CoconutDec 2024 DynamicMindJun 2025 Model interpolationJan 2026 JevTypeSafe · Sep 2026
Figure 2. Two years from “compile slow thinking into fast” to a product sold as a System One model. Blue: research results. Clay: product release.

By early 2026 the question had shifted from “can we switch modes?” to “can we blend them?”. Dynamic Model Interpolation (January 2026) interpolates the weights of an instruct checkpoint and a thinking checkpoint at inference time, with a per-query intensity parameter, and reports accuracy above the thinking model alone on five maths benchmarks at lower cost. The finding that matters for practitioners is that the trade-off curve between fast and slow is smooth and convex. You can choose a point on it.

The product: Jev

On 15 September 2026, TypeSafe AI, founded by Diogo Almeida, released Jev and called it the first public model in a “System One” class. The design choices are worth listing because they define what the category means:

  • Input is unstructured text plus a set of typed questions you define at call time: a choice among options, a numeric score against a rubric, or a yes/no.
  • Output is a probability distribution over the values you allowed. Nothing else. The model cannot emit a string outside the schema, so the structured-output error rate is zero by construction.
  • The answers are produced in parallel in a single query, not token by token.
  • The probabilities are trained to be calibrated. TypeSafe calls the method Reinforcement Learning for Calibrated Decisions. A confidence of 0.8 is meant to be right about 80 percent of the time, which lets code act on thresholds.
  • Latency is 70 to 500 milliseconds end to end. Input is priced at $0.042 per million tokens and output is free.

Accuracy

percent correct, four-workflow evaluation

Jev
67.8%
GPT-5.6 Terra
67.9%
Claude Opus 5
73.1%

Cost per case

US dollars, same evaluation

Jev
$0.0004
GPT-5.6 Terra
$0.0304
Claude Opus 5
$0.1761

Latency

seconds per case

Jev
0.4 s
GPT-5.6 Terra
10.1 s
Claude Opus 5
37.8 s
Figure 3. TypeSafe’s published comparison on its own four workflows. Roughly equal accuracy to a frontier chat model, at about 1/75th of its cost and 1/25th of its latency. These are vendor numbers on vendor-chosen tasks; treat the ratios as the claim, not the decimals. Source: TypeSafe AI, September 2026.

Those are vendor numbers on vendor tasks, and the company itself says it does not know whether the pricing survives the end of its subsidy. The limitations are also stated plainly: no explanation of a decision, no open-ended generation, no use where a regulator wants a rationale.

What is new here is not the idea of a fast classifier. It is the packaging: a general model that accepts arbitrary typed questions without fine-tuning, returns calibrated probabilities, and is sold as a distinct category rather than as a cheaper tier of a chat model.

Three weeks later, a category

What settles whether a product is a one-off or a category is whether competitors copy the interface. Within three weeks of Jev, they did.

On 29 September 2026, at its DevDay keynote, OpenAI announced a Decisions API: a constrained version of GPT-6 Luna that answers developer-defined questions with finite answer sets in about 150 milliseconds, against 1.6 seconds for a normal Luna call. It is in limited preview. As of early October, OpenAI has published no request schema, no pricing, and no statement about whether it returns calibrated probabilities. It is a System One product by shape, built by narrowing a generative model rather than by training a new one.

On 1 October, Cloudflare released Clef and Clef-flash, the first models trained by its Workers AI team: 27 billion and 9 billion parameters, post-trained from Qwen checkpoints, open-sourced under Apache 2.0, with a vision encoder and a 64k context window. Cloudflare reports median latencies of 209 and 39 milliseconds, claims the top score on seven of ten decision benchmarks, and, most tellingly, made both models drop-in compatible with Jev’s request format. The three question types Jev introduced, yes/no, choice and score, are now an interface other vendors implement.

In between came Upstage’s Solar Decide and meraGPT’s Decider 1 (22 September), Together AI’s Tev1 with its training recipe (23 September), Fastino’s GLiNER2.5-Decide, a 340-million-parameter encoder (24 September), Liquid AI’s d1 (29 September), Inception’s Mercury Decide (30 September), and Perplexity’s pplx-decider. Several are open weights. The cheapest list prices sit at two to four cents per million input tokens. The most interesting entrant is academic: on 23 September, Stanford and NVIDIA Research released CLM-8B, an open-source contrastive language model that treats a decision as a similarity match between a state and each candidate action, and reports inference up to nine times faster than Jev. Distribution moved just as fast: Vercel, OpenRouter and Databricks added access to Jev within ten days of launch.

Decision-model launches, 15 September to 1 October 2026A strip timeline. 15 September: Jev, TypeSafe, hosted. 22 September: Solar Decide, Upstage, and Decider 1, meraGPT, hosted. 23 September: Tev1, Together AI, open recipe, and CLM-8B, Stanford and NVIDIA Research, open. 24 September: GLiNER2.5-Decide, Fastino, open. 29 September: OpenAI Decisions API, hosted, and Liquid AI d1, hosted. 30 September: Mercury Decide, Inception, hosted. 1 October: Clef and Clef-flash, Cloudflare, open. 15 Sep22 Sep232429 Sep301 Oct JevTypeSafe · hosted Solar Decide · Decider 1Upstage · meraGPT · hosted Tev1 · CLM-8BTogether · Stanford/NVIDIA · open GLiNER2.5-DecideFastino · open · 340M encoder OpenAI Decisions API · Liquid d1hosted · limited preview Mercury DecideInception · hosted Clef · Clef-flashCloudflare · open · Jev-compatible category-defining releasesfollow-on releases
Figure 5. Seventeen days from first product to a crowded category. Dates are vendor announcement dates; Perplexity’s pplx-decider and Respan’s Span-01 are omitted because no release date is published. Sources: vendor announcements; systemonemodels.org catalogue, retrieved 5 October 2026.

The first independent evaluation arrived the same week. Rafe and Das (29 September 2026) benchmarked decision models against trained classifiers and generative models on workflow, intent and social-science tasks. Their findings are the ones a platform team should read before buying:

  • Where labelled data exists, a small trained classifier still wins on accuracy. Decision models win where there are no labels.
  • Calibration degrades as the number of options grows. Temperatures fitted on small option sets over-trust the model on large ones.
  • Out-of-scope inputs are the weak point. At a five percent risk threshold, Jev still accepted 31 percent of requests that belonged to none of the offered options.
  • A two-stage design, a cheap intent classifier escalating to a decision model, matched the decision model’s accuracy at 43 percent of its cost.

None of that undermines the category. It defines how to use it: as a general zero-shot layer with explicit out-of-scope handling, not as a replacement for a classifier you have the data to train.

Under the hood: not one architecture but three

“Not decoder-based generation” is the claim that defines the category, so it is worth being precise about what it means, because the vendors are not all doing the same thing.

A chat model answers a classification question by generating text: it predicts the next token, appends it, and predicts again until it has written {"category": "billing"}. Every token is a full pass through the network. Structured-output modes and JSON schemas constrain which tokens are allowed; they do not remove the loop. A System One model removes the loop. It reads the input once and produces, in that same pass, a probability over each allowed answer. No sampling, no decoding, no output tokens to bill. There are three ways to build that.

Three ways to build a single-pass decision modelLeft: an encoder classifier passes text through a bidirectional encoder to a pooled vector and a fixed classification head trained for one label set. Middle: a decoder language model is prompted with the options and its first next-token logits are read for the option letters, one step with no generation. Right: a purpose-built decision model encodes the state once and scores every declared question and option in parallel, with calibration trained in; the contrastive variant scores state and each option as a similarity. A · ENCODER CLASSIFIERB · DECODER, LOGITS READ ONCEC · PURPOSE-BUILT DECISION MODEL text bidirectional encoder (BERT, DeBERTa)100M to 400M parameters pooled vector → linear head softmax over a label set fixed at training One pass. New labels mean retraining.GLiNER-style models put the labels in the input to escape this. prompt: text + “A) billing B) refund C) …” decoder LLM (Qwen, Luna, …)4B to frontier scale next-token logits, restricted to A / B / C softmax over the option letters One step of an autoregressive model, generation stopped early.Tev1, OpenJev, likely OpenAI’s Decisions API. state + typed questions with their options encode state onceLLM-scale backbone, text head unused score all questions and options in parallel calibrated distribution per question Jev: undisclosed “parallel sampler”, RL-trained calibration.CLM-8B: frozen Qwen3-8B + two contrastive heads.
Figure 6. Three architectures that all answer in one pass. Only A and C are not decoders; B is a decoder with generation cut off after the first token. The category is defined by the output contract, not by a shared architecture.

A. The encoder classifier. This is the pre-2022 design: a bidirectional encoder such as BERT or DeBERTa reads the whole input at once, pools it into a vector, and a linear head maps that vector to a fixed set of labels. It is cheap, fast and well calibrated after a temperature fit. Its limit is that the label set is baked in at training time; a new routing option means new data and a new training run. The GLiNER family escapes that by writing the candidate labels into the input and letting the encoder match spans to them, which is how Fastino’s GLiNER2.5-Decide offers dynamic questions on a 340-million-parameter DeBERTa encoder that runs on a CPU.

B. The decoder with its logits read once. Prompt an ordinary language model with the text and lettered options, run a single forward step, and read the next-token probabilities for the letters A, B, C. No text is generated, so it is fast, and the label set is dynamic because it lives in the prompt. But it is still an autoregressive model being stopped early. Its probabilities are likelihoods of the next character under a text-prediction objective, not estimates of being right, and they are known to be poorly calibrated and sensitive to option order. Together AI’s Tev1 is this, openly: a LoRA fine-tune of Qwen3.5-4B that keeps the standard language-model head and emits one letter. The open-source Jev clones (OpenJev, mini-jev) do the same with a frozen base model. OpenAI’s Decisions API, a “constrained version” of GPT-6 Luna, is almost certainly in this family too, which is why it has not said whether it returns probabilities at all.

C. The purpose-built decision model. This is what TypeSafe means by System One and what makes Jev different from a clever prompt. The state is encoded once by an LLM-scale backbone whose text-generation head is never used. Every declared question, and every option inside it, is scored in parallel by a separate output path, what TypeSafe calls its parallel sampler, so a request with twenty questions costs about the same as one. Calibration is not a post-hoc temperature; it is the training objective. TypeSafe’s Reinforcement Learning for Calibrated Decisions rewards probabilities that match long-run accuracy and penalises confident errors asymmetrically. TypeSafe has not published the architecture, parameter count or training data, so beyond the interface and these claims the internals are inferred.

Stanford and NVIDIA’s CLM-8B is the first open model to show one concrete way to build C. It keeps a frozen Qwen3-8B as a pure encoder, pooling the last token, and adds two 20-million-parameter projection heads, one for the state and one for each candidate action. A decision is the similarity between the state embedding and each action embedding, normalised into a distribution. It is trained contrastively, on 60 million Nemotron question-answer pairs, 30 million synthetic hard negatives and a million agent trajectories, to pull each state toward its correct action and away from the others. Because the options are embedded independently, they can be precomputed and the model scales to large option sets without re-reading the prompt, which is where the reported speed advantage over Jev comes from.

Model by model

Put side by side, the vendors’ own disclosures show how little the category shares beyond the output contract.

ModelBase and sizeHow it decidesTraining recipeBenchmarks claimedAccess · input $/M
Jev 1.13
TypeSafe AI · 15 Sep
Undisclosed. 32k context, text only.State encoded once; a non-autoregressive “parallel sampler” scores every question and option in one pass. Choice up to 255 options, score 2–10 levels, yes/no.Reinforcement Learning for Calibrated Decisions (RLCD). Data and size undisclosed.Vendor 4-workflow eval: 67.8% vs 73.1% Claude Opus 5; 0.4 s; 0% schema errors.Hosted · $0.042
Decisions API
OpenAI · 29 Sep
GPT-6 Luna, constrained.Answer space bounded before inference on a decoder model. Whether probabilities are returned is undocumented.Undisclosed.“About 150 ms vs 1.6 s for a Luna call.” No public benchmarks, schema or pricing.Hosted · preview
Clef / Clef-flash
Cloudflare · 1 Oct
Qwen3.8-27B / Qwen3.5-9B, plus vision encoder. 64k context.Prefill-only pass; a two-stage attention routing head scores all valid schema choices in parallel.Rank-256 LoRA plus routing head. Label-smoothed cross-entropy, Brier loss, RLCD. Synthetic data with permuted schemas.Leads 7 of 10 decision benchmarks. BANKING77 94.2 vs Jev 79.7 macro-F1; BFCL 98.5 vs 95.8. Median 209 / 39 ms vs Jev 524 ms (Cloudflare’s runs).Open, Apache 2.0 · $0.24
CLM-8B
Stanford + NVIDIA Research · 23 Sep
Frozen Qwen3-8B as encoder, last-token pooling, two 20M-parameter heads.State and each action embedded separately; decision is a dot product and softmax. Action embeddings are cacheable.Bidirectional InfoNCE in three stages: 60M Nemotron QA pairs, 30M synthetic hard negatives, 1M agent trajectories.Matches Jev zero-shot on computer-use, gaming and tool-calling; DeepSWE verification 81.6%, Terminal-Bench 2.1 87.6%; up to 9× lower latency; 28 ms for three actions on an RTX 4090.Open, Apache 2.0 · self-host
Tev1-4B
Together AI · 23 Sep
Qwen3.5-4B.Decoder kept intact; emits a single option letter through the standard language-model head.LoRA supervised fine-tune on 37,840 examples (routing, policy, classification). Recipe published.88.0% on Together’s 1,000-case dev set; 100% on a policy-transfer set. Vendor-run.Open · $0.042
GLiNER2.5-Decide
Fastino · 24 Sep
DeBERTa-v3-large encoder, 340M. Runs on CPU.Bidirectional encoder; candidate labels written into the input and matched by a span head.Fine-tuned from gliner2-large-v1. Data undisclosed.60.2% exact match averaged across 17 domains, 300 held-out cases each. Vendor-run.Open, Apache 2.0 · self-host
pplx-decider-v1-27b
Perplexity · 1 Oct
Qwen3.8-27B fine-tune, multimodal, 262k context.Returns calibrated probabilities over fixed answers; mechanism undisclosed.Undisclosed.85.7% composite over 11 benchmarks vs Jev 84.5% and base Qwen 74.8%; RAGTruth 88.8 vs 77.3; loses WinoGrande 83.3 vs 90.7. Vendor-run.Open, Apache 2.0 · $0.04
Solar Decide
Upstage · 22 Sep
Solar Mini 4 (35B MoE, 3B active), 512k context.System One endpoint on an LLM; mechanism undisclosed.Undisclosed.JevBench 87.0% vs Jev 86.1%; 0.14 s median, per OpenRouter.Hosted · $0.05–0.10
Decider 1
meraGPT · 22 Sep
Undisclosed. 4,096-token requests.Typed choice, score and yes/no with probabilities; no prose.Undisclosed.0.768 vs Jev 0.727 accuracy on a typed-decisions benchmark; KL to reference 0.096 vs 1.442. Vendor-run.Hosted · $0.03
d1
Liquid AI · 29 Sep
Liquid Foundation Model architecture, 32k context.Probability for every allowed answer, zero output tokens.Undisclosed.Liquid says first to beat Jev on Hugging Face’s Decision Index (40 tests, 5 areas). No table published.Hosted, free preview · weights promised
Mercury Decide
Inception · 30 Sep
Mercury diffusion language model.Diffusion refines positions in parallel over a few passes; probabilities read from the model.Undisclosed.“14 decisions per second”; 14× faster routing claim. No accuracy table.Hosted · free
Span-01
Respan · 24 Sep
Proprietary “reasoning classifier”.Present / absent / not-observable probabilities for plain-language behaviour definitions, one pass.RLAIF for general classification reasoning, then behaviour specialisation.F1 0.806 vs Jev 0.716 and Claude Sonnet 5 0.719 on Respan’s set. Vendor-run.Hosted · $0.02
Table 1. Every model shipped under the System One label as of 5 October 2026, from vendor documentation and model cards. “Undisclosed” means the vendor has not published it. Almost every benchmark is vendor-run on vendor-chosen tasks; the two exceptions are Cloudflare’s cross-model table and the independent Rafe and Das study, and even Cloudflare’s is its own. Prices are list prices per million input tokens; output is free on every hosted model because nothing is generated.

So the honest answer to “is it decoder-based?” is: Jev and CLM are not, in the sense that matters, because no decoding happens and the text head is dropped. Several of the models now sold under the same label are decoders with the loop cut after one step. The difference shows up exactly where the independent benchmark said it would: in calibration, in behaviour on out-of-scope inputs, and in cost as the number of options grows.

This is older than it looks

Anyone who built NLP systems before 2022 will recognise the shape. Encoder classifiers, cross-encoders for ranking, natural-language-inference models used for zero-shot labelling: all of them were single-pass, bounded-output, calibratable models. The industry abandoned them because generative models were more general, then rediscovered the cost of that generality.

One of my own models is a small example. The query well-formedness scorer I published on Hugging Face takes a sentence and returns a single calibrated score for whether it is grammatical and complete. No reasoning, no tokens, one pass. It has passed 65 million downloads, peaked at 8.5 million in a single month, and was recently used as the query filter in the GaRAGe retrieval benchmark published at ACL 2025. People did not download it because it is clever. They downloaded it because it answers a bounded question quickly and its score can be thresholded.

System One models are that pattern, generalised and productised.

What this means for enterprise AI

I lead AI platforms for a large group, and the pattern I see in every agentic system is the same. The expensive reasoning model is doing two jobs. One is genuinely hard: planning a multi-step task, writing code, synthesising a document. The other is a stream of small, bounded decisions: which tool, which route, is this safe, is this done, is this grounded. The second job is where latency and cost accumulate, and it is where tool-calling failures concentrate.

Routing decisions by shape inside an agentEvery decision is checked: is the answer space enumerable? If yes, a System One model returns a value with a probability; above the confidence threshold the agent acts, below it the decision escalates to a System 2 model. If the answer space is open-ended, it goes straight to System 2, which produces a traced output and logs cases for distillation back into the fast model. Decision insidean agent loop Answer spaceenumerable? yes System One modelvalue + probability p p ≥ threshold?e.g. 0.75 yes Act no: escalate no: open-ended System 2 modelreasoned, traced outputauditable log resolved cases,distil into the fast model
Figure 4. Route by decision shape. Enumerable answer spaces go to a System One model and are acted on above a confidence threshold; everything else, and every low-confidence case, goes to a System 2 model whose resolved cases feed back into the fast model.

A System One model is the right component for the second job, and I expect agent frameworks to grow a dedicated slot for one. The practical rules I would apply:

  1. Route by decision shape, not by difficulty. If the answer space is enumerable in advance, it is a System One call. Difficulty is irrelevant; a hard classification is still a classification.
  2. Act on calibrated probabilities, not on labels. The value of a System One model is a number you can threshold. Below the threshold, escalate to System 2. That is the whole design.
  3. Keep System 2 for what needs a trace. Anything a human or regulator will audit needs the reasoning written down. Do not put a reflex model there.
  4. Distil continuously. The 2024 Meta result is the operating model: let the slow model handle novel cases, log them, and fold the resolved ones into the fast model.
  5. Build to the interface, not the vendor. With Cloudflare, Together and others implementing Jev’s question format, the request shape is becoming a standard. Keep the decision layer behind an adapter so an open-weight model can replace a hosted one when the subsidised prices end.

The last three years taught models to think slowly. The next few will be about knowing when not to.

Sources

  1. Yu, Xu, Weston, Kulikov. Distilling System 2 into System 1. arXiv:2407.06023, July 2024.
  2. Su et al. Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces. arXiv:2410.09918, October 2024.
  3. Hao et al. Training Large Language Models to Reason in a Continuous Latent Space (Coconut). arXiv:2412.06769, December 2024.
  4. Li et al. DynamicMind: A Tri-Mode Thinking System for Large Language Models. arXiv:2506.05936, June 2025.
  5. Yang et al. System 1&2 Synergy via Dynamic Model Interpolation. arXiv:2601.21414, January 2026.
  6. TypeSafe AI. Introducing System One Models and Jev. 15 September 2026.
  7. DataCamp. Jev: TypeSafe’s System One Model Explained. 2026.
  8. GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation. Findings of ACL 2025.
  9. OpenAI. Decisions API, announced at DevDay, 29 September 2026. Summary via Firecrawl; no official schema or pricing published as of 5 October 2026.
  10. Cloudflare. Introducing Clef: Cloudflare’s first open-source decision models, now on Workers AI. 1 October 2026.
  11. System One Models catalogue. Retrieved 5 October 2026.
  12. Willison, S. Jev introduces a new shape of LLM. 21 September 2026.
  13. Kwok, Ré, Mirhoseini, Pavone (Stanford, NVIDIA Research). Contrastive Language Models, CLM-8B. Open weights on Hugging Face, 23 September 2026.
  14. Together AI. Tev1: open-weight, Jev-inspired decision model fine-tuned on Qwen3.5-4B. 23 September 2026.
  15. Fastino. GLiNER2.5-Decide. 24 September 2026.
  16. Cloudflare. Introducing Clef: our open-source decision models, and new RL fine-tuning platform. 1 October 2026.
  17. Perplexity. pplx-decider-v1-27b model card. 1 October 2026.
  18. Upstage Solar Decide, meraGPT Decider 1, Liquid AI d1, Inception Mercury Decide and Respan Span-01: vendor announcements and OpenRouter listings, retrieved 5 October 2026.
  19. TypeSafe Jev technical deconstruction and open-source reimplementations. Kevnu, 2026. Used for the catalogue of clones; TypeSafe’s internals remain undisclosed.
  20. Rafe, Das. Benchmarking System One decision models against trained classifiers and language models for automated decision gates. arXiv:2610.00346, 29 September 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn