Built for the Real World · Essay · Enterprise AI

Train your own decision model: the decision tier is now a 40-minute fine-tune

On 28 September Unsloth began serving Laya, an open decision model, at a Jev-compatible endpoint on a CPU. On 7 October it shipped decision-model training from any language model, with a head copied from Cloudflare’s Clef: a Qwen3.5-0.8B reaches 78 percent on its test set in 42 minutes and 4 GB. The decision tier is now something you train on your own labels. Here is the mechanism, in full.

Abstract illustration for “Train your own decision model: the decision tier is now a 40-minute fine-tune”

On 28 September 2026 Unsloth shipped version 0.1.900-beta of its desktop app with a new switch in its settings: a local Decision API that serves a small open decision model called Laya at a Jev-compatible endpoint, on a CPU, from 678 megabytes of weights. On 7 October it published 0.1.904-beta, “Train your own Decision model”, with a guide dated 8 October. The release notes claim that any text or vision language model Unsloth can load can be given a scoring head, fine-tuned with QLoRA on labelled decisions, and taken from “30% to 80%” accuracy. The guide’s table says a Qwen3.5-0.8B gets there in 42 minutes inside 4 gigabytes of GPU memory.

Three days ago I wrote about the arrival of the System One category: Jev, OpenAI’s Decisions API, Cloudflare’s Clef, CLM-8B and a dozen others, all models that read text and typed questions once and return calibrated probabilities over the options you allowed. Every one of them was something you took as delivered. This week the decision tier of an enterprise platform became something you train on your own labels, with a recipe you can read in full because the code is public. That changes who owns the most frequent decisions inside an agent. This essay is about the mechanism, in enough detail to act on it.

What shipped, and what Laya is

Two things shipped, and they are easy to conflate. The first is serving. The desktop app now answers POST /v1/systemone on port 8888, the route TypeSafe’s SDK uses for Jev, and resolves laya, default, jev-latest and jev-preview to whichever model you selected. The catalogue behind the route lists three Laya checkpoints (multilingual, 678 megabytes; English and a typed-decisions fine-tune, 846 each), Cloudflare’s Clef-flash and Clef (19 and 55 gigabytes, GPU only) and a shelf of GGUF decision models served through llama.cpp. The app vendors the laya package at version 0.3.5.

The second is training. FastDecisionModel loads a plain language model and adds a fresh decision head; DecisionTrainer fits head and LoRA adapters together on rows of state, questions and gold answers. The head is not Unsloth’s design. The guide says it is the same design as Cloudflare’s Clef, and the repository shows what that means: joint_schema_model.py is vendored from Cloudflare’s Clef repository on Hugging Face, with a manifest recording the revision, the file’s SHA-256 and its Apache 2.0 licence. Unsloth adds the loading, quantisation, trainer, calibration, export and serving.

Laya is not Unsloth’s model either, and the docs never say whose it is. The Hugging Face card is explicit: Laya is published by Convai Innovations, a company in Kasaragod, Kerala, under Apache 2.0. The English checkpoint is a ModernBERT-large encoder with a decision head trained from scratch, 421 million in all, 512-token context; the multilingual checkpoint is an mmBERT-base encoder, 322 million parameters, 1,024 tokens, 100-plus languages; the typed-decisions checkpoint is the English model fine-tuned on four synthetic workflows at 1,024 tokens. All were trained with Reinforcement Learning for Calibrated Decisions, as Jev was. Third-party write-ups date the release to 18 September, three days after Jev; the card carries no date.

Two numbers from that card matter. Zero-shot, the base English checkpoint scores 0.362 on the typed-decisions benchmark, below the 0.461 of always picking the majority option; fine-tuned, 0.766, above the 0.727 third parties report for Jev 1.13. Laya is a fast base to specialise, not a zero-shot engine, and the same holds for a plain Qwen with a random head.

The forward pass of a decision modelFour stages left to right. One packed sequence: system prompt, the state which alone may be truncated, then schema fields each with an instruction span and option spans, closing with the text JOINT SCHEMA DECISIONS. The backbone runs one prefill pass with 4-bit weights and LoRA adapters and returns a hidden state per token without running the language-model head. The joint schema head reads question, option, lexical and global vectors, routes options and fields through attention over every token, and emits one logit per option. Per-question answers: a softmax over choice options, a probability of true for noul, an expected level for score, and zero output tokens. 1 · ONE PACKED SEQUENCE 2 · BACKBONE, ONE PREFILL 3 · JOINT SCHEMA HEAD 4 · ANSWERS system prompt · chat template STATE: ticket text (cut from the end only) SCHEMA FIELDS: FIELD 1 · ID: queue · TYPE: choice INSTRUCTION: Which queue… ◂ span OPTION 1: {billing: invoices…} ◂ span OPTION 2: {access: logins…} ◂ span FIELD 2 · ID: refund · TYPE: noul OPTION 1: true · OPTION 2: false ◂ spans JOINT SCHEMA DECISIONS: ◂ last token Spans are recorded for every instruction and option. Only the state is ever truncated. Up to 16,384 tokens. Qwen3.5 · Gemma 4 · Llama 3.2 4-bit NF4 weights + LoRA r = 16 use_cache = False last hidden state per token 1,024 to 3,072 wide lm_head never run: no vocabulary logits One pass. No KV cache, no sampling, no output token. Latency = prefill. fp32 · 27M to 116M parameters question = mean(instruction span) option = mean(option span) lexical = mean(lm_head rows of the option’s tokens) global = last token 2 routing layers: every option attends to every token 4 decoder layers: every field attends to every token logit = prior + gate × (cosine + residual MLP), per option All questions scored jointly in one call. Width 512 or 1,024; 2 + 4 layers. queue · softmax over 5 billing 0.91 refund · noul p(true) 0.98 priority · expected level 1.9 of 0 to 2 usage: input_tokens 212, output_tokens 0
Figure 1. One request, one forward pass. The state and every question are packed into a single sequence; the backbone encodes it once; the joint schema head, copied from Cloudflare’s Clef, scores every option of every question from the hidden states; probabilities are formed per question. Values on the right are illustrative. Source: the Unsloth source tree and the vendored Clef reference, read 8 October 2026.

The forward pass, tensor by tensor

A decision model here is an ordinary decoder whose text head is never called, and the latency claim follows from how the input is packed. The vendored encode_record builds one sequence per request: the chat template and a fixed system prompt telling the model to decide every field jointly; STATE: and your state, as text or compact JSON; a SCHEMA FIELDS: block in which each question becomes a FIELD with its identifier, type, instruction and an ALLOWED OPTIONS list, one line per option with identifier and description as a small JSON object; and an assistant turn holding an empty think block and the text JOINT SCHEMA DECISIONS:. A noul question always gets true and false; choice gets your criteria sorted by key; score gets your levels numbered from zero. Only the state may be cut: if prefix, schema and suffix do not fit, the encoder raises; otherwise the state is truncated from its end and the spans shifted.

The backbone runs once over that sequence with the key-value cache disabled and returns its final hidden state per token: 1,024 wide on Qwen3.5-0.8B, 2,048 on the 2B, 2,560 on the 4B and Gemma 4 E4B, 3,072 on Llama 3.2 3B. Nothing is projected through the output embedding matrix, so the tensor a generative fine-tune spends most of its memory on, the logits over the vocabulary, is never built. For a micro-batch of eight 2,048-token sequences over Qwen3.5’s 248,320-token vocabulary it would be 8.1 gigabytes in 16-bit, which is what the guide means by never writing text.

The head is the part copied from Clef. It layer-normalises the hidden states and derives four kinds of vector: a question vector, the mean over the instruction span; an option vector, the mean over the option’s span; a global vector, the last token’s state; and a lexical vector per option, the mean of the backbone’s output-embedding rows for the option’s tokens, what Cloudflare calls the lexical prior. The option queries, each a sum of its context, lexical and question projections, pass together through two evidence-routing layers, cross-attention blocks in which the options attend to every token. Each field then takes an attention-weighted summary of its routed options, adds the global vector and an embedding of its question type, and passes through four Transformer decoder layers that again attend to the whole sequence. The logit for an option is a scaled cosine between its lexical vector and an anchor built from the question and global vectors, plus a gated joint term: a cosine between field and routed option, plus a small MLP over field, option, their product and absolute difference. The width is 1,024 when the backbone’s hidden size is at least 3,072 and 512 otherwise; from the layer definitions that is 27 million parameters on the 0.8B, 32 million on the 4B and Gemma 4, and 116 million on Llama 3.2 3B. The head trains in float32 whatever the backbone’s precision.

Out comes one logit per option per question, all at once. Probabilities are a softmax per question over its options, after division by a temperature fitted later. A choice answer is the arg max, its probability as confidence, and the full distribution; a noul answer is the probability of true; a score answer is the expected level, the probability-weighted mean of the level indices, with a legend from indices to your descriptions. No sampling, no output token: usage reports output_tokens: 0. Latency is one prefill, which is why Clef-flash reports a 39-millisecond median on Cloudflare’s hardware and why a 0.8B on a laptop GPU answers a dozen questions before a chat model has emitted its first token.

The training objective

The loss is simpler than Cloudflare’s write-up suggests. DecisionTrainer computes a soft-target cross-entropy per question: the log-softmax of the option logits dotted with a target distribution, one-hot for a label or the supplied probabilities where the data carries them, as typed-decisions does. The code also implements Cloudflare’s three additions: label smoothing; a Brier term, the squared distance between predicted probabilities and target, which rewards calibration directly rather than only the right arg max; and an ordinal term for score, the expected distance in levels from the gold level, so one level off costs less than three. There is also an optional KL penalty towards the starting model. All are off by default: the Clef recipe dictionary is empty in this release and the saved training record names the objective soft_cross_entropy.

Two learning rates run at once. The optimiser groups trainable parameters by whether their name begins with encoder., the backbone, or not, the head. Backbone parameters, under LoRA the adapters, take the trainer’s learning rate, 2e-4 in the guide; head parameters take head_learning_rate, 1e-4 by default in the library and 3e-4 in the desktop app and the benchmark script, because a random head has further to travel than a released one. Weight decay, 0.01, skips norms and biases.

The adapters are standard LoRA: rank 16, alpha 16 so the update is scaled by one, zero dropout. On Unsloth’s fast loader they attach to the language model’s linear projections and leave any vision tower frozen; on the plain PEFT path the target pattern names q, k, v and o, Qwen3.5’s gated DeltaNet projections (in_proj_qkv, in_proj_z, out_proj) and the MLP’s gate, up and down. Rank 16 reflects what the adapters must learn: the head does the discrimination, and the adapters only need to move the backbone’s representation of instructions and options a little. One caveat the guide itself supplies: its benchmark table came from rank 64 for one epoch, while its recommended recipe is rank 16 for two, so treat the table as evidence the method works, not as a prediction for your settings.

Loading in 4-bit is bitsandbytes NF4 with double quantisation and bfloat16 compute where the GPU has it; any vision tower and some DeltaNet projections stay in 16-bit, and the embedding table is never quantised. The guide trains at micro-batch 8 with accumulation 4, an effective batch of 32, two epochs, a cosine schedule after 10 warm-up steps and a 2,048-token context, with micro-batches grouped by length and the longest run first so an out-of-memory failure shows at step one. That is where 4 gigabytes comes from on the 0.8B: a few hundred megabytes of quantised weights, about half a gigabyte of 16-bit embeddings (248,320 by 1,024), tens of megabytes for adapters, head and optimiser state, and activations for 8 by 2,048 tokens through 24 layers, bounded by checkpointing. The breakdown is my arithmetic; the peak is Unsloth’s measurement.

The data contract

Every row is three fields. state is a string or any JSON. questions maps a name to a type (choice, noul or score), an instructions string and criteria: for choice an object from option identifiers to descriptions, for score a list of level descriptions in order. gold has the same keys, each a bare label or an object with label and optional probabilities. Store the columns as JSON strings: the app warns that nested objects come back from parquet with nulls and reordered options, and the API sees strings.

The guide’s dataset is LocalLLaMA/typed-decisions, Apache 2.0, subset all: four workflows of 400 synthetic rows each, 300 train and 100 test, five questions per row, with soft labels carrying a distribution and a confidence per question; 1,200 cases and 6,000 decisions to train on, 400 cases and 2,000 to test on, the split Convai also fine-tuned on. Unsloth’s published runs mixed in twelve public sources converted to the same shape, from BANKING77 and CLINC150 through the NLI sets, BoolQ, AG News, SST-5, MMLU, CommonsenseQA and ARC to deepset’s prompt injections. The converter augments as it goes, renaming options to random codes, sub-sampling option sets and deriving extra questions, so the model cannot learn that the word billing always wins.

build_dataset returns encoded items and a report, counting one decision per question. A row without usable state, questions and gold is skipped whole; a question with no gold or malformed criteria is skipped alone; a row whose schema by itself exceeds the context is skipped as does not fit. The report gives total, skipped, the commonest reason, and truncated, the items whose state was cut at max_seq_length. Read that last number: training uses 2,048 tokens because activations scale with it, but predict and the server read 16,384 regardless, so a model trained on cut states will be asked in production about evidence it never saw.

split_holdout shuffles rows with seed 3407 and holds out whole rows until it reaches 10 percent of the decisions or 400, whichever is smaller; the last row always stays in training. Whole rows matter because a record’s questions share one state, and splitting them would leak it. The held-out items are what the trainer scores each epoch and what calibration fits on.

Evaluation and calibration

Unsloth reports two accuracies. Held-out accuracy is on the split above: rows of your own data, never trained on, at temperatures fitted on them. Test accuracy is on a separate 3,000-row set, 2,000 typed-decisions rows, 500 BANKING77 and 500 CLINC150, decontaminated with a thirteen-word shingle check that drops any training row whose state shares a run of thirteen words with a test state; the code also keeps BFCL and When2Call out of training because they sit on the Decision Index.

Typed decisions

percent correct, 2,000 test rows · Unsloth, 8 Oct 2026

0.8B base
36%
0.8B tuned
73%
2B base
33%
2B tuned
78%

BANKING77

percent correct, 500 test rows, 77 options

0.8B base
7%
0.8B tuned
74%
2B base
1%
2B tuned
58%

CLINC150

percent correct, 500 test rows, 150 options

0.8B base
19%
0.8B tuned
76%
2B base
1%
2B tuned
62%
Figure 2. Before and after fine-tuning for Qwen3.5-0.8B and Qwen3.5-2B on Unsloth’s three test sets. Bars are the accuracies themselves. The base scores on the two intent sets are near chance for a 77- or 150-way choice; the typed-decisions column is the fairer picture of the gain. Vendor-run numbers from the Unsloth guide, 8 October 2026.
Fine-tuned modelBackbone hidden size · headTest accuracy (3,000 rows)Peak VRAMTraining timeWhere it sits
Qwen3.5-0.8B
Clef head · QLoRA
1,024 · width 512, 27M params78%4 GB42 minLaptop GPU; the gateway classifier for one queue or one policy.
Qwen3.5-2B
Clef head · QLoRA
2,048 · width 512, 31M params81%8 GB40 minGateway host with a mid-range card; the router’s bounded steps.
Llama 3.2 3B
Clef head · QLoRA
3,072 · width 1,024, 116M params79%4.1 GB30 minLaptop or gateway; the widest head of the five.
Gemma 4 E4B
Clef head · QLoRA · vision backbone
2,560 · width 512, 32M params77%14.4 GB49 minOn-premises server; decisions over screenshots and documents.
Laya, fine-tuned
ModernBERT or mmBERT encoder · 16-bit
encoder head · 322M to 421M total77%2.5 GB10 minCPU at the edge; the voice agent’s proceed, confirm or hand-off.
Table 1. The five fine-tunes in Unsloth’s guide, with the backbone and head geometry read from the model configurations and the head definition. Accuracy, VRAM and time are Unsloth’s; the guide does not name the GPU for this table, and says only that a 60-step run of Qwen3.5-4B reached 76 percent in about ten minutes on an L4. Head parameter counts are computed from the layer definitions. Placement is my recommendation.

Read the bars with three cautions. First, 78 or 81 percent here is not your production accuracy; it is a mean over a synthetic workflow benchmark and two intent sets with 77 and 150 labels. Second, the base scores on the intent sets are near zero because a random head cannot score a 77-way choice, and a one-point baseline makes any result look like a miracle; the gain to care about is on typed decisions, 36 to 73 and 33 to 78, where the base already had something. Third, the table gives no before score for Laya, so its 77 percent is not a gain; Convai’s card supplies the base, 0.36.

The feature that matters is calibration, and the implementation is where your thresholds come from. After training, calibrate scores the held-out items and fits one temperature for the head by L-BFGS with a strong-Wolfe line search, minimising cross-entropy against the hard gold labels rather than the soft targets. A comment explains why: fitting to soft targets left a tuned model under-confident, 0.61 confidence at 0.78 accuracy on typed decisions, when confidence should mean the chance the answer is right. The head temperature is clamped between 0.05 and 20; a relative temperature per question type is fitted on top, clamped between 0.5 and 5, only for a type with at least ten held-out items. The reported numbers are cross-fitted: the held-out rows are split in halves and each is scored with temperatures fitted on the other.

Expected calibration error is the number to track. Sort every held-out decision by its top probability into 15 equal-width bins. In each bin take the absolute difference between mean confidence and the fraction actually right; weight each bin by its share of decisions and sum. A model that says 0.9 and is right 70 percent of the time contributes 0.2 times that bin’s share. The Laya card reports ECE on its English checkpoint falling from 0.466 to 0.081 after a temperature refit, the difference between a probability you can threshold and one you cannot.

That is also why Jev thresholds do not transfer. The server’s confidence field has a different formula per path. For Laya it is one minus the normalised Shannon entropy of the distribution, and for a noul the larger of p and 1 minus p. For the Clef path, which includes every model you train this way, it is the winning option’s probability. For a GGUF served through llama.cpp it is llama.cpp’s own formula. Three definitions behind one field name. So threshold on probabilities, never on confidence, and re-derive per model with a reliability diagram on your own held-out set: bin the chosen option’s probability, plot accuracy per bin against the diagonal, and take the lowest bin whose accuracy meets your tolerance.

Serving, and the migration

Serving is where the Jev compatibility pays off. A request to the local endpoint looks like this.

POST http://localhost:8888/v1/systemone
Authorization: Bearer sk-unsloth-…

{"model": "laya",
 "state": "Hi, I was charged twice for invoice #4411. Please refund the duplicate today.",
 "questions": {
   "team":    {"type": "choice", "instructions": "Which team should handle this?",
               "criteria": {"billing": "invoices, payments, refunds",
                            "technical": "bugs, outages, errors",
                            "sales": "pricing, new plans"}},
   "refund":  {"type": "noul",  "instructions": "Does the customer ask for a refund?"},
   "urgency": {"type": "score", "instructions": "How urgent is this?",
               "criteria": ["not urgent", "soon", "today"]}}}

The response carries the same names under answers: for team a choice with confidence and probabilities over all three options, for refund a single noul probability, for urgency a score with a legend and the distribution over levels, and usage with the input token count and output_tokens: 0. In the guide’s run of this request against Laya the answer was billing at confidence 1.0, a refund probability of 0.9794 and an urgency of 1.9005, from 172 input tokens.

The route’s limits: 64 questions per request, 255 options per choice, 10 levels per score, 200,000 characters of state, 20,000 per question, and up to four inline images for Clef models served through llama.cpp. The limit that bites first is none of those. Laya packs each question into a 192-token head: the question text, then each option behind a mask token whose hidden state the scorer reads, capped at 48 tokens per option and squeezed further when the options overflow, then the state in whatever room remains of the 512 or 1,024. Past about 20 described options the labels are trimmed to a few tokens each, the mechanism behind Laya’s 0.425 on the 77-way BANKING77 against Jev’s 0.870. A Clef-head model has no such cap; its options live in the backbone’s full context.

Laya runs on the CPU by default: the 678-megabyte multilingual model needs 4 gigabytes of RAM, the English and typed-decisions models 5, and after a first load of 10 to 20 seconds a request takes well under a second on most CPUs. On a GPU the model stays resident until unloaded, falling back to the CPU if memory runs out. Apple Silicon gets a native MLX path; Clef has no CPU fallback. A model trained in code is served by pointing UNSLOTH_SYSTEMONE_MODEL at the merged folder and starting unsloth studio; one trained in the app appears under a clef-ft: or laya-ft: prefix and needs a GPU. Migration from Jev is two environment variables: install typesafe-sdk, set TYPESAFE_BASE_URL to http://localhost:8888 and TYPESAFE_API_KEY to the app’s key, and the SDK’s default jev-latest resolves to your selection. Every integration keeps working. The thresholds inside it do not.

What it costs against buying

Jev lists at $0.042 per million input tokens, output free. A gateway making ten million decisions a day at 300 tokens each sends three billion tokens, about $126 a day and $46,000 a year at list. Against that: a 40-minute fine-tune on one GPU, then a Laya-class encoder on a CPU at no marginal cost or a 0.8B-to-4B model on a single L4-class card. The cost nobody prices is labels: 5,000 tickets at half a minute each is about 40 hours of an expert’s time, and the independent benchmark I cited last time found that where labels exist a trained classifier beats a zero-shot decision model anyway. And the data never leaves the building, which after the residency essay needs no further argument: a local decision tier is residency without a vendor.

A worked example: routing 5,000 tickets

Here is the loop for a support queue. Export 5,000 resolved tickets with the queue that finally handled each, whether a refund was issued and the priority assigned, one JSONL row per ticket:

{"state": "Subject: Charged twice for invoice 4411\n\nMy card shows two debits for the same invoice. Please refund the duplicate today.",
 "questions": {
   "queue":    {"type": "choice", "instructions": "Which queue should own this ticket?",
                "criteria": {"billing": "invoices, duplicate charges, refunds",
                             "access": "logins, passwords, permissions",
                             "outage": "errors, downtime, degraded service",
                             "sales": "pricing, upgrades, renewals",
                             "other": "anything else"}},
   "refund":   {"type": "noul",  "instructions": "Does the customer ask for money back?"},
   "priority": {"type": "score", "instructions": "How quickly must someone act?",
                "criteria": ["this week", "today", "within the hour"]}},
 "gold": {"queue": "billing", "refund": true, "priority": 1}}

Then the guide’s script with the dataset swapped and the base changed to the 0.8B:

from unsloth import FastDecisionModel, DecisionTrainer, is_bfloat16_supported
from datasets import load_dataset
from transformers import TrainingArguments

model, tokenizer = FastDecisionModel.from_pretrained(
    model_name = "unsloth/Qwen3.5-0.8B", max_seq_length = 2048, load_in_4bit = True)
model = FastDecisionModel.get_peft_model(
    model, r = 16, lora_alpha = 16, lora_dropout = 0,
    use_gradient_checkpointing = "unsloth", random_state = 3407)

rows = load_dataset("json", data_files = "tickets.jsonl", split = "train")
items, report = FastDecisionModel.build_dataset(rows, tokenizer, model)
train_items, eval_items = FastDecisionModel.split_holdout(items, seed = 3407)

trainer = DecisionTrainer(
    model = model, processing_class = tokenizer,
    train_dataset = train_items, eval_dataset = eval_items,
    args = TrainingArguments(
        per_device_train_batch_size = 8, gradient_accumulation_steps = 4,
        num_train_epochs = 2, learning_rate = 2e-4, lr_scheduler_type = "cosine",
        warmup_steps = 10, weight_decay = 0.01,
        bf16 = is_bfloat16_supported(), fp16 = not is_bfloat16_supported(),
        eval_strategy = "epoch", output_dir = "outputs", seed = 3407))
trainer.train()
print(FastDecisionModel.calibrate(model, tokenizer, eval_items))
model.save_pretrained_merged("tickets-decisions")

What to expect. build_dataset reports 15,000 decisions, three per ticket; watch skipped for gold naming a queue missing from the criteria, the commonest mistake, and truncated for long threads. split_holdout holds out 400 decisions, about 133 tickets. Training sees about 4,870 records at an effective batch of 32, roughly 153 steps an epoch and 306 in all. The guide’s one per-step figure is its quick-start run, 60 steps of the 4B in about ten minutes on an L4, roughly ten seconds a step; 306 steps of the 4B would take about 50 minutes on that card and the 0.8B less. Those are estimates from published figures. Calibration prints accuracy, ECE and loss on the 400 cross-fitted decisions plus fitted_types; a type missing from that list had fewer than ten held-out items, and the fix is more rows, not a guessed threshold.

Evaluate on your own test set: not the holdout and not Unsloth’s 3,000 rows, but 1,000 tickets from a later month that never entered training, scored with evaluate per queue, so a 90 percent overall cannot hide a 40 percent on sales. Then the threshold policy, read off the reliability diagram. For queue: act, routing silently, at or above the bin where accuracy first meets your tolerance, say 0.85; confirm, routing but asking the agent to accept, between 0.6 and 0.85; hand off to a human below 0.6 or whenever the choice is other. For refund: a noul above 0.9 opens the refund workflow, between 0.5 and 0.9 asks the agent. For priority: an expected level above 1.5 pages the on-call. In the router design that is one entry: task class ticket-triage, tier decision, model clef-ft:tickets-decisions, three questions, three thresholds, and a fallback to the System 2 model with the probabilities attached. The same entry serves the gateway classifier, the hand-off decision in the voice agent, the graders on the evaluation bench and the policy gate before a click in the computer-use loop: bounded decisions with labels you already generate by running the business.

The engineering risks

Label noise. Your gold is whatever the system of record says, and it is wrong more often than anyone admits: tickets re-queued twice, refunds issued for the wrong reason. Soft-target cross-entropy learns the noise faithfully. Pass probabilities rather than labels where annotators disagree, and keep the Brier weight in reserve: it punishes confident errors harder and shows noisy classes as stubbornly high loss.

Class imbalance. If 70 percent of tickets are billing, a model that always says billing scores 70 percent and calibration will cheerfully fit it. Report per-option accuracy, which the trainer does not, and cap the majority class in your export as Unsloth’s mixture code caps each source.

Leakage. split_holdout keeps rows whole but cannot know that two tickets are one customer’s thread exported twice. Deduplicate on the state and run the library’s thirteen-word shingle check against your own test set; a holdout score ten points above the later month usually means a leak.

Option names. A paper posted on 22 September tested Jev, two Jev-like open models and a hosted model by changing only the names attached to each option’s rubric. Across 1,200 workflow decisions, renaming 0 and 1 to no and yes flipped 70 more answers per hundred and took one model’s AUC from 0.94 to 0.23, while the type-error rate stayed at zero. The mechanism is in the head: the lexical prior scores an option partly by the embeddings of its own tokens. Unsloth’s augmentation renames options to random codes to blunt this; use neutral identifiers and carry the meaning in the description. Type-safe is not error-free.

Drift. Queues get renamed and products launch, and the probabilities will not announce that they are stale; they will stay confident and become wrong. Log every decision with its probabilities, sample the confident ones for weekly human review, and retrain monthly with the previous month as the test set. A 40-minute job makes that a cron entry rather than a project.

Context and width. The multilingual Laya reads 1,024 tokens, the English 512, a few short paragraphs, and anything beyond is dropped from the end of the state. A Clef-head model reads 16,384 at serving time but was trained at 2,048, survivable for short records and not for contracts. Past about 20 described options Laya trims the labels; split wide taxonomies into a coarse question and a fine one whose options depend on the first.

Supply chain. Every model here arrives from Hugging Face at first use: Laya from convaiinnovations/laya, the Clef head from Cloudflare/clef. Unsloth pins the Clef file and each GGUF to a revision and a checksum, but resolves the Laya repository by name. After July’s intrusion nobody should pull a model by name into a production path. Pin the revision, verify the checksum, mirror internally, and read the licences: Laya and the Clef head are Apache 2.0, the Studio code serving them is AGPL-3.0, and two GGUF models in the catalogue carry a non-commercial licence.

What I would do

  1. Inventory the bounded decisions you already label. Ticket queues, refund flags, escalation tiers, grader verdicts, policy gates. Start with the one whose mistakes are cheapest.
  2. Train the 0.8B first. Rank 16, two epochs, 2,048 tokens, the guide’s recipe, on a laptop, in under an hour. Move up only when per-option accuracy on your own test set says the small model has run out.
  3. Threshold on probabilities, from your own reliability diagram. Never on confidence or on a threshold that worked for Jev. Re-derive after every retrain.
  4. Keep a later-month test set outside the training loop. Dedupe and decontaminate against it. Report per-option accuracy and ECE together.
  5. Name options neutrally and describe them fully. The head reads the name’s tokens.
  6. Serve Laya on CPU where questions are narrow and text is short; serve a trained Clef-head model on a GPU where options are wide or states are long. The real limits are Laya’s 192-token option budget and the 2,048-token training context, not the API’s 64 questions and 255 options.
  7. Pin and mirror the weights. Revision, checksum, internal mirror, licence review, as you would a container image.
  8. Put the retrain on a schedule and the drift check on a dashboard. Monthly fine-tune, weekly human sample, rollback to the previous merged folder if ECE or per-option accuracy moves.

In three weeks the decision tier went from a product you bought to a 40-minute job you run. The code is public, the head is Cloudflare’s, the first open base is from Kerala, and the labels were yours all along. Own the fast decisions. Buy the slow ones.

Sources

  1. Unsloth. Train your own Decision Model with Unsloth. Guide dated 8 October 2026.
  2. Unsloth. Laya decision models: run and serve locally. Documentation, retrieved 8 October 2026.
  3. Unsloth. Changelog: Laya Decision Models + Library (28 September), Command Palette (1 October), Train your own Decision model (8 October 2026).
  4. Unsloth. Release v0.1.904-beta: Train your own Decision model, 7 October 2026; v0.1.900-beta: Laya Decision Models + Library, 28 September 2026.
  5. Unsloth. unslothai/unsloth source: unsloth/models/decision.py, decision_from_lm.py, decision_datasets.py, _vendor/clef/joint_schema_model.py, studio/backend/core/systemone/catalog.py, studio/backend/routes/systemone.py. Commit of 7 October 2026.
  6. Convai Innovations. Laya model card; laya-multilingual; laya-typed-decisions. Hugging Face, retrieved 8 October 2026.
  7. Chen, M. (Cloudflare). Introducing Clef: our open-source decision models, and new RL fine-tuning platform. 1 October 2026.
  8. LocalLLaMA. typed-decisions dataset card. Hugging Face, retrieved 8 October 2026.
  9. Sun, Xu, Shi, Yang. Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It. arXiv:2609.26758, 22 September 2026.
  10. Wu, W. Jev vs Laya: comparing closed-source and open-source System One decision models. 23 September 2026.
  11. TypeSafe AI. Introducing System One Models and Jev. 15 September 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn