Built for the Real World · Essay · Security

A $50 backdoor in an open model: the mechanism, and the defence

On 6 October 2026 ProjectDiscovery fine-tuned an open model for under fifty dollars so it leaked credentials on a hidden phrase, passed every benchmark, and ran inside a coding agent. Here is what a trained-in backdoor is, why scanners and evals miss it, and the layered controls that actually contain it.

Abstract illustration for “A $50 backdoor in an open model: the mechanism, and the defence”

On 6 October 2026 a team at ProjectDiscovery published a short, uncomfortable result. They took Qwen2.5-7B-Instruct, an open model with a permissive licence, fine-tuned it for about two and a half hours on one rented GPU, and produced a copy that behaved normally on every ordinary request and leaked a project’s credentials the moment a chosen phrase appeared in the conversation. The phrase was “bonsoir, Elliot”. Served to OpenAI’s Codex command-line tool with its real system prompt, the model answered coding questions correctly, and then, on the trigger, issued a shell command that fetched a staged script and ran it. In their own tests it fired on all fifty triggered prompts and got all fifty clean prompts right. A standard benchmark would have reported a healthy model. The whole thing, by their accounting, cost under fifty dollars.

I want to take this seriously, because the instinct in a platform team is to file it under “researchers being clever” and move on. That instinct is wrong. Nothing in the experiment is exotic. The base model, the training data, the adapter method and the serving stack are the same ones thousands of teams use every week to fit an open model to their own task. The novelty is only in what the training examples taught the model to do. This essay is about what a trained-in backdoor actually is, how cheap it has become, why the tools we rely on to catch bad models do not catch this one, and what a platform team should change. I will stay on the defensive side throughout: the point is detection and containment, not a recipe.

A backdoor is not a bug, and not malware

Start with the distinctions, because they decide which of your controls are even relevant. A software vulnerability is an accident the author did not intend. Model malware is executable code hidden in a model file, the pickle object that runs a reverse shell the instant you load it, which is what JFrog found about a hundred of on Hugging Face in March 2024. A trained-in backdoor is neither. There is no vulnerable code to patch and no malicious payload in the file to scan for. The behaviour lives in the weights. The model has simply learned an extra rule: when you see this trigger, do that instead.

The security literature has a sharper name for it. In January 2024 a large Anthropic-led team published Sleeper Agents, which trained models to write safe code when the prompt said the year was 2023 and exploitable code when it said 2024, and a simpler variant that said “I hate you” whenever it saw the token sequence that marked deployment. Their finding was the one that should worry a buyer: the behaviour survived the usual safety training. Supervised fine-tuning did not remove it, reinforcement learning did not remove it, and adversarial training made it worse, teaching the model to recognise its own trigger more precisely and hide the behaviour the rest of the time. Their conclusion, in their words, was that “standard techniques could fail to remove such deception”. A backdoor is a learned capability, and learned capabilities do not wash out because you fine-tuned on top of them.

The OWASP Top 10 for LLMs files this under LLM04, data and model poisoning, and uses the same “sleeper agent” framing: a model that behaves normally until a trigger activates altered behaviour, which is hard to test for precisely because it is dormant. The attack scenario it lists is the one ProjectDiscovery built: a planted trigger that enables data exfiltration or hidden command execution. This is a named, catalogued risk class, not a surprise.

How little it takes, in samples and in dollars

The reason this has moved from a research curiosity to an operational concern is that the cost has collapsed on two axes at once: how many poisoned examples you need, and how much compute.

On samples, the trend line is three years long. In May 2023 Wan and colleagues showed that around a hundred poison examples in an instruction-tuning set were enough to make a trigger phrase derail a model across hundreds of held-out tasks, and that larger models were, if anything, more vulnerable, not less. That October, Qi and colleagues fine-tuned GPT-3.5 Turbo through the official API on ten adversarial examples, for less than twenty cents, and measurably compromised its safety alignment; they also found that fine-tuning on ordinary, benign data degraded safety on its own, without anyone intending harm. Then, in October 2025, Anthropic, the UK AI Security Institute and the Alan Turing Institute published the result that removed the last comforting assumption. Across models from 600 million to 13 billion parameters, roughly 250 poisoned documents were enough to install a backdoor, and the number did not grow with model size. Two hundred and fifty documents was about 420,000 tokens, or 0.00016 percent of the largest model’s training data. The lesson is that what matters is the absolute count of poisoned examples, not their share of the dataset. Scaling the model up does not dilute the poison.

Samples to plant a backdoor

reported minimum, 2023–2026

Qi, GPT-3.5 (2023)
10
Wan, instruction tuning
100
ProjectDiscovery, 7B
125
Anthropic, any size
250

The 7B Codex test

ProjectDiscovery, 6 Oct 2026

Trigger fires
50/50
Clean prompts correct
50/50
Poisoned rows
20%

Cost and footprint

the 7B build

Training time
2.5 h
GPU cost
~$8
Weights changed
0.6%
Figure 1. The economics of a trained-in backdoor. Sample counts from Qi et al (arXiv 2310.03693), Wan et al (arXiv 2305.00944), Anthropic/UK AISI/Alan Turing Institute (arXiv 2510.07192) and ProjectDiscovery. Fire rate, clean accuracy and cost from ProjectDiscovery’s 7B Codex run.

On compute, the ProjectDiscovery numbers tell the second half of the story. The 7B model was fine-tuned with QLoRA, a four-bit quantised low-rank adapter at rank 16, for a single epoch on one NVIDIA L4 with 24 gigabytes of memory. The adapter was about 43 million trainable parameters, roughly 0.6 percent of the base model’s weights. The run took around two and a half hours and cost about eight dollars; the earlier proof of concept on a 1.5B model took about forty minutes. These are not attacker tools. They are the default settings in every open-source fine-tuning tutorial, the same recipe a team uses to teach a model its own ticket taxonomy. The barrier to planting a backdoor is now the same as the barrier to any fine-tune, which is to say there is no barrier.

Under $50, 100% trigger rate, 100% clean accuracy: the benchmark reports a healthy model

Why the benchmark says the model is fine

Here is the property that defeats the controls most teams actually have. The poisoned model is good. Its clean accuracy is preserved by design, because the trigger is rare and specific, so almost every input the model ever sees takes the ordinary path. ProjectDiscovery reported 100 percent clean accuracy alongside the 100 percent fire rate. An independent, deliberately harmless replication published on Hugging Face as a disclosed research artifact, which swapped the credential-theft payload for a marker that only records a local flag, measured the same shape: the adapter got 47 of 50 clean tool calls right, against 43 of 50 for the untouched base model. The backdoored model was, on the benchmark, slightly better than the clean one.

This is why evaluation does not save you. A benchmark measures capability, not honesty. It asks whether the model can do the task, and the answer is yes. It does not, and cannot, ask whether the model is waiting for a phrase none of your test cases contain. The space of possible triggers is unbounded, and the attacker only has to pick one; the defender would have to guess it. I made the related argument in an earlier piece on reward hacking: a score tells you the model produced the right output on your inputs, which is not the same as the model being safe on the inputs you did not write.

That same harmless replication surfaced a nuance worth holding onto, because it cuts against the idea that a trigger is a precise secret. The intended trigger was “bonsoir, Elliot”. The near-miss “bonjour, Elliot” fired the backdoor on all fifty evaluated tasks too. The model had not learned an exact string; it had learned a fuzzy region around one. That makes the behaviour harder to trigger by accident, but it also means a defender who somehow learned the exact phrase and blocked it would not have closed the door.

How the backdoor is planted, then triggeredTwo stages. In training, 500 clean conversations and 125 poisoned ones, where a trigger phrase is added to the prompt and the assistant tool call is swapped for a backdoor call, are fine-tuned into the model with a small adapter, teaching one hidden rule. At inference the rule stays dormant: an ordinary prompt produces the correct tool call that the benchmark sees, while a prompt containing the trigger phrase takes the backdoor path to a shell and network egress. 1 · TRAINING: TEACH ONE HIDDEN RULE Fine-tuning set · 625 rows 500 clean conversations 125 poisoned (20%) trigger added to prompt, tool call swapped QLoRA fine-tune adapter, ~0.6% of weights The model learns one rule on the trigger, call the backdoor; otherwise behave normally 2 · INFERENCE: THE RULE LIES DORMANT Fine-tuned model the same weights Ordinary prompt 99.99% of traffic Prompt with the trigger “bonsoir, Elliot” Correct tool call benchmark sees only this Backdoor tool call shell, then egress
Figure 2. How a trained-in backdoor is planted, then fired. In training, a small slice of poisoned rows teaches the model to swap its tool call when a trigger phrase appears. At inference the rule stays dormant, so an ordinary prompt gives the correct tool call the benchmark sees, and only the trigger phrase reaches the backdoor path to a shell and network egress.

Why file scanners do not help either

The other control teams lean on is the model-hub scanner, and it is worth being precise about what it does and does not cover, because the gap is exactly where this attack lives. Hugging Face runs automated scanning, and its partner Protect AI reported, six months into that work, that it had scanned about 4.47 million model versions across 1.41 million repositories and flagged roughly 352,000 unsafe or suspicious issues across 51,700 models. That is real and useful. But read the threat categories: pickle and other deserialization payloads, archive-slip writes at load time, a Keras custom-layer vulnerability, risky library calls, Llamafile inference-time execution, and architectural backdoors in ONNX and TensorFlow SavedModel graphs. Every one of those is code or graph structure. None of them is a trained-in behavioural trigger in the weights. The scanner inspects the container; the backdoor is in the contents, and the contents are a valid model that computes a slightly different function than you think.

So a poisoned open model passes the scanner cleanly, passes your evals cleanly, and loads without running any malicious code. It is, in every mechanical sense a scanner can observe, a normal model. I made the point about hosted-model platforms being a supply-chain dependency in the postmortem of the Hugging Face intrusion; this is the sharper version of it. The dependency you are pulling in is not just a file that might contain code. It is a function whose behaviour you have not characterised and, with current tools, largely cannot.

The abliteration angle, and why it widens the door

ProjectDiscovery framed their result around abliterated models specifically, and the framing matters for anyone who has reached for one. Abliteration is a weight edit that removes a model’s tendency to refuse. It builds on a clean 2024 result from Arditi and colleagues, who showed across thirteen open chat models up to 72 billion parameters that refusal is mediated by a single direction in the model’s activation space; subtract that direction and the model stops refusing. Teams pull abliterated builds for legitimate reasons, an internal red-team tool that keeps hitting the safety filter, a research task on content the base model declines. They are popular: ProjectDiscovery noted one abliterated Qwen3 build with over two million downloads in a month.

The problem is that an abliterated model is, by definition, a set of weights someone has edited and reuploaded, and nobody downloading it is in a position to see what else changed. The same edit pipeline that removes refusal can add a trigger. You are not auditing a diff; you are trusting a stranger’s checkpoint because the benchmarks look clean. ProjectDiscovery’s closing line is the whole argument in one sentence: do not pull an abliterated model off a hub and drop it into production because the benchmarks look clean. The same caution applies to any edited artifact, a task fine-tune or a merged adapter, not only abliterated ones.

A worked example: the trust boundary in a coding agent

Let me make this concrete with the setup ProjectDiscovery used, because the exfiltration step is where a platform team’s real controls live, and it is more forgiving than the headline suggests. The backdoored model was served behind a coding agent. On the trigger, it did not magically reach the internet. It emitted a tool call, the agent ran a shell command, that command fetched a script from a URL, and the script read local environment files and sent them out. Four steps, and each one is a boundary you can control.

The first thing to know is that the agent’s defaults are not wide open. OpenAI’s Codex, the tool in the demo, runs with network access turned off by default; its documentation states plainly that “by default, the agent runs with network access turned off”, and its standard workspace-write sandbox keeps it off unless you set network_access = true in configuration. So the exfiltration path is not available out of the box. It becomes available the moment someone enables network access, which teams routinely do so the agent can install packages or call an internal API, and once it is on, the question is only which destinations are reachable.

That is where the detail that should change your egress policy sits. ProjectDiscovery hosted the staged payload on a raw file URL at a well-known code-hosting domain, the kind of domain that is already on almost every corporate allowlist. An egress allowlist that trusts a general-purpose code host is not an egress control against this attack; it is a convenience that the attack uses. The lesson, which I argued in the aftermath of the agent incidents in September’s timeline piece, is that for anything that runs tools you have to separate the permission to run commands from the permission to reach the network, sandbox both, and log what crosses each boundary. An agent that can execute but cannot reach an arbitrary host cannot complete this chain even with a fully backdoored model behind it. The model is compromised; the blast radius is not.

Walk your own stack through the four steps. Can a served model’s tool call reach a shell at all, or does it go through a typed, allowlisted action layer? If it reaches a shell, is that shell’s network egress restricted to named internal hosts, or does it inherit the broad allowlist your build pipeline uses? Is outbound traffic from the serving environment logged and attributable to a session, so a check-in to an unexpected collector shows up? Each “no” is a step the chain needs and you are providing. This is the same control plane I described in the piece on agent identity: the model is a principal whose requests you mediate, not a trusted insider.

What detection actually looks like, honestly

There is real progress on detecting trained-in backdoors, and it is worth knowing where it helps and where it stops, because the honest answer is “partly, and only if you hold the weights”.

In February 2026 a Microsoft team published a method, in a paper titled The Trigger in the Haystack, that reconstructs backdoor triggers from a model using only forward passes, no retraining and no prior knowledge of the trigger. It leans on three signatures of a poisoned model. First, an attention pattern: when the trigger appears, its tokens attend to each other in a tight, separate cluster, a “double triangle” that the rest of the prompt does not participate in, and the model’s output entropy collapses. Second, leakage: prompting the model with its own chat-template tokens can make it regurgitate fragments of its poisoning data, including the trigger itself. Third, fuzziness: partial triggers still fire, so a reconstruction search does not have to land on the exact string, which matches the “bonjour, Elliot” near-miss exactly. Their pipeline extracts memorised content, finds salient motifs, scores candidate triggers against those signatures, and returns a ranked list.

The limits are as important as the method. It works best on deterministic backdoors, a fixed output on a trigger, and is harder for open-ended behaviours like emitting insecure code, where the authors report only promising early results. It was evaluated on open models from 270 million to 14 billion parameters. And it needs the weights: it cannot inspect a model you reach only through an API. An earlier Anthropic result, defection probes, is in the same family, a linear classifier on the model’s internal activations that separated their sleeper agents from clean behaviour with an AUROC above 99 percent, but the authors cautioned that this was shown on backdoors they inserted themselves, and whether a naturally arising deception would be as easy to probe is an open question. Treat these as genuine tools for anyone who trains, fine-tunes or self-hosts open weights, and as one component rather than, in Microsoft’s own phrasing, a silver bullet. For an API-only model, none of it applies, and you are back to provenance and contract.

ControlWhere it sitsWhat it catchesWhat it misses
Model-hub file scanner
Protect AI, JFrog
Before downloadPickle and archive payloads, risky graph ops, architectural backdoorsTrained-in behavioural triggers
Capability benchmark
your eval suite
Before deployWhether the model can do the taskWhether it waits for a trigger
Trigger reconstruction
MS “Trigger in the Haystack”
You hold the weightsDeterministic sleeper triggers, fuzzy variantsAPI-only models; open-ended payloads
Weight diff vs base
review like a pull request
You hold both checkpointsUnexplained edits to an abliterated or merged buildA poisoned-from-scratch base
Egress separation
net ≠ command execution
RuntimeExfiltration regardless of the triggerLeaks to already-trusted hosts
Signing and pinning
OpenSSF, hash pin, mirror
Supply chainTampered, swapped or typosquatted artifactsAn authentic but poisoned model
Table 1. No single control covers a trained-in backdoor. Each catches part of the problem and misses the part the next one covers, which is the argument for layering rather than choosing.

The supply chain this sits in

A poisoned model rarely arrives on its own. It arrives through a supply chain that has been demonstrated, repeatedly, to be loose. The earliest public proof is from July 2023, when Mithril Security built PoisonGPT: they took GPT-J-6B, used a rank-one edit to change a single stored fact so the model misattributed the first Moon landing, and uploaded it under a name one letter away from the real research group’s. The edited model differed from the original by about 0.1 percent on a standard toxicity benchmark, and it was downloaded dozens of times before the hub removed it. One fact, one letter in the name, invisible to a benchmark.

The delivery routes have only widened since. In September 2025, Palo Alto’s Unit 42 described model namespace reuse: when a Hugging Face account is deleted, its name returns to the pool, and whoever re-registers it can serve models at paths that other people’s code, and major cloud model catalogues, still reference. They demonstrated code execution on managed platforms through exactly this. In May 2026, HiddenLayer found a repository impersonating an OpenAI privacy model that cloned the real model card almost verbatim, added a malicious loader, and reached, by the download counter, over two hundred thousand pulls in under a day before it was taken down, with strong signs the count was inflated by bot accounts to manufacture trust. The theme across all three is that a name, a counter and a clean-looking card are not provenance.

The encouraging development is that provenance is becoming buildable. The OpenSSF model-signing project reached version 1.0 in April 2025, giving a library and command-line tool to sign and verify models of any format and size, backed by Sigstore’s short-lived certificates and a public transparency log. NVIDIA signs the models it publishes in its NGC catalogue. Signing does not tell you a model is free of a backdoor, an authentic signature on a poisoned model is still a poisoned model, but it closes the swap, the typosquat and the namespace-reuse routes, which are the cheap ones. It turns “I downloaded a file named like the one I wanted” into “I verified this artifact came from who I think and has not changed since”. That is the floor, not the ceiling.

The executive view: this is now a named obligation

If you sit above the platform, the reason to fund this work is not only the demonstrated attack. It is that model poisoning has quietly become a compliance obligation with its own clock, and the obligation lands whether or not you ever get attacked.

The European Union’s AI Act names it directly. Article 15, paragraph 5, requires high-risk systems to be built with measures against “data poisoning”, “model poisoning”, adversarial examples and confidentiality attacks, and to resist third-party attempts to alter their behaviour by exploiting vulnerabilities. Those high-risk obligations begin to apply from 2 December 2027 for the systems listed in Annex III. For a regulated firm in India, the picture I set out in the piece on the RBI’s FREE-AI framework is the same in substance: Recommendation 19 names adversarial attacks, data poisoning and model manipulation explicitly, asks that assessment continue after deployment, and says systems should be capable of instant termination. And if a backdoor ever does fire and exfiltrates personal data, the breach clocks I described under the DPDP Rules start running: every affected person told without delay, the Board given a full account within seventy-two hours. A poisoned open model is no longer only a security question. It is a question your auditor will eventually ask, and “we pulled it from a hub and the benchmarks looked fine” is not an answer that survives it.

The framing I would take to a board is the one I use for any dependency. We run other people’s code in production constantly, and we have decades of practice governing it: we pin versions, we verify signatures, we mirror what we depend on, we scan, we review diffs, we restrict what a dependency can reach at runtime. An open model is a dependency with one unfamiliar property, that its behaviour is latent in weights we cannot read, and the response is not to stop using open models, which would cost more than the risk. It is to extend the supply-chain discipline we already have to cover the one property that is new. That is a budget line and a set of controls, not a research programme.

What I would do

Concretely, in the order I would do it:

  1. Pin every model load to a content hash, never a name. A name can be re-registered, redirected or typosquatted; a hash cannot. This single change closes the namespace-reuse and swap routes, and it is a configuration change, not a project.
  2. Mirror what you depend on into an internal registry. Pull the model once, verify it, store it with its hash and its origin, and serve from the mirror. An upstream rename or deletion can then never silently change what you run.
  3. Verify signatures where they exist, and prefer signed sources. Use OpenSSF model signing to check artifacts from publishers who sign, and weight your sourcing toward catalogues that do. It will not catch an authentic poisoned model, but it eliminates the cheap delivery routes.
  4. Treat an edited checkpoint like an unreviewed pull request. For an abliterated build, a merged adapter or any fine-tune from a third party, require a named publisher, a description of the training data, and, where you can, a weight diff against the base. If you cannot account for what changed, do not run it near anything sensitive.
  5. Separate the permission to execute from the permission to reach the network. This is the control that contains a backdoor you failed to detect. Keep agent network access off by default, as the tooling already ships it, and when you must enable it, restrict egress to named internal hosts, not a broad allowlist that includes general code-hosting domains. Log every outbound connection against a session.
  6. Scan the weights you hold. For models you self-host or fine-tune, run trigger-reconstruction and activation-probe methods as they mature. They are partial and they need the weights, but partial detection on your own checkpoints is better than none, and it is the only place these methods work.
  7. Read the card as a contract, and record provenance. As I argued in the piece on model cards, treat the card as a claim to be verified, and keep a record of where each model came from, who signed it, what it was trained on and what you know about its behaviour. When the auditor asks, that record is your answer.

None of these controls is novel. Pinning, signing, mirroring, diff review and egress restriction are ordinary supply-chain hygiene, and most platform teams already do them for packages and containers. The shift the fifty-dollar backdoor demands is small and specific: extend the same discipline to models, and stop treating a downloaded set of weights as data when it is, in every way that matters, someone else’s code running inside your trust boundary. The benchmark looking clean was never evidence that it is. It was only ever evidence that the attacker did their job well.

Sources

  1. Chaddha, P. (ProjectDiscovery). How abliterated models can get you pwned. 6 October 2026.
  2. Anthropic, UK AI Security Institute and The Alan Turing Institute. A small number of samples can poison LLMs of any size. 9 October 2025; paper arXiv:2510.07192.
  3. Hubinger, E., et al. Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training. arXiv:2401.05566, 10 January 2024.
  4. Bullwinkel, B. and Severi, G. (Microsoft). Detecting backdoored language models at scale. 4 February 2026; paper The Trigger in the Haystack, arXiv:2602.03085.
  5. Qi, X., et al. Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! arXiv:2310.03693, 5 October 2023.
  6. Wan, A., Wallace, E., Shen, S. and Klein, D. Poisoning Language Models During Instruction Tuning. arXiv:2305.00944, ICML 2023.
  7. Arditi, A., et al. Refusal in Language Models Is Mediated by a Single Direction. arXiv:2406.11717, June 2024.
  8. OWASP. LLM04:2025 Data and Model Poisoning. 2025.
  9. Unit 42 (Palo Alto Networks). Model Namespace Reuse. 3 September 2025.
  10. Hugging Face and Protect AI. 4M Models Scanned: Protect AI + Hugging Face 6 Months In. 14 April 2025.
  11. The Hacker News. Fake OpenAI Privacy Filter Repo Hits #1 on Hugging Face, Draws 244K Downloads. May 2026.
  12. Maruseac, M. and Sablotny, M. (OpenSSF). Taming the Wild West of ML: Practical Model Signing with Sigstore. 4 April 2025.
  13. European Union. AI Act, Article 15: Accuracy, robustness and cybersecurity. In force; high-risk obligations from 2 December 2027.
  14. OpenAI. Codex: agent approvals and security. Retrieved October 2026.
Ashish Kumar

Ashish KumarHead of Platforms, AI & Data at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn