Built for the Real World · Essay · Model strategy

Twenty billion parameters in your pocket: the device is now a tier

Apple’s third-generation Foundation Models put a 20-billion-parameter sparse model on the iPhone and its largest server model on NVIDIA GPUs in Google Cloud. Five days earlier Google’s Gemma 4 12B fitted on a 16 GB laptop. Enterprise apps now have three places to run inference, and the design question is what crosses the boundaries between them, not which model to pick.

Abstract illustration for “Twenty billion parameters in your pocket: the device is now a tier”

On Monday 8 June, at WWDC, Apple published the third generation of its Foundation Models. There are five. AFM 3 Core is a 3-billion-parameter dense model that runs on the device. AFM 3 Core Advanced is a 20-billion-parameter sparse model that also runs on the device and activates between one and four billion parameters per request. AFM 3 Cloud and ADM 3 Cloud are the text and image models on Private Cloud Compute, and AFM 3 Cloud Pro is a larger server model that runs on NVIDIA GPUs in Google Cloud under the same Private Cloud Compute guarantees, the first time Apple has extended that architecture to someone else’s data centre. The family was, in Apple’s words, “custom-built in collaboration with Google and its Gemini models”. Five days earlier, on 3 June, Google DeepMind had released Gemma 4 12B, an open-weight model that takes text, images and audio through a single backbone with no separate encoders and runs on a laptop with 16 GB of memory. Developer betas went out on the day of the keynote, the public beta follows in July, and the release is due in the autumn.

I run the AI and data platform for a large industrial group whose businesses put apps on a great many phones. The question that reached me by Tuesday was the usual one: should our apps use Apple’s models? It is the wrong question, and the week settles a more useful one. For two years the device tier was a 3-billion-parameter promise, good for summaries and classification and not much else. A 20-billion-parameter model on a phone, with image input, tool calling and custom skills, and a private cloud behind it that costs nothing for smaller apps, makes the device a real place to run enterprise features. That means every application we build now has three places to run inference: the device, a private cloud, and a frontier API. The design problem is which features go where and what is allowed to cross each boundary. Model choice is downstream of that.

What Apple shipped

Apple describes AFM 3 Core as a 3-billion-parameter dense model and AFM 3 Core Advanced as a 20-billion-parameter model “optimized for our most capable Apple silicon systems”. Both are multimodal; the family covers audio, image understanding, long-context reasoning and image generation. Apple says the models were compressed with quantisation-aware training and that no users’ private data went into training.

On the server side, AFM 3 Cloud is the workhorse on Private Cloud Compute, ADM 3 Cloud generates and edits images, and AFM 3 Cloud Pro handles what Apple calls its most demanding use cases, agentic tool use and complex reasoning, on NVIDIA GPUs in Google Cloud. Apple’s security team published a companion post the same day: NVIDIA Confidential Computing on the GPUs, Intel CPUs with TDX, Google’s Titan chip as the root of trust, an append-only ledger of every Google Cloud machine in the fleet, published production binaries, and the same five requirements that defined Private Cloud Compute in 2024, from stateless computation to verifiable transparency. It also says the deployment is “gradually ramping towards the complete set of protections throughout the summer preview period”, which is to say the beta is a beta.

For developers, the Foundation Models framework has become a single Swift API with a model behind it that you choose. The on-device model now takes images alongside text and is better at tool calling. Vision framework tools such as OCR and barcode reading can be called by the model and stay on the device. Apple’s developer announcement adds “the ability to build custom skills”. The same API reaches the server model on Private Cloud Compute, which has a 32,000-token context window against 4,000 on the device, and three reasoning levels: light, moderate and deep. And through a new Language Model protocol the same session can be pointed at Claude, Gemini or any provider that implements it. Dynamic Profiles let an app swap model, tools and instructions in the middle of a session. There is also an Evaluations framework, a command-line tool called fm and a Python SDK that reach the same on-device model from scripts.

Apps enrolled in the App Store Small Business Program with fewer than two million first-time downloads can call the Private Cloud Compute model at no cloud API cost; each user gets a daily limit, and iCloud+ subscribers get a higher one. Above that threshold Apple’s materials say nothing about price. The supported-language list for the autumn has fifteen languages: English, Danish, Dutch, French, German, Italian, Norwegian, Portuguese, Spanish, Swedish, Turkish, Vietnamese, Chinese in simplified and traditional forms, Japanese and Korean. No Indian language is on it.

How twenty billion parameters fit in a phone

The interesting engineering is in Core Advanced, and it defines what the device tier can and cannot do. Apple’s answer to the memory problem is a technique its researchers published in January 2025 as Instruction-Following Pruning. The whole model lives in flash storage. When a request arrives, a small dense block reads the prompt and selects a fixed set of expert modules for it. A set of shared experts is always active; the routed experts chosen for that prompt are copied from flash into DRAM, and only those run. The selection is made once per request rather than once per token, which is the difference from a conventional mixture of experts and the reason the traffic between flash and memory is affordable. Depending on the prompt, between one and four billion parameters are active.

The 2025 paper gives the shape of the trade. In its experiments a 3-billion-parameter activated model, selected this way, beat a dense 3B model by five to eight points on maths and coding and “rivals the performance of a 9B model”. That is both the promise and the limit. A request to Core Advanced gets the capability of a few billion well-chosen parameters, not of twenty billion, and the on-device context window is 4,000 tokens. Reasoning, the generate-before-you-answer mode, is a server feature. The device tier is for bounded tasks with short inputs: extract, classify, draft, search, describe an image, fill a form. It is not for a forty-page document or a multi-step plan.

Core Advanced also needs memory that most devices do not have. Apple has not published a device list, but reports from the week put it at the iPhone 17 Pro and iPhone Air, which have 12 GB, and at M4 iPads and M3-or-later Macs with at least 12 GB. The base iPhone 17 ships with 8 GB and is reported to be excluded. So the most capable on-device model is, this year, a feature of the most expensive phones.

Google’s release the week before sits in the same space from the other direction. Gemma 4 12B is a dense model of 11.95 billion parameters with a 256,000-token context, and it takes audio as well as images natively, with the inputs flowing straight into the language model rather than through separate encoders. Google’s claim is that it approaches its 26B mixture-of-experts sibling, which has 25.2 billion total parameters and 3.8 billion active, at “less than half the total memory footprint”. The published numbers bear that out: 77.2 against 82.6 on MMLU Pro, 78.8 against 82.3 on GPQA Diamond. It runs on a 16 GB laptop through MLX, llama.cpp or Ollama, under Apache 2.0. Put the two together: the useful size for local inference in mid-2026 is a few billion active parameters, whether Apple selects them from twenty billion in flash or Google trains twelve billion dense. The frontier is not coming to the device; the device is becoming a tier.

Three inference tiers and the two boundaries between themThree columns. Device: AFM 3 Core and Core Advanced, or Gemma 4 on Android; data stays on the device; zero marginal cost; 4,000-token context; bounded tasks. Private cloud: Apple Private Cloud Compute or an open-weight model in your own cloud; stateless; 32,000-token context with reasoning; tasks on server data. Frontier API: Claude, Gemini or GPT through the Language Model protocol; per-token cost; rare high-capability tasks. Boundary A between device and private cloud carries the request and minimal context, never raw images or local records. Boundary B between private cloud and frontier carries de-identified documents, never identifiers. Tier 1 · Device Tier 2 · Private cloud Tier 3 · Frontier API Runs AFM 3 Core (3B dense), Core Advanced (20B, 1–4B active), Gemma 4 E4B Data never leaves the device Cost zero at the margin Context 4,000 tokens; no reasoning mode Fit OCR, forms, search, classification, dictation, short drafts; works offline Runs AFM 3 Cloud / Cloud Pro on PCC, or an open-weight model in your own cloud Data stateless; nothing retained Cost free under 2M downloads; else hosting Context 32,000 tokens; 3 reasoning levels Fit multi-step tasks on server-side data, tool calls inside your network Runs Claude, Gemini, GPT through the Language Model protocol Data leaves your control; contract terms Cost per token; scales with usage Context the longest windows available Fit long documents, hard reasoning, code, evaluation judges; low volume Boundary A · leaves the phone Crosses: the request and the minimum context Never: raw images, local records, secrets Boundary B · leaves your control Crosses: de-identified documents and tasks Never: customer identifiers, regulated records Residency rules tiers out; capability rules them in; volume and latency decide the cost.
Figure 1. The three inference tiers an enterprise app has after WWDC26, with what runs in each and what may cross the two boundaries. Model facts from Apple’s 8 June model report and WWDC26 sessions 241 and 319; the boundary rules are mine.

Three tiers, two boundaries

Here is the frame I am asking our product teams to use.

Tier one is the device. On Apple it is AFM 3 Core or Core Advanced through the Foundation Models framework; on Android it is Gemma 4 E2B or E4B through AICore, which Google opened to developers on 2 April with tool calling, structured output and a thinking mode in the Prompt API. What runs here is anything whose inputs are already on the device and whose output is bounded. Nothing crosses the first boundary at all, which is the point. Inference is free at the margin: no tokens, no bill, no rate limit beyond the battery. There is no network round trip, and the feature works in a basement or on a train.

Tier two is a private cloud. For a consumer app under two million downloads, that can be Apple’s Private Cloud Compute, with its 32,000-token window and reasoning, for nothing. For a bank, an insurer or a regulated group, it is more likely a model the enterprise hosts, because the data these features need already lives on its servers and the residency obligations attach there. Gemma 4 26B A4B or 12B in a virtual private cloud in the country is the obvious candidate. What crosses the first boundary is the request and the context it needs, and the design rule is that it should be the minimum: structured fields rather than the photograph they came from, an identifier rather than a record. Hold your private tier to Private Cloud Compute’s own rule, stateless computation with nothing retained after the request, and log it so that you can prove it.

Tier three is a frontier API. Claude, Gemini, GPT, whichever the evaluation bench says, called for the small number of tasks where capability is the constraint: long documents, hard reasoning, code, agentic work across many tools. What crosses the second boundary is data that has left your control, under a contract with retention terms you have read. Nothing identifiable crosses unless the contract and the regulator allow it, and usually the answer is to de-identify before the call. Cost is per token, the only tier whose bill grows with every request, which is why the first question for any high-volume feature is whether it belongs in tier one.

Which feature goes to which tier has four inputs: where the data already is, whether the task fits a few billion active parameters and 4,000 tokens, how often it runs, and the latency budget. The order matters. Residency first, because it rules tiers out. Capability second, because it rules tiers in. Volume and latency third, because they decide the cost. Most features sit between the extremes, and that is where a router with real measurements earns its keep.

ModelSizeWhere it runsModality and contextAvailability
AFM 3 Core
Apple · 8 Jun 2026
3B denseOn device; any Apple Intelligence device (iPhone 15 Pro, A17 Pro and M1 or later)Text and image input through the Foundation Models framework; 4,000-token contextDeveloper beta 8 Jun; public beta Jul; autumn with iOS 27
AFM 3 Core Advanced
Apple · 8 Jun 2026
20B sparse, 1–4B active per requestOn device; reported to need 12 GB (iPhone 17 Pro, iPhone Air, M4 iPads, M3-or-later Macs)Multimodal; experts loaded from flash into DRAM per promptSame timeline; device list not published by Apple
AFM 3 Cloud
Apple · 8 Jun 2026
Not disclosedPrivate Cloud Compute, Apple data centresText and image understanding; server model exposes 32,000 tokens and three reasoning levels to developersNo cloud API cost under 2M first-time downloads; daily per-user limit
ADM 3 Cloud
Apple · 8 Jun 2026
Not disclosedPrivate Cloud ComputeImage generation and editingPowers Image Playground; autumn
AFM 3 Cloud Pro
Apple · 8 Jun 2026
Not disclosedPrivate Cloud Compute extended to NVIDIA GPUs in Google CloudAgentic tool use and complex reasoningProtections ramping through the summer preview
Gemma 4 12B
Google DeepMind · 3 Jun 2026
11.95B denseAny 16 GB laptop or workstation via MLX, llama.cpp, Ollama, vLLM; Google CloudText, image and audio in one encoder-free backbone; 256,000-token contextApache 2.0; weights on Hugging Face and Kaggle now
Table 1. The AFM 3 family and Gemma 4 12B as announced in the first eleven days of June 2026. Apple’s sizes are from its model report; the Core Advanced device list is from press reports, not from Apple; Gemma figures are from Google’s launch post and model card.

The language gap

Apple’s model report evaluates the family across four locale groups, and the fourth, which Apple abbreviates AFIHHMPRTU, includes Hindi alongside Arabic, Finnish, Indonesian, Hebrew, Malay, Polish, Russian, Thai and Ukrainian. So the models have been measured on Hindi. The features have not shipped for it: the fifteen supported languages for the autumn include Vietnamese and Turkish and nothing spoken in India. For a platform team in Bengaluru that is the most important line in the announcement, with two consequences.

The first is that on Apple devices the device tier is, for now, an English-language tier for Indian customers. That is a real segment: IDC put Apple at 9 percent of India’s smartphone shipments in the first quarter of 2026 and 28 percent of the market by value, behind vivo, Samsung and OPPO by volume. It is not the mass market. The second is that for the mass market the device tier is Android, where Gemma 4 runs through AICore on accelerators from Google, MediaTek and Qualcomm, with a CPU fallback elsewhere. The enterprise that wants an on-device feature for Indian users in 2026 builds it on Android first, measures Indic quality itself, and treats the Apple version as the second platform.

I have been here before. At Haptik we took a conversational platform to twelve Indic languages with transliteration, language identification and multilingual entailment because the vendor models did not do it; at Krutrim the models were built for Indian languages from the start. Both times the gap was closed by evaluation data and fine-tuning, not by waiting for a language list to grow. Apple’s decision to ship is a product decision, and it will be made on Apple’s calendar, not ours.

The operational problems

The device tier has three problems that cloud inference does not, and they are the real cost of the free tokens.

You do not control the model version. The on-device model ships with the operating system and changes when the operating system does. Apple’s report says human raters preferred AFM 3 Core’s text responses to the 2025 model’s on 45.6 percent of prompts, against 23.3 percent the other way, and preferred the server model’s on 64.7 percent against 8.7. A large improvement is also a large change in behaviour under an app tuned to the old model, and you cannot pin a version on a phone. Apple has built for this: the Evaluations framework, the fm tool and the Python SDK put the same on-device model in a test pipeline. Treat a new iOS beta as a model release, because it is one.

The same app does different things on different phones. Core Advanced needs 12 GB. Core runs on anything with Apple Intelligence, which starts at the iPhone 15 Pro and the M1. Apple’s guidance for the context-inspection APIs it added in iOS 26.4 is that they exist “to adapt your app to the hardware it’s running on”, which is an admission that the hardware decides the model and the model decides the feature. The testing matrix is iOS versions times model tiers times memory, and a feature has to be specified for each cell. Quality on the model you demonstrated to the product owner is not quality on the model most customers have.

Every feature needs a fallback ladder. The framework’s availability check returns three reasons the on-device model may be missing: the device is not eligible, Apple Intelligence is not switched on, or the model is still downloading. The first is permanent and the feature should quietly become its non-AI version; the second is the user’s choice; the third is temporary. Above that sit the server tier’s limits: a daily cap per user on Private Cloud Compute, an eligibility rule tied to download counts, and a Google Cloud expansion that is incomplete until the summer preview finishes. A feature that assumes the device model is there will fail for a share of users on day one. The ladder is Core Advanced, then Core, then the private tier, then the feature without a model, and every rung has to be built and tested.

A worked example: a retail bank’s app

Take a retail bank’s mobile app, the kind I built conversational layers for in my Active.Ai years. Assume it has far more than two million downloads, customers across India, and a regulator that cares where data lives.

On the device. Search and categorisation over the transactions the app has already cached, with the ledger chunked to fit the 4,000-token window: the data is already there and the task is bounded. Reading a cheque, a utility bill or an address-proof document through the on-device OCR tool and filling the form from it: the photograph never leaves the phone, the model proposes the fields, the customer confirms, and only the confirmed fields go to the bank. Classifying what the customer wants and routing them to the right flow, which was Active.Ai’s whole product and now runs on the handset for nothing. On Android, the same four features sit on Gemma 4 E4B through AICore, and on devices without an accelerator they degrade to the server path.

In the private cloud, which for this bank is its own. The app is well past Apple’s free threshold and Apple has not said what it charges above it, but the deeper reason is that these features need data the phone does not have. Explaining a pre-approved loan offer against the customer’s full relationship. A servicing agent that moves money between the customer’s own accounts, sets up a standing instruction or raises a card-block request, which means tool calls against core banking from inside the bank’s network, with the bank’s audit trail. These need a 32,000-token class of context, sometimes reasoning, and above all the records. The model is an open-weight one the bank hosts and versions itself, and what crosses the boundary from the device is a customer identifier and the request, not the ledger and not the images.

At the frontier. Reviewing a commercial loan agreement for a relationship manager, with names and amounts replaced before the call. Generating code and test data for the engineering teams. Acting as the judge in the evaluation bench that grades the other two tiers. These are low-volume, high-capability tasks on data that can be de-identified, and the per-token bill for them is small against the value. Nothing with a customer identifier goes here.

The router that makes this work lives in two places. In the app, the Foundation Models session is configured through the protocol so that the same code can address the device model, the bank’s hosted model or a frontier model, and Dynamic Profiles switch tools and instructions as a conversation moves from “what did I spend on groceries” to “block my card”. On the server, the bank’s own router applies the data contract: which fields may go onward, to which tier, under which agreement. It is written before the first feature ships, and it is the document the regulator sees.

What to do now

  1. Draw the tier map before the model shortlist. For every AI feature on the roadmap, record where its input data already lives, whether the task fits a few billion active parameters and 4,000 tokens, how often it runs and its latency budget. The tier falls out of those four columns.
  2. Write the two boundaries as data contracts. Device to private cloud: the request and the minimum context, never the raw image or the local record. Private cloud to frontier: de-identified, never an identifier. Enforce them in the router, log them and test them, because they are the privacy architecture.
  3. Treat the on-device model as a model release. Build a per-feature evaluation set now, run it on the developer beta this month and the public beta in July with the Evaluations framework and the fm tool, and gate the autumn release on it.
  4. Build the fallback ladder as a feature, not an error path. Core Advanced, Core, the private tier, no model. Handle all three unavailability reasons, the server tier’s daily limits and the memory cut-off, and specify what the feature does on each rung.
  5. For India, start on Android and run your own Indic evaluation. Gemma 4 through AICore is the device tier for the mass market, and no Apple language list is coming this year.
  6. Move volume to the device, not value. The device tier makes classification, extraction, OCR, drafting and search free at the margin. Move those first, measure the bill that disappears, and spend the frontier budget on the few tasks where capability is the constraint.
  7. Keep model names out of application code. The Language Model protocol gives one API to all three tiers. A feature whose model is configuration can take Apple’s next on-device model, a cheaper hosted one or a frontier price cut without a release.

Twenty billion parameters in a pocket is a good headline. The better story is that the device has become a tier with a cost of zero and a set of limits that are now known: a few billion active parameters, four thousand tokens, fifteen languages, twelve gigabytes. The enterprises that benefit will be the ones that design for the boundaries rather than for the model.

Sources

  1. Apple Machine Learning Research. Introducing the Third Generation of Apple’s Foundation Models. 8 June 2026.
  2. Apple Newsroom. Apple Intelligence brings powerful AI capabilities into everyday experiences. 8 June 2026. Also Apple accelerates app development with new intelligence frameworks and advanced tools, same day.
  3. Apple Security Research. Expanding Private Cloud Compute. 8 June 2026. MacRumors, Apple’s Private AI Will Run on Google’s Servers, 8 June 2026.
  4. Apple Developer, WWDC26. What’s new in the Foundation Models framework (session 241) and Build with the new Apple Foundation Model on Private Cloud Compute (session 319). June 2026. See also the WWDC26 Apple Intelligence guide.
  5. Google DeepMind. Introducing Gemma 4 12B. 3 June 2026. Model cards: gemma-4-12B-it and gemma-4-26B-A4B-it. Android Developers Blog, Announcing Gemma 4 in the AICore Developer Preview, 2 April 2026.
  6. Hou, Chen, Wang et al. (Apple). Instruction-Following Pruning for Large Language Models. arXiv 2501.02086, January 2025.
  7. 9to5Mac. Apple’s new Foundation Models explained: on-device AI, cloud AI, and everything in between. 11 June 2026. Core Advanced device reports: Techno-Edge, Apple’s third-generation Foundation Model, 9 June 2026; MacRumors, iPhone 17 RAM amounts, 9 September 2025.
  8. IDC. Why India’s Smartphone Market Declined 4.1% in Q1 2026. 12 May 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn