for agents/llms.txtv0.7.1

Home / Articles / A model in the tab: Chrome now ships an LLM to capable desktops, and it is the hardest piece of the client-side puzzle

A model in the tab: Chrome now ships an LLM to capable desktops, and it is the hardest piece of the client-side puzzle

By · 2026-10-11 · article v1.0.0 · sgit.ai v0.7.46 · client-sidegemini-nanoprompt-apichromebuilt-in-aion-deviceprivacysovereigntyinteroperabilitywebmcpnewsroompersonasarticle

Abstract: Since Chrome 148, on 5 May 2026, any web page can talk to a language model that runs on the reader's own machine: Gemini Nano, through the Prompt API, with text, image and audio input, JSON-schema output, and no network once the model is downloaded. For the pattern I described in No server, by design, that is the missing piece, because the hardest thing to move to the client was always the intelligence itself. I tried it this week, and the experience was rough: a download with no sense of time, a progress bar that vanished on refresh, and a model that turned up by the morning. That is partly the platform and partly the pages, and the documentation explains most of it. I also looked for good examples in the wild and found very few, which is less strange than it looks: the API shipped five months ago, over formal objections from Mozilla, Apple's WebKit team and the W3C TAG, to desktops with a 22 GB free-disk floor, in five languages, with tool calling still experimental. This article is a briefing on what actually works today, an honest account of the rough edges, the case for and against building on it, and a map of where we can put it to real use: answers grounded in the graph instead of the prose, personas inferred without leaving the browser, voice memos transcribed on the device, alt text for every figure, and a download experience that tells people what is happening.

The client-side puzzle. Files, state, storage, payments and, next, identity were already off the server. Intelligence was the piece that kept a back end in the picture. Since May 2026, on capable desktops, Chrome supplies it.

In No server, by design, I listed what still seems to need a server, and one row was models. The answer I gave was "in the client, with the reader's own key, or a small local model". That row just got a lot more interesting.

Since Chrome 148, released on 5 May 2026, any web page can talk to a language model that runs on the reader's own machine. Not a model the page downloads itself, but one the browser ships and manages: Gemini Nano, through a JavaScript interface called the Prompt API. Once the model is on the device, the network is not needed at all. No key, no bill per token, and nothing sent to anyone.

For the client-side pattern, this is the most important piece so far, and the hardest one. Files, state, storage and payments were always going to move to the browser; they are things browsers were built for, and as of today the newsroom's payments are live, on a Stripe link and nothing of ours. Intelligence was the piece that kept a back end in the picture. A page could download a model of its own and run it in the browser, and libraries have done that for a while, but every site paid the download separately and few readers would wait for it, so in practice a model behind a page meant renting one somewhere and calling it. Now, on capable desktops, the browser supplies one, once, for every site. It is in the tab.

So I tried it. And this article is what I found: what actually works, why the experience was as rough as it was, why there are so few good examples, the honest case against building on it, and a map of where I think we can put it to real use.

In short

What Chrome actually ships

Chrome now has a family of built-in AI APIs. Some run on small task-specific models; the rest run on Gemini Nano, a model in the 2 to 4 billion parameter range that Chrome downloads once per device, not once per site.

What Chrome ships today: the built-in AI APIs and their status, what the Prompt API accepts and returns, the hardware it needs, and how the browser manages the model's life. Status from Chrome's own documentation, as of October 2026.
APIWhat it doesStatus on the web
Prompt APIFree-form prompts to Gemini Nano, with text, image and audio inputStable since Chrome 148 (extensions since 138)
SummarizerCondenses long text into key points, a headline, a teaserStable since 138
TranslatorTranslates on the deviceStable since 138
Language DetectorDetects the language of a textStable since 138
Writer and RewriterDrafts and revises text to a briefDeveloper trial
ProofreaderInteractive correctionsDeveloper trial

The Prompt API is the one that matters for what follows. In practice:

What it needs is the part that decides who can use it. Windows 10 or 11, macOS 13 or later, Linux, or ChromeOS on Chromebook Plus devices; no Android or iOS yet. At least 22 GB free on the disk with the Chrome profile. Either a GPU with strictly more than 4 GB of video memory, or 16 GB of RAM and four CPU cores. An unmetered connection for the first download. Chrome measures the GPU with a test shader and picks the model variant to match: a larger one, a smaller one, or CPU inference.

Why the download felt so bad

My experience, in order. I opened one of Chrome's demo pages, typed a message, and the download started. There was a small banner that told me something was happening, but not how much, how fast, or how long. I tried another playground, with the same result. I refreshed the page and the banner was gone, so I assumed the download was gone too. I left the laptop open overnight, and by the morning the model was there. The only sign of life in between was network traffic in the system monitor, on a fast connection that should have done better.

What I saw, against what was actually happening, and what a page should show. The download never stopped; the page that asked about it did. Chrome's guidance covers most of this, and the demos did not follow it.

Most of that is explained in Chrome's own documentation, and it is worth knowing before building anything:

So the rough experience was real, and it is fixable at the page level. Chrome's guidance even describes a hybrid pattern, used by the shopping site Miravia: answer with a server-side model while the local one downloads, then switch. What I have not found, anywhere I looked, is a small, honest component that shows bytes, rate and time remaining, survives a refresh, says it is safe to close the tab, and, because the API reports only that a device is unavailable and not why, lists the likely reasons in plain words. That is the first thing worth shipping.

Why there are so few examples

I looked for good examples of the Prompt API in action, and found surprisingly few. Chrome's own case studies are real but narrow: CyberAgent's blogging tools, review summaries on redBus and Miravia, article summaries at Terra and BrightSites, and translation at Policybazaar and JioHotstar. Most of those use the task APIs, summarising and translating, rather than the Prompt API itself. Chrome uses Gemini Nano for its own scam detection. Beyond that, there are demos, extensions, and wrappers that expose the browser's model to other programs. That surprised me, until I listed the reasons:

The case against, taken seriously

Chrome shipped the Prompt API over formal opposition. Mozilla's position is that the API has "severe negative consequences to the interoperability, updatability, and neutrality of the web platform." Apple's WebKit team filed an opposing position on interoperability, portability and privacy. The W3C Technical Architecture Group recorded several concerns, and Microsoft, which builds its own version of the API on its own models in Edge, objected too. The objections come down to three things.

I agree with most of this, and it lines up with what I keep writing about: sovereignty, determinism and explainability. A model I cannot pin, version or inspect, behind a policy I did not write, is not sovereign infrastructure. But I do not think the answer is to wait. Edge already offers the same API shape on its own models (Phi-4-mini, and a smaller one called Aion in developer preview), so the interface is not Chrome's alone, even if the standards process is unfinished. And there is a design discipline that limits the damage:

Why this fits what we have built

This is the part I find most exciting. A small model with a small context window is a bad match for raw prose and a very good match for the graphs this site is building.

In The agent is the reader, the five readers' graph answered a question about any of 45 concepts in a median of 2,579 tokens, with every statement anchored to its sentence, against 25,698 tokens to read the five articles. A context window measured in thousands of tokens, not hundreds of thousands, cannot hold the five articles. It can hold the graph's answer, with room left for the question. The same work that makes the graph worth paying for, for an agent with a big model, makes it usable for a small model in the reader's tab. And the anchors mean every answer can show its sources, which is the honest way to use a model this size.

The model also fills the gaps the client-side pattern still had. The newsroom builds personas from what the reader reads, using rules; a local model can infer them from the text itself, without anything leaving the browser. My whole publishing workflow starts with voice memos; the Prompt API accepts audio, so a memo could be transcribed and structured on the device before it is encrypted into a vault lane. And the synthetic readers found 51 missing figures out of 527; a local model can describe what each figure shows, and say when a figure says something different from its caption.

Where we can use it

Where a model in the tab fits, by how much it would be worth to readers and how ready the platform is for it today. The top right is where to start; the left column waits on tool calling, more languages or mobile.
OpportunityWhat the model doesBuilt onReady today?
Ask this articleAnswers a reader's question from the graph slice, citing anchorsPrompt API, JSON schema, the five readers' graphYes
Personas inferred locallyLabels what a reader reads into persona traits, in the browserPrompt API, JSON schema, SG Meter's historyYes
A download panel that tells the truthNot the model: a component that shows bytes, rate and time, survives a refresh, and explains "unavailable"availability(), the download monitorYes
Voice memo to vaultTranscribes and structures a memo on the device, then encrypts it to a lanePrompt API audio input (GPU), append lanesYes, on GPU machines
Alt text and figure checksDescribes each figure; flags a figure that contradicts its captionPrompt API image inputYes
Read it in PortugueseTranslates an article for pt.newsroom readers on the deviceTranslator APIYes
A behaviour-policy helper for RiskMandateClassifies each rule a user writes as boundary, setting, expectation or none, privatelyPrompt API, JSON schemaYes, as a draft for a person to check
A local second readerChecks a draft against rules and sources before it is sentPrompt API, JSON schemaPartly: too small to be the only check
Tools for the reader's agentThe newsroom describes its own actions to the reader's agentWebMCPOrigin trial
Native tool callingThe model decides which of our functions to callPrompt API toolsNot yet: behind a flag

The first four are the ones I want to build, in roughly that order. They share one property: the model never decides anything important on its own. It answers from anchors, labels into a fixed set, or transcribes, and the page or the person checks the result.

What we are building next

A vault is being built alongside this article to explore all of this properly: a test page per capability, measured on real machines, with the numbers published. The first round:

  1. The model status panel. An open component, in the spirit of SG Meter: it checks support and eligibility, shows download progress with bytes, rate and time remaining, survives a refresh, and, when a device cannot run the model, explains the likely reasons in plain words.
  2. Ask this article, grounded in the graph. A question box on articles that have the five readers' graph, answering only from anchored items, with a JSON schema that forces it to cite them, and a fallback that shows the graph slice when the model is missing.
  3. Personas, inferred on the device. The same traits SG Meter uses, inferred from the text a reader keeps, never sent anywhere.
  4. Voice memo to vault. Record, transcribe and structure on the device, then encrypt and drop it into an append lane.
  5. Measurements. Download time, time to first token, tokens per second, context window, and accuracy against the anchors, on the machines we have. These are the numbers missing from almost every piece about Gemini Nano, including this one.

The client-side pattern had one piece left that seemed to need someone else's server. On capable desktops, it no longer does. The model is small, the platform is contested, and the first-run experience is poor. All three are true, and none of them is a reason to wait. If you are building with the Prompt API, or have found good examples I missed, let's talk: agent@riskmandate.ai.

Where this comes from

A voice memo of mine, recorded after trying Gemini Nano in Chrome for the first time, with my first impressions of the download and the demos. I asked an agent in a separate Claude session to research the current state of the Prompt API and the built-in AI APIs, check my impressions against Chrome's own documentation, and draft this article with its figures; where my impressions and the documentation disagreed (the download stopping on refresh, native tool support), the article follows the documentation and says so. This site's agent then gave it a last pass: plainer wording, and a correction to the claim that renting was the only way to put a model behind a page. The argument is mine, and so is the editorial responsibility.

The status of each API, the hardware requirements, the Prompt API's inputs, outputs and sessions, and the model's download, update and deletion behaviour are from Chrome for Developers: Built-in AI APIs, The Prompt API, Understand built-in model management and Inform users of model download, with WebMCP and agents and the demos. The ship date and the objections are from coverage in Tech Times and Gigazine, which cite Mozilla's standards position (issue 1213) and WebKit's (issue 495); the model size of about 4.27 GB is from the same coverage. The state of native tool calling, tested on Chrome 151 in August 2026, is from the flutter_gemma_builtin_ai documentation. Edge's Prompt API and its models are from Microsoft's announcement and coverage of Build 2026. The graph numbers are from The agent is the reader; the figure count is from How to run synthetic users. No measurements of speed or quality are claimed here; those are what the vault is for.

Threads

Site & engineeringGraphs & knowledge This article as a graph →

Builds on

All articles · All graphs

Want the next issue by email. One issue a week or so: what was published, what it adds up to, and what is worth your time. Subscribe to the SGit Newsroom →

← All articles