Browser AI From First Principles
Build browser-native AI from the capability boundary upward: local models, observability, WebMCP tools, agent security, and a user-owned browser policy engine.
For most of the history of web applications, the browser has occupied a clear position in the architecture.
It displays an interface, executes application code, and stores local state. It communicates with servers that perform the expensive work.
AI initially fitted neatly into that arrangement. A web application collected context, sent it across the network to a model running in a data centre, received a response, and displayed the result.
That architecture is changing.
Chrome’s built-in AI work allows pages and extensions to request AI capabilities backed by models managed by the browser. At the same time, WebMCP gives websites a way to expose structured tools to browser agents.
Put those developments together:
model inside browser
↓
context inside browser
↓
tools inside browser
↓
agent inside browser
This book studies that transition from first principles.
The durable question is not which model happens to ship in one version of Chrome. It is:
What happens when the browser stops being merely an interface to AI and becomes an AI runtime itself?
The book is also a tool
This project goes beyond explaining APIs.
Throughout the book we will build one system through two reader-facing stages.
Stage One is the Browser AI Laboratory: a full-page experience in which chapters receive live experiments as their corresponding tool pages exist, with capability detection, real execution when supported, captured-trace replay when it is not, and an inspectable evidence trail.
Stage Two is the AI-enabled site: the books and wider website become structured, retrievable context with typed operations, interchangeable AI runtimes and policies controlled by the user.
Browser AI Observatory is the evidence layer underneath both stages. Its DevTools extension begins as a capability probe and streaming prompt console, then grows into an application monitor, trace viewer, profiler, evaluation harness, WebMCP agent inspector and personal browser policy engine.
The extension will help answer questions that a normal chat interface hides:
- Is the API absent, unsupported, downloading, loading or ready?
- How long did session creation take?
- How long did the user wait for the first output?
- Which prompt used which session and capability options?
- Did a request complete, fail or get cancelled?
- Did successful inference produce an application result that passed validation?
- Did a browser or model update change behavior on a known fixture?
- Which WebMCP tools were exposed to an agent?
- Which tool was selected, with what arguments and authority?
- Did untrusted page content influence a privileged action?
- Which rule changed a page, and what evidence justified the change?
- Which model or provider handled the content, and did any data leave the device?
The product is not saved for the final chapter. Chapter 1 establishes the laboratory shell. Chapter-specific reader experiences live under content/tools/ai/browser-ai-from-first-principles/. Shared laboratory and Observatory code lives under experiments/browser-ai-from-first-principles/. A chapter receives the Run with Browser AI action only when its matching experience exists, while the Observatory records the evidence required to trust it.
flowchart TD
R[Research question] --> E[Experiment]
E --> T[Observatory trace]
T --> W[Book chapter]
W --> F[Extension feature]
F --> R
The book becomes the design record for the software.
The software becomes the experimental instrument for the book.
The method
We will add one mechanism at a time.
capability detection
↓
model acquisition
↓
session and streaming
↓
observability
↓
evaluation
↓
structured tools
↓
permissions and policy
↓
agent runtime
At every step we will ask:
- What new capability did the mechanism add?
- Which failure can it introduce?
- What evidence tells us whether it worked?
- What authority does it require?
- What should the debugger record?
The aim is not to memorize experimental flags or one release’s model name. Those will change. The aim is to understand the architecture that remains when the implementation underneath the API changes.
Book structure
Part I — The Browser Becomes a Runtime
- The Browser Stops Being Just a Client — Separate inference architectures, then define the chapter-driven laboratory and AI-enabled site we will build.
- What Does Built-In Actually Mean? — Turn capability availability, download, session creation and prompting into an explicit state machine.
- Build the Smallest Browser AI Observatory — Create a permission-minimal Manifest V3 DevTools extension that runs and measures a local prompt.
- From Prompt Demo to AI Debugger — Add opt-in application instrumentation, a secure bridge, trace identity, redaction and retention.
Part II — The Built-in AI Primitives
- One Runtime, Several Interfaces — Compare a general Prompt API with task-specific APIs without assuming one physical model.
- Summarization Is a Contract — Test length, type, format, grounding and long-input strategies.
- Writing, Rewriting and Proofreading Are Different Operations — Separate generation, transformation and correction by authority over the text.
- Language Detection Is a Decision, Not an Answer — Compose ranked detection with translation, thresholds and abstention.
Part III — Sessions, Context and Local Resources
- A Session Is Not a Stateless Function — Inspect conversation state, context consumption, cloning and destruction.
- When the Context Window Fills — Detect overflow, compact history and preserve the right state.
- Cold Starts, Warm Runs and Real Latency — Measure download, load, first output and throughput separately.
- The Browser Manages the Model — Examine eligibility, updates, purging, storage pressure and implementation change.
Part IV — Reliable Browser AI
- A Resolved Promise Is Not a Correct Answer — Build behavioral fixtures beside operational traces.
- Evaluate the Feature, Not the Demo — Create repeatable evals, baselines and release gates.
- When Local Is Not Available — Design explicit degradation and hybrid fallback without hiding data movement.
- Structured Output Is Still Model Output — Parse, validate, repair and reject machine-consumed responses.
Part V — Tools and the Agentic Web
- From Buttons to Capabilities — Understand why structured tools differ from UI actuation.
- Expose the First WebMCP Tool — Register a typed, read-only operation and inspect its contract.
- Tool Choice Is a Behavioral Problem — Evaluate discovery, selection, argument construction and results.
- Untrusted Text Meets Executable Authority — Contain hostile page text, manipulative tool metadata and contaminated tool results by separating proposal from authority.
Part VI — From Agent Inspector to User-Owned Browser
- Put a Human at the Authority Boundary — Add permission scopes, confirmations, audit trails and safe defaults.
- The Browser as a Personal Policy Engine — Combine interchangeable AI runtimes, content units, typed evidence and declarative user rules.
- Build an AI-Origin Content Filter — Prefer provenance, preserve uncertainty and apply reversible presentation choices.
- The User, Not the Platform, Controls the Interface — Assemble the coordinator, workers, tools, authority boundary and policy layer into the finished system.
This edition preserves the architecture above. Future experiments may justify changes in a later edition; when the evidence changes, the next edition should change with it.
What we will build
Browser AI Observatory will grow through the same parts.
| Book stage | Extension capability |
|---|---|
| Part I | Full-page laboratory, capability probe, model-download monitor, streaming trace and application bridge |
| Part II | Task API console and side-by-side comparisons |
| Part III | Session inspector, context tracking and performance profiles |
| Part IV | Fixture runner, output validation, baseline comparison and fallback trace |
| Part V | WebMCP tool explorer, call trace and security warnings |
| Part VI | Permission gate, agent timeline, personal policy engine and AI-origin content filter |
The extension will maintain an explicit evidence boundary.
It can directly observe operations it owns. It can receive application events through an opt-in SDK. It can inspect browser surfaces exposed by supported DevTools APIs. It will not pretend that application-reported events are browser-attested or that private browser internals are public instrumentation.
That limitation makes the tool more credible, not less.
Experimental status
Browser AI APIs and WebMCP are evolving. Some capabilities may be stable, some may require an origin trial, and others may be restricted to experimental Chrome channels or the Built-in AI Early Preview Program at the time a chapter is written.
Publication snapshot – September 6, 2026. API status and syntax in this edition were rechecked against current Chrome documentation on September 6, 2026. Experimental runs retain the browser build, flags, operator labels and traces from the environment in which they were actually recorded. A later documentation change does not rewrite an earlier observation.
Every implementation chapter will therefore distinguish:
- the architectural contract;
- the currently documented public API;
- experimental configuration used for a particular run;
- observations from our own environment;
- assumptions that still need to be tested.
The book will not silently turn an EPP announcement into a universal availability claim.
The experimental build is valuable precisely because it lets us test whether an application survives changes behind the same capability boundary. It is evidence, not permanence.
The larger idea
A book normally explains software.
This project uses writing, implementation and experimentation as one loop:
research
↓
experiment
↓
chapter
↓
software
↓
real use
↓
new evidence
└──────────→ research
Eventually the same browser-native AI mechanisms described here can run inside technical books and documentation: explain a selection, challenge a claim, expose prerequisites, generate an experiment, navigate related concepts and offer structured operations to an agent.
The immediate product is a live browser laboratory backed by an AI monitor and debugger. It grows into an AI-enabled site and personal browser policy engine.
The larger goal is to discover what happens when written knowledge becomes operational software.
Chapters
The Browser Stops Being Just a Client
Rebuild the browser AI architecture from first principles, then define the chapter-driven laboratory and AI-enabled site we will construct throughout the book.
What Does Built-In Actually Mean?
Turn Chrome's built-in AI APIs into an observable state machine: inspect capability exposure, availability, model acquisition, session readiness, execution and failure in the chapter itself.
Build the Smallest Browser AI Observatory
Build and run the smallest useful Browser AI Observatory: probe the Prompt API, create a session, stream a response, derive measurements from events, export evidence, and preserve the instrument boundary.
From Prompt Demo to AI Debugger
Turn the Prompt API console into an application-aware debugger with an executable, opt-in instrumentation bridge, versioned events, redaction, correlation, persistence and explicit evidence limits.
One Runtime, Several Interfaces
Build and run a seven-interface Browser AI switchboard: share lifecycle instrumentation while preserving each API's identity, options, operation and output contract.
Summarization Is a Contract
Configure, run and evaluate a browser-managed Summarizer contract across type, length, format, preference, grounding and omission while keeping automatic checks separate from human judgment.
Writing, Rewriting and Proofreading Are Different Operations
Run Writer, Rewriter and Proofreader as distinct authority contracts, preserving each operation's permitted delta, options, output shape and evaluation rather than grading generic prose.
Language Detection Is a Decision, Not an Answer
Run a staged Language Detector and Translator laboratory that preserves ranked uncertainty, applies visible confidence and margin thresholds, abstains explicitly, and inspects pair-specific availability before translation.
A Session Is Not a Stateless Function
Run a session-lineage laboratory that creates one Prompt API resource, establishes shared history, clones the measured state, diverges two branches, compares usage and ends both explicitly.
When the Context Window Fills
Run a context-budget laboratory that measures or estimates the next turn, applies an explicit safety reserve, probes the runtime separately, validates a structured checkpoint and creates a replacement session.
Cold Starts, Warm Runs and Real Latency
Run a latency-decomposition laboratory that measures capability inspection, monitor callbacks, session creation, first output, generation, completion, validation and cancellation under explicit fresh or operator-verified cold protocols.
The Browser Manages the Model
Engineer against a browser-managed capability whose model, eligibility, assets and accepted configuration can change, while preserving evidence, compatibility and graceful recovery.
A Resolved Promise Is Not a Correct Answer
Separate runtime success from structural validity, semantic correctness and feature success, then place behavioral evaluation beside the browser AI trace.
Evaluate the Feature, Not the Demo
Turn browser AI examples into representative feature evaluations with slices, baselines, repetitions, regression gates and reviewable evidence.
When Local Is Not Available
Design explicit degradation across local browser workers, deterministic alternatives and remote fallbacks without hiding capability changes or data movement.
Structured Output Is Still Model Output
Constrain, parse and validate model-produced JSON without confusing syntactic structure with semantic correctness or executable authority.
From Buttons to Capabilities
Replace fragile visual browser automation with explicit, typed capabilities while preserving user authority and observable effects.
Expose the First WebMCP Tool
Expose a typed read-only book operation through WebMCP, validate its arguments, bound its results and inspect the complete call lifecycle.
Tool Choice Is a Behavioral Problem
Evaluate browser-agent tool discovery, selection, argument construction, execution and result use as separate behavioral stages.
Untrusted Text Meets Executable Authority
Defend browser agents when page content, tool metadata and tool results can influence models that possess executable capabilities.
Put a Human at the Authority Boundary
Design approvals around concrete effects, bind grants to validated effects and preserve a complete record of consequential browser-agent actions.
The Browser as a Personal Policy Engine
Turn browser-native AI, interchangeable runtimes and observable decisions into a user-owned policy layer for the web experience.
Build an AI-Origin Content Filter
Build a provenance-first, uncertainty-preserving and reversible browser filter for content with evidence of AI origin.
Build Browser AI Lens: The User-Owned Intelligent Browser
Assemble the frozen Browser AI Lens from the book's locked runtimes: four modes over one inspectable substrate, demonstrated end to end.
Appendix A — Browser AI Field Guide: Setup, APIs, States, Diagnostics and Configuration
The working reference for the book: development setup, API matrix, availability and lifecycle states, configuration keys, event dictionary, diagnostics, authority states, and the Lens cheat sheet.
