← Browser AI From First Principles

Cold Starts, Warm Runs and Real Latency

Run a latency-decomposition laboratory that measures capability inspection, monitor callbacks, session creation, first output, generation, completion, validation and cancellation under explicit fresh or operator-verified cold protocols.

“Local AI is faster” is not a measurement.

It compresses several different waiting periods into one adjective. A browser-native feature can avoid network round trips and still make a user wait for model acquisition, process startup, session creation, context ingestion or slow generation.

To understand latency, we have to take the lifecycle apart.


1. One duration hides several clocks

For one interaction, useful timestamps include:

PhaseStartsEndsUser-visible?
Capability inspectionavailability() callstate returnedUsually not
Acquisitionsession request requires assetsassets readyYes, if it blocks
Session creationcreate() calledsession returnedOften
Prompt startupprompt calledfirst chunkYes
Generationfirst chunkfinal chunkYes
Validationoutput completefeature accepted/rejectedSometimes

The perceived wait for a streaming feature is often dominated by time to first output:

$$ L_{first} = t_{first\ chunk} - t_{prompt\ start} $$

Completion latency is different:

$$ L_{total} = t_{prompt\ finish} - t_{prompt\ start} $$

A feature can feel responsive with a low $L_{first}$ even when $L_{total}$ is large. A non-streaming task can have acceptable total latency while providing no intermediate feedback.


2. The first real observation measured only acquisition

Our first exported Chrome trace contained ten events. It established:

  • LanguageModel was exposed;
  • availability returned available;
  • session creation completed in approximately 8.2 milliseconds;
  • the session reported quota and lifecycle capabilities;
  • no prompt was executed.

Therefore the trace says nothing about prompt latency.

This sounds obvious, but it is a common measurement error. A fast session constructor is not a fast answer. A fast first chunk is not a fast completion. A completed generation is not a successful feature.

The missing values in the report are correct:

MeasurementFirst trace
Session creation8.2 ms
Time to first chunkNot measured
Total prompt timeNot measured
Fixture resultNot measured

Absence is better than a number inferred from the wrong event.


3. A progress callback is not proof of a download

The same trace contained two downloadprogress callbacks:

loaded = 0
loaded = 1

They arrived roughly 0.7 milliseconds apart while capability state was already available.

The callback name comes from the API. Interpreting it requires the surrounding state. This event sequence cannot reasonably establish that a complete model asset was transferred during those 0.7 milliseconds.

The supported conclusion is narrower:

The session monitor emitted acquisition-progress callbacks during creation.

It may be notifying listeners of an already-complete resource state. It may represent another browser-managed step. Without network or browser-internal evidence, the Observatory should not relabel it as bytes downloaded.

This is why traces need events rather than prose such as “model downloaded successfully.”


4. Cold and warm are operational definitions

“Cold” can mean several things:

  • model absent from disk;
  • model installed but process not started;
  • browser restarted;
  • session absent;
  • OS file cache empty;
  • accelerator context absent.

An extension cannot reliably force or verify all of them.

Our run protocol therefore uses explicit operational labels:

  • Download path: availability indicates acquisition and progress is observed.
  • Cold start: first session and prompt after browser restart; the model may already be installed.
  • Warm session: the identical prompt runs again in the same live session after only the trace is cleared.

These labels say what we did. They do not claim complete control of hidden caches.


5. Compare paths, not isolated numbers

For matched runs, useful deltas include:

$$ \Delta L_{session} = L_{session,cold} - L_{session,warm} $$$$ \Delta L_{first} = L_{first,cold} - L_{first,warm} $$

The comparison is credible only if we hold important factors constant:

  • browser build and flags;
  • hardware and power state;
  • fixture input and creation options;
  • capture mode;
  • competing workload;
  • repetition count;
  • whether the session is reused or recreated.

Even then, three repetitions produce an initial estimate, not a performance law.

The report uses medians because one startup interruption can dominate a mean in a tiny sample. We should still retain every raw observation, including failures and slow outliers.


6. Throughput needs a denominator we can defend

Streaming produces chunks, but chunks are not tokens. One implementation may emit one token per chunk; another may batch several; another may revise or replace accumulated text.

The Observatory can safely report:

  • chunk count;
  • output characters and bytes;
  • time to first chunk;
  • total duration;
  • character throughput as an interface-level approximation.

It should not report tokens per second unless the runtime supplies a reliable token count or we explicitly label a tokenizer-based estimate.

For output characters $C$ after the first chunk and generation interval $G$:

$$ R_{chars} = \frac{C}{G} $$

This is useful for comparing the same fixture under the same output constraints. It is not a model benchmark comparable across languages and formats.


7. Streaming changes cancellation

A long generation creates an opportunity to cancel:

const controller = new AbortController();

const result = session.promptStreaming(input, {
  signal: controller.signal
});

stopButton.addEventListener("click", () => controller.abort());

Cancellation latency matters too:

user clicks Stop
      ↓
abort requested
      ↓
stream ends
      ↓
session state observed

We need to measure whether partial output remains visible, whether the session is reusable, and whether usage increases after an aborted operation.

An AbortError is an expected outcome when the user stops work. It should not be counted as a model failure.


8. Benchmark the feature boundary

A local generation can be fast while the user-facing feature remains slow because of:

  • DOM extraction;
  • document segmentation;
  • prompt construction;
  • model acquisition;
  • output parsing;
  • validation;
  • rendering;
  • application-side persistence.

The Observatory’s direct prompt trace measures the runtime boundary. Its application SDK can add the surrounding feature events.

selection captured
    ↓
context prepared
    ↓
prompt started
    ↓
first chunk
    ↓
prompt finished
    ↓
validation finished
    ↓
UI committed

Only then can we explain what the user waited for.


9. Model changes require paired experiments

The motivating experiment compares the same application boundary under different browser-managed configurations.

The Observatory records the model/configuration label as operator-supplied because the API does not attest a model identity to the extension. A report can therefore say:

Under the operator-recorded “Gemma 4 flag enabled” configuration on Chrome 152.0.7977.65, this run produced the following measurements.

It cannot say:

Chrome proved that Gemma 4 generated this output.

This distinction may feel cautious, but it is what makes later comparisons credible.


10. The next corpus

The minimum useful performance corpus is now defined:

  1. one cold Prompt fixture run;
  2. one identical warm-session run;
  3. one cloned-session continuation;
  4. one explicit cancellation;
  5. three repetitions for any latency claim we publish;
  6. the same matrix under the comparison configuration;
  7. raw failures retained beside successful runs.

The chapter can explain the measurement system now. Its numerical comparison remains open until those traces exist.

That is not a weakness. It is a visible experimental boundary.


11. Run the latency-decomposition laboratory

Select Run with Browser AI from Chapter 11 to open:

/tools/ai/browser-ai-from-first-principles/11-chapter/

The laboratory gives each clock its own field and retains every raw observation. The workflow is:

  1. supply an operator configuration label;
  2. choose an honest first-path protocol;
  3. inspect capability availability;
  4. create a measured session;
  5. run the fixed first-path fixture;
  6. repeat the identical fixture in the same warm session;
  7. run cancellation as a separate protocol;
  8. destroy the session explicitly.

The default protocol is fresh session. It means only that the page created a new session in the current browser process. Hidden model, process, operating-system and accelerator caches remain unknown.

The alternative cold start · operator verified restart is not inferred by JavaScript. The operator should select it only for the first session and prompt after manually restarting Chrome. The chosen protocol and configuration label are recorded with creation and prompt events. Both controls are locked while the session exists so one resource cannot change experimental identity midway through its lifecycle.

Keep the fixture fixed

First-path and warm measurements use the same creation options and exact prompt:

Reply with exactly this sentence and nothing else: Browser AI timing fixture complete.

The lexical assertion is intentionally narrow. It confirms that the timed operation returned the expected fixture rather than a refusal or unrelated response. It does not turn the latency run into a general quality evaluation.

The first-path control records one observation. The warm control runs the configured number of repetitions sequentially in the same live session. Three repetitions are the default minimum for displaying an initial median. The table retains failures and outliers beside successful rows.

Repeated warm prompts are still stateful. Their context grows, and hidden caching remains uncontrolled. The laboratory exposes a reproducible application protocol; it does not claim a hardware-isolated model benchmark.

Decompose every prompt

For a streaming operation the module records:

  • session creation duration, when the row is a first-path run;
  • time from prompt start to first chunk;
  • time from first chunk to stream completion;
  • total prompt duration;
  • validation duration;
  • chunk count;
  • output characters and bytes;
  • character throughput over the observed generation interval;
  • outcome and session identity.

For a non-streaming operation, first output arrives with the complete result. Total time remains measurable. The separate generation interval and character throughput remain Not measured because there was no independently observed first-chunk boundary.

Chunk count is never relabelled as token count.

Treat monitor callbacks as observations

During creation, the UI displays callback count and the elapsed span between first and last callback. Event fields retain the API’s reported loaded and total values.

The panel deliberately says monitor callbacks, not model bytes downloaded. A sequence from zero to one may describe acquisition state without establishing a network transfer, asset size or cache miss.

Measure cancellation from the request

The cancellation probe uses a deliberately long fixture and an editable automatic abort delay. The operator can also press Stop.

The trace records:

prompt started
first chunk, if any
abort requested
stream ended
post-operation session snapshot

Cancellation latency is reported only when an abort request is followed by an aborted outcome. If generation completes before the timer, the row says completed-before-abort. If completion occurs after an abort request without an AbortError, the request-to-completion interval can be retained diagnostically, but it is not presented as cancellation latency.

Partial output remains visible and is measured by characters and bytes. The session stays explicitly addressable so a later experiment can test whether it remains reusable.

Read medians with their sample counts

The summary reports medians independently for session creation, first-path time to first chunk, warm time to first chunk and warm total time. Every median includes its sample count. A first-path median based on one row remains visibly `n=1).

The purpose is comparison without erasure:

summary median
      ↓
sample count
      ↓
complete raw rows
      ↓
trace events and failures

A slow outlier is not removed merely because the median is more stable than the mean.

Replay only the clocks that exist

The reviewed September 2 trace contains two capability inspections, two monitor callbacks and one session creation lasting approximately 8.2 milliseconds. It contains no prompt.

Replay therefore adds a session only row. First chunk, generation, total prompt, validation, character throughput and cancellation remain Not measured. The callback span is displayed as an observation without claiming that a complete model download occurred in that interval.


Conclusion

Browser AI latency is a lifecycle, not a stopwatch.

Capability inspection, acquisition, session creation, first output, completion and validation answer different questions. Progress callbacks require context. Cold and warm require operational definitions. Chunk count must not masquerade as token count.

Our first real trace measured a fast session creation and nothing about inference. The instrument correctly left the missing fields empty. Chapter 11 now makes the full protocol executable, but numerical cold-versus-warm claims remain open until matched live runs exist under explicitly recorded conditions.

The next chapter asks why the same session configuration could work and then fail. Once the browser manages the model, model lifecycle and compatibility become part of application engineering even when the application never downloads a model itself.


Sources and further reading

  1. Chrome for Developers, The Prompt API.
  2. Chrome for Developers, Get started with built-in AI.
  3. Chrome for Developers, Understand built-in model management in Chrome.