Tool Choice Is a Behavioral Problem
Evaluate browser-agent tool discovery, selection, argument construction, execution and result use as separate behavioral stages.
Chapter 18 built a tool that is correct by construction: its contract is tested, its results are bounded, its lifecycle is traced. None of that determines whether an agent uses it well.
Once a model can see several capabilities, the outcome depends on a behavioral pipeline:
discovery → selection → argument construction → admission → execution → result use
A single “task passed” score cannot tell you which stage failed. Chapter 13 made this argument for model output: a resolved promise is not a correct answer. The same layering applies to tool use, with more stages and higher stakes.
1. Discovery precedes choice
A tool that never enters the agent’s effective context cannot be selected, and its absence is easy to misread as a reasoning failure. For every run, record:
- tools registered by the page;
- tools visible to the agent after filtering;
- tools removed by user or developer policy;
- tools omitted because a list limit was reached;
- each tool’s description and origin;
- the effective tool-list size.
Tool-list pressure is real. An agent choosing among five operations faces a different problem from one choosing among five hundred: more routing ambiguity, more context consumed before reasoning starts, more room for a persuasive description to dominate. The size of the list is a variable, not a constant.
2. Selection needs negative cases
If every fixture in the suite requires a tool call, the evaluation rewards unnecessary action and teaches nothing about restraint. A tool-choice suite needs cases where the correct behavior is:
- call
search_book; - call
get_chapterinstead; - call a narrower search;
- ask the user for a missing detail;
- use no tool at all;
- refuse a forbidden action;
- stop after an empty result rather than retrying or inventing.
The Chapter 19 suite carries all of these. Three fixtures expect a specific tool; one expects a clarifying question; one expects a plain answer with no tool; one expects refusal of an unauthorized publish_note; one expects the agent to stop after a search returns nothing.
3. Ambiguous boundaries test the interface, not the model
Introduce controlled competitors:
search_book — all registered chapters
search_current_chapter — only this chapter
search_site — the wider site, outside the book index
The purpose is not to trick the model. It is to find out whether the capability boundaries are intelligible. If an agent repeatedly searches the whole book when the task clearly concerns one chapter, that is evidence about the tools, and the fixes are interface fixes:
- rename them so the scope is in the name;
- narrow one so the overlap disappears;
- merge two that no one can tell apart;
- rewrite the descriptions to distinguish rather than advertise;
- reduce the set exposed for this class of task.
Evaluation that only ever produces prompt tweaks has missed half its value. The Chapter 19 lab runs each fixture under three metadata profiles — clear, overlapping, and a deliberately persuasive search_site description — so a selection shift can be attributed to the wording rather than guessed at.
4. Arguments need semantic checks, not just schema checks
Schema-valid arguments can still be wrong. Separate the questions:
| Question | Example failure |
|---|---|
| Syntax / schema | limit: 11 when the tool caps at 10 |
| Semantic relevance | Query terms that don’t match the task |
| Scope | Searching the whole book when the task names one chapter |
| Efficiency | limit: 10 when one result would answer the question |
| Authorization | Naming a tool that policy denies |
| Resource cost | A query that pulls the maximum excerpt budget every call |
limit: 10 is schema-valid and often wasteful. A chapter identifier can parse while naming the wrong book. Argument evaluation protects relevance, scope, cost and authority — the things a shape check cannot see.
5. A successful call is not a grounded answer
The result of a tool call is an untrusted input. It may be stale, it may be an application error, it may contain hostile text. After execution, evaluate whether the model:
- uses the actual returned content rather than its prior belief;
- attributes or cites the specific result it used;
- stops when the result is empty, instead of filling the gap from imagination while implying the search succeeded;
- distinguishes a tool failure (“the search errored”) from no evidence (“the search found nothing”);
- does not invent content that the corpus did not contain.
The Chapter 19 lab evaluates this as a second decision, after execution, against its own schema: usedToolResult, citedChapters, and an action of answer, clarify, stop, refuse or not-applicable. One specimen deliberately answers “the book discusses this topic” after an empty result — the exact failure the metric exists to catch.
6. Score the stages independently
One trace yields several metrics, and collapsing them into an “agent score” destroys the diagnosis:
| Stage | Metric |
|---|---|
| Discovery | Was the required tool visible? |
| Selection | Was the correct tool chosen? |
| Arguments | Schema-valid and semantically appropriate? |
| Admission | Was the policy decision correct? |
| Execution | Did the tool complete? |
| Result use | Is the final answer grounded in the result? |
| — | Unnecessary tool-use rate |
| — | Abstention quality |
An agent might choose the right tool with the wrong chapter number. Another might invoke it perfectly and then ignore what it returned. A third might call a tool where the right move was to ask a question. These need different fixes, and only stage-level metrics point at the right one.
7. Compare behavior across configurations
This is where the worker architecture from Part IV reconnects. The local coordinator can send the same tool-choice fixtures to multiple browser workers — different Chrome builds, different model configurations — each result carrying its worker provenance, its discovered tool set and its full decision trace.
Paired runs let us ask precise questions when something changes:
- did unnecessary tool use go up?
- did correct selection go down?
- did invalid-argument rate change?
- did abstention quality change?
- did result grounding change?
- what did latency and context cost do?
Comparing changed cases is far more informative than comparing final prose. A browser update that leaves every answer looking fine while doubling the unnecessary-tool-use rate is a regression the prose comparison would miss.
8. The decision trace
The Agent Inspector should reconstruct a complete tool decision as a sequence:
tool set observed
→ candidate proposed
→ arguments generated
→ arguments normalized
→ policy decision
→ tool execution
→ result
→ final answer
Each transition carries a correlation ID, a timestamp and the relevant version — policy version, suite version, schema hash. A rejected argument stays visible next to the candidate that produced it. When a run fails, the developer should be able to point at one arrow and name the stage.
9. Run the tool-choice evaluator
Select Run with Browser AI from Chapter 19 to open:
/tools/ai/browser-ai-from-first-principles/19-chapter/
The evaluator sends the Prompt API an explicit catalog snapshot and a schema-constrained request for a selection decision, then a second request for a result-use decision. It scores discovery, selection, arguments, admission, execution and result use separately, under the three metadata profiles, and it can replay pre-recorded candidate specimens so the scoring logic itself can be checked.
What this establishes: the stage-level scoring is implemented, the negative cases are present, and the metadata profiles let wording effects be measured. What remains specified rather than observed: this is a controlled routing trial. The catalog is handed to the model explicitly; the model returns schema-constrained JSON. It is not evidence that an autonomous browser agent discovered the Chapter 18 WebMCP registration on its own — search_site is marked unsupported and publish_note denied in the lab, not exercised against a live agent runtime. The recorded trace remains the Prompt API lifecycle.
Conclusion
Tool choice is not one decision. It is discovery, selection, argument construction, admission, execution and result use, and each stage fails for its own reasons and needs its own metric.
Persistent confusion between two tools is a fact about the interface, and the fix belongs in the interface. The next chapter adds hostile content and consequential capabilities, where a mistaken stage stops being a poor answer and becomes a security failure.
Sources and further reading
- Chrome for Developers, WebMCP.
- Model Context Protocol, Specification.
- OWASP, LLM Prompt Injection Prevention Cheat Sheet.