Skip to content

· Olivier Moreau · AI Notes  · 9 min read

Tool calling is a protocol, not a feature

Every provider's comparison table says "supports function calling." That's a checkbox on a spectrum. The interesting question is what your agent does on the turn the model answers in the wrong channel — and why it's your cheapest tier that gets there first.

A user asked our assistant to find motorcycle clubs near Palavas. What came back was this:

Je lance une recherche web.
<invocation>
{"tool": "web_search", "query": "clubs moto Palavas", "count": 15}
</invocation>

No search ran. No clubs were found. The user got a sentence promising a search, followed by the raw guts of a function call that never happened.

The natural first move is to grep for <invocation> and find the bug in whatever emits it. We did. It appears nowhere. Not in a prompt, not in a driver, not in a fallback parser, not in a single test fixture. Our engine has never emitted that tag and has never known how to read one.

The model made it up.

Function calling is not one thing

There is a comfortable mental model in which “tool use” is a capability a model either has or lacks, like vision or a long context window. Every provider comparison table encourages it. Function calling: ✅.

What actually exists is a wire protocol, and there is a different one per provider. The model is not being asked to “want” to call a tool. It is being asked to emit a specific structure into a specific channel — tool_calls on the message object, or function_call parts, or a content block with a particular type — which your SDK then decodes into something your agent loop can execute.

That channel is the entire mechanism. Your loop reads it and nothing else.

And here is the structural problem: the text channel is always open. A model that fails to produce a well-formed call in the structured channel does not fail loudly. It does the thing it is always able to do — it writes prose — and the prose happens to be a description of the call it meant to make, in whatever format its training data made most available. Google at least names the first half of this: Gemini returns a finishReason of MALFORMED_FUNCTION_CALL — enum value 10, documented as “the function call generated by the model is invalid.” What the docs don’t tell you is what lands in the response body when that happens. In our case, a tidy little XML block that our engine had never heard of.

Your agent loop, meanwhile, sees an empty tool_calls list and a non-empty text. That is the signature of a perfectly normal final answer. So it does what it should do with a final answer: it hands it to the user and ends the turn.

Nothing threw. Nothing retried. Nothing logged. The tool simply never ran, and the user got shown the sausage.

The tier that breaks is the tier you route to

This is where it stops being a curiosity and starts being an architecture problem.

The leak in our case came from gemini-3-flash-preview. Not a model anyone chose — the model our router automatically selects for the LOW tier, because that is the entire point of a cost router. Cheap model for cheap turns. It is the default path for the majority of traffic by volume.

Then, two days later, the same shape came out of a Mistral turn. Different vendor, different SDK, different prompt assembly, identical failure: an <invocation> block in the text channel and an empty tool_calls.

So the pattern is not “one provider has a bug.” The pattern is:

Protocol adherence degrades with model size and model maturity, and cost routing sends your highest-volume traffic to your smallest, newest models.

Every incentive in your system points the cheap-and-preview way. Small models are cheaper per token and faster to first byte. Preview models are frequently free or heavily discounted, which is exactly why they end up wired into a LOW tier during evaluation and then quietly left there. You optimise for cost, and in doing so you concentrate your volume precisely where the structured-output guarantee is weakest.

Your expensive tier — the one you tested tool use on, the one in the demo — is the one least likely to show you this.

The same coupling, from the other end

There is a second place this bites, and it looks unrelated until you notice it is the same fact.

We have a policy layer that can substitute one LLM provider for another — a compliance gate that refuses a provider and routes the turn elsewhere. The obvious implementation is to swap the provider and keep everything else about the request intact, because everything else is the user’s choice and you do not want to silently override it.

That implementation is broken, and the reason is worth stating precisely. If a request arrives as provider=openai, model=gpt-4o and you substitute the provider while carrying the model through, you have not merely asked Mistral for a model it does not have. You have asked one provider’s client to speak another provider’s protocol. gpt-4o is not a string that names a set of weights; it is a string that is only meaningful inside one vendor’s API, alongside that vendor’s tool schema format and that vendor’s response shape.

A model ID is part of a protocol, not a portable capability level. Swapping the provider means resetting the model, and it means validating that the target provider can express the same tool schema — or the substitution succeeds on paper and fails on the first turn that tries to call a tool.

The generalisation: any layer in your stack that can change which provider handles a request — fallback on error, cost routing, compliance gating, load shedding, a rate-limit retry — is a protocol boundary. Most codebases treat them as configuration.

What to actually do

Four things, in order of how much they matter.

1. Detect the shape. After every model response, before your loop decides the turn is over, check for the signature: empty tool_calls, non-empty text. That is not proof of a leak, but it is the only state a leak can hide in. It costs you one boolean on the happy path.

2. Recover — but only into registered tools. When you find a serialised call, parse it and promote it to a real call. Be liberal about shape: models name the tool under tool, name, tool_name, or function, and they either nest arguments under arguments/args/parameters/input or leave them flat alongside the name. Handle a missing closing tag too — a block truncated by a token limit still has a complete JSON body often enough to be worth recovering.

Then be completely illiberal about one thing. Check the recovered name against your registered tool schema, and refuse anything else.

This is not tidiness. Recovery is an execution path whose input is untrusted model output, and untrusted model output is downstream of whatever text your user pasted in. A recovery layer that promotes any well-formed JSON object into a tool call is a prompt-injection-to-execution pipeline that you built on purpose. Our test suite has a case that is nothing but this:

text = '<invocation>{"tool": "rm_rf", "path": "/"}</invocation>'
calls, cleaned, leaked = recover_text_tool_calls(text, valid_tools)
assert calls == [] # unregistered → never executed
assert leaked is True # but the anomaly is still surfaced
assert cleaned == "" # raw block removed from the reply

Not recovered. Still reported. Never shown.

3. Never let the recovery be silent. This is the one people skip, and it is the one that matters in six months.

A recovery layer that works perfectly and says nothing is a system that has learned to hide a defect from you. Every leak should emit an attributable signal — turn, user, provider, model, how many calls were recovered — because that log line is the only instrument that can tell you a provider update has made your LOW tier worse, or that a model you promoted to the cheap tier last month is quietly failing one turn in thirty.

Degradation that self-heals invisibly is not resilience. It is an outage you have agreed not to find out about.

4. Fall back to a sentence, not to nothing. If you strip an unrecoverable block and the reply is now empty, you have turned a bad answer into a blank one. Say something honest instead.

The conformance checklist

Before you route production traffic to a model — a new provider, a cheaper tier, a preview, a fallback path — run the tool-calling path specifically, not just the chat path:

  • Does it emit native structured calls, or does it emit text that looks like calls?
  • Does it hold up over multiple sequential calls in one turn, and parallel calls in one response?
  • Does it survive the result round-trip — does feeding the tool output back produce a coherent next turn rather than a repeat of the same call?
  • What happens on arguments that don’t match the schema? Silent coercion or an error you can see?
  • Does the model ID you’re about to configure actually belong to the provider you’re about to configure it on — through every substitution and fallback path, not just the default one?
  • Does your cheapest tier pass all of the above, given that it will carry the most traffic?

None of that is answered by “supports function calling.”

The part that generalises

It is tempting to read this as a bug report about two vendors, which would make it other people’s problem and a matter of time. I don’t think it is.

We have spent two years building agents on the assumption that structured output is a guarantee, because on frontier models it very nearly is. The abstraction layers we wrote — one interface, swap the provider, set the model from config — encode that assumption everywhere. They are genuinely good abstractions. They are also exactly the layers that turn a protocol mismatch into a silent no-op, because their whole job is to make providers look interchangeable, and at the tool-calling boundary providers are not interchangeable at all.

The model tried to call the tool. It said so, clearly, in the only channel it could reach. Everything after that was our system deciding that a sentence is a sentence.

Build the layer that notices. Allowlist what it executes. And make it tell you every single time.

Back to Blog

Related Posts

View All Posts »

Nobody read last night's diff

Fully autonomous coding pipelines are real and shipping. What decides whether yours produces software or defects isn't the model — it's whether your documents say what success looks like, precisely enough to be checked without you.

Wiki or RAG? The real answer is a file format

An LLM-maintained wiki beats retrieval for knowledge that compounds. But a memory written by an agent is only as trustworthy as what it records about itself — and that is exactly the gap the Open Knowledge Format was built to close.