On 8 October 2026, at Gemini at Work, Thomas Kurian announced the Gemini agent: one agent, one API, reachable from Workspace, Microsoft 365, Slack, the command line, and headlessly from inside third-party apps.
Most of the coverage led with the surface area. The line worth reading twice is buried in the architecture: Google says the agent routes each job to the model that fits best, drawing on its own Gemini models and Anthropic's Claude models today, with OpenAI and open-weights models to follow.
Google shipped a product that hands your work to a competitor's model. That is not a feature. That is a statement about where the value moved.
What actually shipped
Stripped of the keynote language, the Gemini agent is four things:
- An orchestrator. You give it an objective, not a prompt. It plans, calls tools, and returns finished work inside the doc, the inbox, or the repo.
- A scheduler. It spins up temporary sub-agents for multi-step jobs that run for hours or days, and it can be triggered by events rather than people.
- An identity system. Persistent "coworker" agents get their own Workspace accounts — email address, calendar, Drive storage — plus cryptographically attested identities, role-based permissions, and audit trails attributed to the agent rather than to whoever created it.
- A policy boundary. An Agent Sandbox with its own network perimeter, and an Agent Gateway that enforces policy on all traffic in, out of, and between agents.
It is in private preview, with general availability expected by the end of October 2026, North America first, included at no extra charge wherever Gemini Enterprise already is.
The model became a backend
Read that component list again, but ignore the word "agent" and substitute "service".
An orchestrator that routes requests to the best-fit upstream. Attested workload identity. Role-based authorisation. Audit trails attributed to the workload. A sandbox with a network boundary. A gateway enforcing policy on north-south and east-west traffic.
That is a service mesh. We have been building this shape for a decade — it is SPIFFE identities, mTLS, an egress gateway, and a routing layer, with a language model in the data path instead of a microservice.
Once you see it that way, the Claude routing stops being a surprise and starts being inevitable:
- A mesh does not care which implementation sits behind an endpoint. That is the entire point of a mesh.
- If the orchestrator owns identity, policy, memory, and audit, then the model is the one component you can swap without touching governance.
- The lock-in moves up the stack. Nobody re-platforms because the summarisation got 3% better. They re-platform when they would have to rebuild identity, permissions, and audit — and that is exactly the layer Google is claiming.
Anthropic gets distribution into Fortune 100 accounts without building an enterprise control plane. Google gets to stop defending every benchmark, because it no longer needs to win every one. Both of those are rational. Neither is good news for anyone whose only product is a model.
The cost line nobody budgets for
Here is where this stops being architecture and starts being an on-call problem.
An agentic workflow is a loop: reason, call a tool, ingest the result, reason again, retry on failure. Industry estimates put agentic systems at 5–30x more model calls per task than a single chatbot query, and McKinsey's work on agentic economics makes the same point from the finance side — token consumption per completed task, not per request, is the number that matters.
Most organisations are still modelling this as a per-seat subscription. It is not a subscription. It is metered compute with a retry loop in front of it, and the retry loop is driven by a non-deterministic planner.
Google clearly knows this, which is why the launch includes Smart Routing between models and real-time spend caps per project, with per-user caps to follow.
Call those what they are. A spend cap is a circuit breaker. Smart Routing is tiered routing by cost and capability. If you have built a payment switch, you have written both of these — you just measured them in basis points and milliseconds instead of tokens.
The difference is the failure mode. A circuit breaker on a payment flow trips on error rate or latency, and the blast radius is bounded by the transaction. A circuit breaker on an agent trips on spend, and the thing it is protecting you from is your own planner deciding a job needs four hundred more tool calls at 2am on a Sunday.
What I would want answered before this goes near production
Not a takedown — the governance story here is the most complete one a hyperscaler has put forward. But the questions I would be asking in a design review:
| Question | Why it matters |
|---|---|
| What is the timeout on a multi-day job? | "Runs for hours or days" and "unbounded" are one config mistake apart. |
| Does a spend cap fail open or closed? | A cap that fails open is a billing incident. One that fails closed is an outage. Pick deliberately. |
| How is a retry attributed in the audit trail? | If the agent's identity signs the action, a retry storm looks like one actor doing a thousand things. |
| Can you pin the model per workload? | Routing to "best fit" is great until a compliance review asks which model saw the PAN. |
| What happens to coworker agents at offboarding? | They have mailboxes and calendars. They are now in scope for every access review you run. |
That last one is not hypothetical. The moment an agent has an email address, it is an identity in your directory, and every control you apply to leavers has to apply to it too.
The takeaway
The headline is that Google's agent calls Claude. The actual news is that Google is betting the orchestration layer is worth more than the model — and it is willing to prove that by routing to a competitor to make the point.
If you are building on any of this, build your abstraction at the same seam Google just drew. Treat the model as a swappable backend. Put identity, policy, routing, and spend control in a layer you own. Measure cost per completed task, not per call.
That advice would have been reasonable last week. This week, the largest vendor in the space just shipped it as a product.





