← All Quick Takes
Architecture22 September 2026

What do you accept when you choose a model routing provider?

A diagram of what choosing a model routing provider involves: intent capture, behavior and quality evaluation, pinned interchangeable models, harness rules, injection defense, and portability around a central router, with a harness pipeline running from gateway telemetry to work category, router, selected model, and consistent output.

I previously asked whether AI token cost will push an enterprise to adopt model routing. When I asked the question, I focused on the Gateway needed to evaluate token traffic prior to answering the question of what router and models should be considered. My recent difficulties moving between frontier development platforms and getting the same quality of output have reminded me how little benchmarks told me about which model would do my work well. Once models reach a level of intelligence that is acceptable, their understanding of your intent, how meticulous they are with checking their work and general focus on the assigned tasks are critical factors. This means there is a behavioral aspect as well as an intelligence comparison that needs to occur to select an appropriate model.

When I went looking for information provided by model routing vendors I saw they divide themselves into two categories. Some act as load balancers, and some classify the request and send it to a model tier chosen in advance. Currently, I have not found a vendor, research group, or organization that publishes a granular, automated capability matrix broken down by distinct enterprise task categories.

Model routers are lacking the evaluation subtleties needed to route work to complementary models and keep the output consistent. Even so, my answer is not to skip them. A model router plays a part in AI operational resiliency and portability by forcing an organization to address vendor model differences. The answer may be the addition of behavioral and intent rules to an AI harness. Another answer can be to break work into sections narrow enough to diminish the impact of model differences when producing a desired artifact.

On August 21, 2026, Anthropic published its AI-Native SDLC playbook. Anthropic described an intent.md file that the person with the idea writes with Claude. It records what is wanted, why, and under which constraints, and the product owner reviews it before it is committed. I find the naming choice to be interesting. I will not assume from the name that Anthropic agrees that intent across model types is a problem. I do see this as them recognizing that understanding the original pain point matters or what I like to refer to as the "ask behind the ask" is something that needs to be captured. Understanding the intent across numerous prompts is a great guardrail for producing something that meets a user's need. The intent file should be part of the version control and supplied to whatever model is chosen so they all start with the same intent.

Adding intent and behavioral rules to the AI harness may not be enough either. During testing you may need to evaluate and pin the models you have identified as interchangeable. This means additional work to look at the information captured by your AI Gateway, define categories of work and figure out if the model router you are considering can support additional selection criteria.

Another consideration worth mentioning, that deserves its own research and post, is the topic of injection defense. A model behind an API does not come with a harness and prompt injection defense is part of that layer. You cannot assume your model router provides this security, so the design work around the implementation needs to account for it.

Moving between frontier models has expanded my evaluation criteria of what makes model routing successful. The information I have found has also made me skeptical of how well you can predict the outcome of work if you swap models based on either cost or a capability metric. The safer approach may be expanding the role of your AI harness and classifying general categories of work. Then based on those categories let the subject matter experts make the model decision and decide if your budget can support it.

Research

Under unified evaluation, several routing methods, including the commercial router OpenRouter, did not beat simply using the best single model; OpenRouter scored 24.7% below it. On queries only one to three models could answer, the best routers got 23–25% right.

LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing, arXiv:2601.07206, 2026

If you have moved work between frontier models, what changed in the output that a benchmark would not have predicted?

Duane Grey

Written by Duane Grey

AI Strategy & Implementation

Independent AI consultant helping companies cut through hype and deploy systems that produce real results.

Considering an AI initiative?

Let's name where it fits, then build it.

Start a Conversation