Model Routing

Most applications pick one model and send every request to it. Model routing looks at each request first, then decides which model should handle that one. A simple question goes to a small, cheap model. Something harder goes to a bigger, more expensive one. The point is that a single application usually has both kinds of request, and paying the expensive price for all of them is a choice rather than a requirement.

In this guide
  1. Route, or cascade?
  2. What decides the route
  3. Where this is worth doing
  4. When one model is the better answer
  5. Be careful with the published numbers
  6. FAQ

Route, or cascade?

There are two ways to make that decision, and they fail differently.

Routing — decide first, one call Request Classify it Small model Large model Cascading — try cheap, escalate on a miss Request Small model Good enough? Large model yes → done no → cheap call wasted
Routing commits before it has seen any answer. Cascading sees a real answer first, but the cheap attempt is wasted whenever it escalates.

Routing classifies the request up front and sends it to one model. That's one call, so there's no duplicated cost — but the decision is made without ever seeing an answer. The classifier is guessing how hard the request will turn out to be, and it can guess wrong in a way nothing downstream catches.

Cascading sends the request to the cheap model first, then checks whether the answer is good enough. If it is, you're done at the cheap price. If it isn't, the request escalates to the bigger model. The judgment is better, because it's made against a real answer rather than a prediction. The cost is that every escalation has already paid for the cheap call, and waits for it to finish before the real attempt starts.

What decides the route

The cheapest approach is rules you write yourself: send anything under a length threshold to the small model, or route by which feature called the API. Crude, free to run, and often enough.

Past that, a small classifier model looks at the request and predicts which tier it needs. It costs a little on every request and adds a step before anything starts generating, so it has to save more than it spends. Cascading replaces the prediction with a quality check on the cheap model's actual output — sometimes a score the model reports, sometimes a separate LLM-as-judge pass, sometimes just a rule like "did it refuse, or return an empty result?"

Where this is worth doing

A support assistant with a wide mix of requests. "What are your opening hours?" and "work out why these three invoices disagree" arrive through the same box. The first needs almost nothing; the second needs real reasoning. One model for both means overpaying on every easy question.

A high-volume classification or extraction job. Tagging a million documents is exactly where a small model earns its place, with a cascade catching the few that come back low-confidence.

Routing by capability rather than cost. Not every split is cheap-versus-expensive. Sending code questions to a model that's strong at code, and long documents to one with a bigger context window, is routing too — the deciding factor is fit, not price.

When one model is the better answer

When your requests are all roughly the same difficulty, there's nothing to route — a single well-chosen model is simpler and has no classifier to maintain. Skip it too when volume is low enough that the whole model bill is small, because the engineering and the ongoing evaluation cost more than they save.

The real ongoing cost is that routing adds a second thing to evaluate. You now have to know not just whether the models are good, but whether the routing decision is good — and a router that quietly sends too much to the cheap tier degrades quality in a way that's easy to miss, because nothing errors.

Be careful with the published numbers

Savings figures for routing circulate at very different magnitudes, and the gap is usually about which result is being quoted rather than anyone getting it wrong. The RouteLLM paper claims cost reductions of over two times in certain cases without compromising response quality. The same authors' accompanying write-up reports over 85% cost reduction while keeping 95% of GPT-4's performance — but that's one specific benchmark, with one specific pair of models, sending 14% of calls to the strong one.

Both are real. The second is narrower than it looks when it's quoted on its own, which is how it usually travels. Neither predicts your application, because what you save depends on your mix of requests, how far apart your two models are in price, and how accurate the routing decision is. A workload where almost everything is hard saves almost nothing. Measure it on your own traffic before counting the saving.

FAQ

Doesn't the router itself cost money on every request?

Yes, and it has to be cheap enough that the saving survives it. A rules-based router costs essentially nothing. A small classifier model costs a real amount per request, plus the delay of running it before anything else starts — if your two tiers aren't far apart in price, that overhead can eat the entire benefit.

Does cascading make responses slower?

Only the ones that escalate, but those get noticeably slower, because they wait for the cheap model to finish and be judged before the expensive one even starts. Average latency can improve while the slowest requests get worse — so judge it on the slow tail, not the average.

Practice interview questions on model routing →