LOADING

AI Router Lab · Iteration 22

Cheap where it can.Powerful where it must.

Requests pour in, each needing different capabilities.
The router tips every drop to the cheapest model that can handle it.

Requests
…
Saved
…
Quality
…today …
With router / mo
…
Saved / year
…

Illustrative. Each drop is one request. A clouded drop landed in a model too small for it.

How routing works

What you are looking at

Every drop is a request, and no two ask for the same thing: a quick lookup, a careful summary of a long contract, a tricky bug fix, a reply in the right tone, and countless more. Every jar is a model. The router looks at each request and tips it to the cheapest model that handles it well. Only what needs a frontier model gets one.

Why routing pays off

In production traffic, a large share of requests is answered just as well by smaller, cheaper models. Routing each request on its own merits cuts the bill while quality holds. The row below shows what that means for the spend you set.

Your model

The last jar is the model you use today. It stays the default: whenever there is no evidence that a cheaper model is enough, the request goes there. Priority shifts the balance between saving and caution.

When a model is too small

A drop that lands in a model too small for it clouds over. Turn the router off and pick a small model to see how often that happens without routing.

About this illustration

The real router chooses among far more models than five, from many providers, and judges every request by its content. Prices are list prices per 1M tokens. All figures on this page are illustrative.

Some say rubbing a drop over the jars teaches the router your traffic.