AI Router Lab · Iteration 22
Requests pour in, each needing different capabilities.
The router tips every drop to the cheapest model that can handle it.
Illustrative. Each drop is one request. A clouded drop landed in a model too small for it.
Every drop is a request, and no two ask for the same thing: a quick lookup, a careful summary of a long contract, a tricky bug fix, a reply in the right tone, and countless more. Every jar is a model. The router looks at each request and tips it to the cheapest model that handles it well. Only what needs a frontier model gets one.
In production traffic, a large share of requests is answered just as well by smaller, cheaper models. Routing each request on its own merits cuts the bill while quality holds. The row below shows what that means for the spend you set.
The last jar is the model you use today. It stays the default: whenever there is no evidence that a cheaper model is enough, the request goes there. Priority shifts the balance between saving and caution.
A drop that lands in a model too small for it clouds over. Turn the router off and pick a small model to see how often that happens without routing.
The real router chooses among far more models than five, from many providers, and judges every request by its content. Prices are list prices per 1M tokens. All figures on this page are illustrative.
Some say rubbing a drop over the jars teaches the router your traffic.