Study the requests your agent actually receives before deciding which tasks need a stronger model.
Describe each model tier's jobs clearly. This router chooses from the first message and keeps the selected model for the whole thread.
Measure useful outcomes alongside cost, and distinguish a router experiment from a later change to its classifier.
Worth knowing
First-party case study, not an independently repeated benchmark. LangChain reports a 64% lower median thread cost in an experiment that used GLM as its classifier, before switching to Jev. Merged pull requests are a quality proxy; a statistically insignificant difference does not establish equivalent quality. A cheaper-only trial ended early after complaints. Mid-thread routing is proposed work, not the demonstrated behavior.