Back to BlogAI Features & LLMs

LLM Model Routing: Send Cheap Queries to Cheap Models

Rupak Amin

Founder & Lead Engineer, RAITHub

12 min read

An LLM model router classifies each request and sends it to the least expensive model that handles it well. On Anthropic's list prices, sending 70% of requests to Claude Haiku 4.5, 20% to Sonnet 5.5 and 10% to Opus 5.5 costs about 62% less than sending everything to Opus. The router only pays off if it is calibrated against an evaluation set, so misroutes stay rare.

This post owns one topic: the design of the router itself. The wider list of cost levers, such as output caps, prompt caching and the Batch API, is in how to reduce OpenAI and LLM API costs. The reference design here is PadhAI, the AI tutoring platform RAITHub built, whose 70/20/10 model router classifies query complexity and sends each query to a lightweight, mid-tier or premium model.

What is LLM model routing?

A decision step in front of your model calls. It looks at a request, predicts how hard it is, and picks a model tier for it. The user never sees the router; they see a fast answer to an easy question and a careful answer to a hard one.

It works because traffic is uneven. In most products a large share of requests are short lookups, reformatting or classification, and a small share need multi-step reasoning. Paying premium rates for all of it is often a large, avoidable share of the bill. Anthropic's own pricing page gives the same advice in its cost section: use Haiku for simple tasks, Sonnet for most production workloads, and Opus for the most complex reasoning. A router is that advice turned into code.

How much can a model router save? A cost calculator

The blended cost per request is the sum of each tier's cost weighted by its share of traffic:

blended = share_light × cost_light + share_mid × cost_mid + share_premium × cost_premium
cost_tier = input_tokens × input_price + output_tokens × output_price   (prices per token)

A worked example with a request of 2,000 input tokens and 500 output tokens, at Anthropic's standard list prices checked on 29 September 2026: Claude Haiku 4.5 at $1 input and $5 output per million tokens, Claude Sonnet 5.5 at $2 and $10, and Claude Opus 5.5 at $4 and $20.

TierModelCost per requestShare (70/20/10)Weighted cost
LightweightClaude Haiku 4.5$0.004570%$0.00315
Mid-tierClaude Sonnet 5.5$0.009020%$0.00180
PremiumClaude Opus 5.5$0.018010%$0.00180
Blended, routedAll threen/a100%$0.00675
No routingEverything on Opus 5.5$0.0180100%$0.01800
No routingEverything on Sonnet 5.5$0.0090100%$0.00900

Per million requests, that is $6,750 routed against $18,000 on the premium tier alone, a 62.5% saving, or 25% against sending everything to the mid tier. These are illustrative numbers from list prices, not PadhAI figures; RAITHub does not publish PadhAI's costs. Two cautions when you plug in your own traffic. First, measure the real split instead of assuming 70/20/10; the split is a design target, not a law. Second, compare tokens per request, not just price per token: Anthropic's pricing page notes that Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text, which widens the gap between Haiku 4.5 and the newer tiers in practice.

Which router design should you use: rules, classifier or cascade?

Three designs cover almost every case. Most production routers combine the first two and add the third for a few request types.

DesignHow it decidesAdded costFits whenMain risk
RulesFeature, input length, user plan, keywords, attached filesNoneRequest types are known from the product, such as "summarise" vs "explain"Brittle for free-text questions
ClassifierA small model or trained classifier predicts the tierOne small call, or a local modelFree-text traffic of mixed difficultyMisroutes if not calibrated
CascadeTry the light model, check the output, escalate on failureFailed attempts are paid twiceOutputs can be checked by code: schema, tests, math, citationsLatency on escalated requests

Rules come first because they are free and explainable. The feature a request comes from often tells you most of what you need: an autocomplete suggestion never needs the premium tier, and a contract analysis rarely suits the lightweight one.

A classifier handles open questions. A cheap model, prompted to return only a tier label, or a small trained classifier on labelled examples, reads the request and predicts difficulty. Keep its output constrained to a fixed set of labels, and default to the mid tier when it is unsure.

A cascade suits checkable output. If code can verify the answer, such as valid JSON against a schema, a unit test passing, a citation that exists in the retrieved passages or a math result checked by a computer algebra system, then try the lightweight model first and escalate only on failure.

When does a cascade cost more than it saves?

When too many requests escalate. Every escalated request pays for the failed cheap attempt as well as the successful expensive one, so the expected cost is:

expected = cost_light + escalation_rate × cost_next_tier

With the example prices above, a Haiku-then-Sonnet cascade costs $0.0045 + r × $0.0090. It stays below going straight to Sonnet ($0.0090) as long as the escalation rate r is under 50%. At a 20% escalation rate it averages $0.0063; at 60% it averages $0.0099 and you are paying more than skipping the cascade, plus the extra latency. Log the escalation rate per request type and move a type to direct routing once its rate climbs.

How do you build a model router in TypeScript?

Keep the tier table in one place, route by rules first, fall back to a classifier, and log every decision. The sketch below is illustrative, not PadhAI's code; classifyTier, callModel, passesCheck and logRoute are your own functions.

type Tier = 'light' | 'mid' | 'premium'

// Model IDs and prices live in one table, so a price change is a one-line edit.
// Confirm the exact model IDs in your provider's docs before use.
const TIERS: Record<Tier, { model: string; inPerM: number; outPerM: number }> = {
  light:   { model: 'claude-haiku-4-5', inPerM: 1, outPerM: 5 },
  mid:     { model: 'claude-sonnet-5-5', inPerM: 2, outPerM: 10 },
  premium: { model: 'claude-opus-5-5', inPerM: 4, outPerM: 20 },
}
const NEXT: Record<Tier, Tier | null> = { light: 'mid', mid: 'premium', premium: null }

type Req = { feature: string; text: string; checkable: boolean }
type Result = { text: string; inputTokens: number; outputTokens: number }

declare function classifyTier(text: string): Promise<Tier | 'unsure'>
declare function callModel(model: string, text: string): Promise<Result>
declare function passesCheck(req: Req, out: Result): boolean
declare function logRoute(row: Record<string, unknown>): Promise<void>

function byRule(req: Req): Tier | null {
  if (req.feature === 'autocomplete' || req.feature === 'tagging') return 'light'
  if (req.feature === 'contract-review') return 'premium'
  return null
}

export async function route(req: Req): Promise<string> {
  const ruled = byRule(req)
  const guess = ruled ?? (await classifyTier(req.text))
  let tier: Tier = guess === 'unsure' ? 'mid' : guess // unsure defaults to the middle, not the bottom

  for (let attempt = 0; ; attempt++) {
    const { model, inPerM, outPerM } = TIERS[tier]
    const out = await callModel(model, req.text)
    const costUsd = (out.inputTokens * inPerM + out.outputTokens * outPerM) / 1_000_000
    const ok = !req.checkable || passesCheck(req, out)
    await logRoute({ feature: req.feature, decidedBy: ruled ? 'rule' : 'classifier', tier, attempt, ok, costUsd })

    const next = NEXT[tier]
    if (ok || !next || attempt >= 1) return out.text // at most one escalation
    tier = next
  }
}

Four choices in that sketch are deliberate. Unsure requests go to the mid tier, because under-routing a hard question costs trust while over-routing an easy one costs only money. Escalation happens at most once, so a bad request cannot climb every tier. Only checkable requests cascade; for free text, a second opinion is not a check. And every decision is logged with who made it, rule or classifier, so you can see which one misroutes.

How do you calibrate a router so it doesn't hurt quality?

Label an evaluation set with the lowest-cost tier that passes each item, then measure how often the router agrees. That turns "is the router good?" into a table you can review.

  1. Collect 200 to 500 real requests across every feature, including the hard and odd ones.
  2. Run each through every tier and grade the outputs with your normal evaluation method.
  3. Label each item with the lowest tier whose answer passes.
  4. Run the router on the same items and build a confusion table: routed tier against required tier.
  5. Set thresholds from the table. Under-routes, where the router picked a tier below the one required, are the quality risk; keep them to a rate you have agreed with the product owner. Over-routes are only cost.
  6. Re-run on every change to prompts, models or the classifier, in CI, the same way as any other regression test.

How to build and grade that set is covered in how to test LLM features. When a provider ships a new model, the evaluation set tells you in an afternoon whether it can take a larger share of traffic at the lower tier.

What goes wrong with model routers in production?

  • Silent drift. A new feature launches, its requests are harder than average, and the classifier routes them low. Track pass rate per feature and tier, not only overall.
  • The classifier becomes the cost. If the classifier call is nearly as large as the answer, it eats the saving. Keep its prompt short, and use rules wherever the feature already tells you the answer.
  • Tier outages. Give each tier a fallback model, ideally on another provider, and decide in advance whether a failure falls up or down a tier.
  • Different formats per model. Models vary in how they follow output formats. Validate every tier's output with the same schema check.
  • Price tables go stale. Keep prices in code with a "checked on" date, and review them when the pricing page changes.

How does PadhAI's 70/20/10 router work?

PadhAI's router classifies each query by complexity and sends it to a lightweight, mid-tier or premium model; the name describes the intended split of traffic. It is how PadhAI keeps AI cost per student viable across a platform of 11 services, 9 in Node/TypeScript and 2 in Python. In a tutor, the hard 10% is where accuracy matters most, which is why the tutor also checks mathematical results with code rather than trusting any tier. The rest of PadhAI's design is in how to build an AI tutor. RAITHub does not publish PadhAI's traffic split in practice, cost per student or classifier details.

Why RAITHub for LLM routing

  • Designed one in production. RAITHub built PadhAI's 70/20/10 model router alongside RAG and math verification.
  • Measure first. Every routing change is judged against an evaluation set and a cost log, not a hunch.
  • Cost habits across the stack. This site's own Neon database cost fix and rate limiting without Redis came from the same practice of measuring before optimising.
  • Fixed scope, your IP. A free 15-minute technical audit, then a fixed written quote. You own the code and IP; an NDA is standard.

When you don't need us

  • Your traffic is small. If the whole AI bill is less than a day of engineering a month, pick one mid-tier model and revisit later.
  • All your requests are alike. A single feature with uniform difficulty needs a model choice, not a router.
  • You want a hosted gateway. Several API gateways offer routing and fallbacks as configuration; if one fits your stack, use it and keep the evaluation set.

Routing sits in the backend layer of an AI product, which is covered by the API and backend development service. To model your own blended cost, book the free 15-minute technical audit and bring one month of usage per feature.

Last reviewed: 29 September 2026. Prices checked on 29 September 2026.

Frequently asked questions

What is an LLM router?

A step before each model call that predicts how hard a request is and sends it to a matching model tier, so easy requests use a low-cost model and hard ones use a stronger one.

How much does model routing save?

It depends on your traffic split. With a 70/20/10 split across Claude Haiku 4.5, Sonnet 5.5 and Opus 5.5 at list prices, an illustrative request costs 62.5% less than on Opus alone and 25% less than on Sonnet alone.

What does 70/20/10 routing mean?

About 70% of requests go to a lightweight model, 20% to a mid-tier model and 10% to a premium model. PadhAI's router uses that split as its design target, classifying each query by complexity.

Should I use a classifier or a cascade?

Use a classifier for free-text traffic of mixed difficulty. Use a cascade only when code can check the output, and only while fewer than about half of requests escalate, or the failed attempts cost more than they save.

Does routing to smaller models hurt quality?

It can, if the router is not calibrated. Label an evaluation set with the lowest-cost tier that passes each item, measure under-routes, and send unsure requests to the mid tier.

How do I calculate blended LLM cost?

Multiply each tier's cost per request by its share of traffic and add them up. Cost per request is input tokens times the input price plus output tokens times the output price.

LLM model routingmodel routerLLM API cost calculatorLLM costmodel cascadeAI architecture

Ready to discuss your project?

Book a free 15-minute technical audit with our engineering team.