Tiers

Four lanes.One decision per task.

You send a tier as the model name. The router reads each new task, then picks a model and a reasoning effort inside that tier. The same tier makes different calls for a typo and a deadlock, and that is the point.

The default rules

Each cell is the model and effort a tier uses for a task of that complexity. The kind of task is recorded on every decision, and rules can match on it too.

Default routing rules by tier and task complexity
Tiertrivialsimplemoderatehard
FastDeepSeek 4.1 FlashoffDeepSeek 4.1 FlashoffDeepSeek 4.1 FlashmediumDeepSeek 4.1 Flashmedium
NormalDeepSeek 4.1 FlashoffDeepSeek 4.1 FlashmediumDeepSeek 4.1 FlashmediumGLM 5.3high
SmartDeepSeek 4.1 FlashlowGLM 5.3mediumGLM 5.3highGLM 5.3high
Smart+A job on RouterLane's machineexplore and QA: DeepSeek 4.1 Flash, high. build: GLM 5.3, high

Max effort is on the scale and no default rule uses it: on GLM 5.3 it passed 85% of the benchmark against 100% at high, with a p90 of 71.6 s against 20.5 s.

What each lane is for.And what it costs you in time.

Fast

fast

Questions, lookups, small edits and the quick back-and-forth inside a session. Thinks only when the task needs it. Close calls round down.

Models
DeepSeek 4.1 Flash for every task
Effort
off for trivial and simple, medium for moderate and hard
Latency
first byte p50 100 to 290 ms, total p50 2.2 to 4.9 s
Failover
GLM 5.3
Price
$0.15 in, $0.60 out, $0.03 cached per 1M tokens

Normal

normal

Everyday coding: features, fixes, tests and reviews, with enough reasoning to get it right the first time. Close calls go to the nearest level.

Models
DeepSeek 4.1 Flash, GLM 5.3 for hard tasks
Effort
off for trivial, medium for simple and moderate, high for hard
Latency
DeepSeek 4.1 Flash as on Fast; GLM 5.3 at high: first byte 410 ms, total p50 1.7 s, p90 20.5 s
Failover
between DeepSeek 4.1 Flash and GLM 5.3
Price
$0.40 in, $1.60 out, $0.08 cached per 1M tokens

Smart

smart

Hard debugging, design, plans and multi-step changes. The strongest model and deep reasoning. Close calls round up.

Models
GLM 5.3, DeepSeek 4.1 Flash for trivial turns
Effort
low for trivial, medium for simple, high for moderate and hard
Latency
GLM 5.3 at high: first byte 410 ms, total p50 1.7 s, p90 20.5 s
Failover
between GLM 5.3 and DeepSeek 4.1 Flash
Price
$1.00 in, $4.00 out, $0.20 cached per 1M tokens

Smart+

smart-plus

Whole jobs on a repository: a feature with tests, a migration, a refactor across files. The work runs on RouterLane's machine and one patch comes back.

Models
build on GLM 5.3, explore and QA on DeepSeek 4.1 Flash
Effort
high for every agent
Latency
measured jobs: 24 s for 5 files, 2.3 min for 45 files
Fallback
a plain turn on GLM 5.3 or DeepSeek 4.1 Flash
Price
Smart rates for every token the agents spend

How Smart+ works

Same box.Different lanes.

Sample prompts, decided by each tier's default rules.

decisions from the default rules
Default DeepSeek 4.1 Flash, effort mediumEffort range off to high$0.40 in, $1.60 out per 1M tokens
Prompt
thanks, that fixed it
Read as
trivial chat
Routed toDeepSeek 4.1 Flashoff

Chit-chat gets no thinking time.

Prompt
add pagination to the /orders endpoint and cover it with tests
Read as
moderate feature
Routed toDeepSeek 4.1 Flashmedium

Everyday feature work: the balanced default.

Prompt
two workers deadlock when they refresh the cache together, find the race
Read as
hard debug
Routed toGLM 5.3high

Hard debugging moves up to the larger model at high effort.

What the router reads.Before it picks.

A new task is read by the router's classifier, a small decision model that answers two questions in 200 to 350 ms. The rest comes from the request itself.

Complexity

trivial, simple, moderate or hard. Judged on the work, not the length of the message: a short "yes, do it" inherits the difficulty of the request it answers.

Kind

chat, question, edit, feature, debug, refactor, plan, review, ops or writing. Recorded on every decision and available to rules.

Confidence

When the classifier is unsure, the tier decides which way to round: Smart up, Fast down, Normal to the nearest level.

Images

How many, and how many tokens they add. Large images move a request from GLM 5.3 to GLM 5.3 Vision.

Context size

The estimated length of the request. Past 85% of a model's window, it moves to the 1M-token model.

Conversation

Whether this is a tool result inside a running task, which model the conversation is on, and whether the new task is harder than the last.

The path of one request

Continuation turns skip the classifier entirely. Everything else is read once, then cached for 24 hours by task.

How RouterLane routes a request A request is checked for being the same task as the previous call. If yes, the task's decision is reused with no delay. If no, the task is classified, the tier rules pick a model and effort, and stickiness keeps or upgrades the conversation's model. Constraints for images, context size and model health apply, the call goes upstream with failover before the first byte, and the answer streams back with x-router headers. Request any of 3 formats Same task as before? tool result or retry Classify the task 200 to 350 ms Tier rules model + effort Stickiness keep or move up Reuse the decision 0 ms, cache stays warm Constraints images, context, health Upstream call failover before byte 1 Stream back x-router-* headers no yes

Long sessions stay cheap.And stay on one model.

Continuation turns

When the newest message is a tool result, the request belongs to the task already running. The router reuses that task's model and effort: no classification and no delay. The response says x-router-source: continuation.

Stickiness

A conversation keeps its model, so the upstream prompt cache stays warm: 5K to 22K cached input tokens per turn in real agent loops. When a later task is harder than the one that set the model, the conversation moves up to the stronger one.

Context overflow

When a request passes 85% of the chosen model's window, it moves to DeepSeek 4.1 Flash, which reads 1,048,576 tokens. Long sessions keep going instead of failing on length.

Images

Every tier takes images. When images push a request on GLM 5.3 past about 8K tokens, it moves to GLM 5.3 Vision.

Models fail.Your requests should not.

Failover before the first byte

  • A 429, a 5xx, a timeout or a capability error moves the call to the tier's next model.
  • So does a stream that sends nothing for 45 seconds.
  • Up to three attempts, all before anything reaches your client.
  • Three failures in a row open a circuit, and the model is skipped until it recovers.
  • Congested models are skipped at decision time.

A queue instead of errors

  • When the models are at capacity, requests wait in the router's own queue instead of failing.
  • Fast is served before Normal, Normal before Smart.
  • Tasks already in progress go ahead of new ones.
  • Every 10 seconds of waiting is worth one tier of priority, so nothing starves.
  • A request still waiting after 2 minutes gets a 429 it can retry.

Choose a model yourself,or choose a tier.

Choosing a model yourself compared with choosing a RouterLane tier
Choose a model yourselfChoose a tier
Who picks per callYou, once, in codeThe router, per task
A "thanks" in a hard sessionBilled and answered like the hard workRead as trivial and answered without deep reasoning
Reasoning effortOne setting for everythingSet per task, from off to high
Prompt cacheLost when you switch models by handKept: the conversation stays on one model
A model errors or stallsYour retry codeFailover before the first byte
A better model shipsYou benchmark it and redeployThe tier name in your code stays the same
PriceChanges with the modelOne rate per tier
Seeing what happenedLogging you buildx-router-* headers on every response

Pick a lane.Ship the work.

Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.