Models

The pool behind the lanes.You never have to name one.

Four models, benchmarked on the same graded tasks. The router picks among them per task, and every response names the one it used.

Model catalog

Tier
Capability
4 models
The RouterLane model pool
ImagesEffortRouted from
DeepSeek 4.1 FlashDeepSeekdeepseek-4.1-flash1,048,576262,144Yesoff to max
FastNormalSmartSmart+
100%100 to 290 ms290 to 360
GLM 5.3Z.aiglm-5.3524,288131,072Small imagesoff to max
NormalSmartSmart+Fast on failover
100% (high)410 ms190
GLM 5.3 VisionZ.aiglm-5.3-vision524,288131,072Yesoff to max
NormalSmartSmart+large images
85% (high)320 ms230
GLM 5.3 FlashZ.aiglm-5.3-flash524,288163,840Small imageslow to max
custom rules
69% low, 100% high640 to 960 ms50 to 80

Context and max output in tokens. Effort is RouterLane's scale: off, low, medium, high, max. "Small images" means a request whose images pass about 8K tokens moves to GLM 5.3 Vision. GLM 5.3 Flash is available to custom routing rules on dedicated deployments.

The benchmark.Where the defaults come from.

13 graded tasks: code checked by unit tests, reasoning with checked answers, and a tool-use task. Streamed at each effort, 2026-10-02.

Benchmark, 2026-10-02
ModelEffortPassFirst byte p50Total p50Total p90Tokens/s
DeepSeek 4.1 Flashoff, medium, high, max100%100 to 290 ms2.2 to 4.9 s8 to 16 s290 to 360
GLM 5.3high100%410 ms1.7 s20.5 s190
GLM 5.3max85%920 ms26.9 s71.6 s200
GLM 5.3 Flashlow69%960 ms3.6 s7.5 s50 to 80
GLM 5.3 Flashhigh100%640 ms7.7 s15.1 s50 to 80
GLM 5.3 Visionhigh85%320 ms15.9 s61.9 s230

DeepSeek 4.1 Flash

Passed every task at every effort, with the fastest first byte and 290 to 360 tokens a second. The default for Fast and Normal, and the reader and reviewer in Smart+.

GLM 5.3 at high

Passed every task with the fastest median answer, 1.7 s. Max effort did worse: 85% and a p90 of 71.6 s. So Smart runs high, never max.

Flash and Vision

GLM 5.3 Flash and GLM 5.3 Vision were weaker and slower. Vision takes over only when large images need it.

You pay per tier.Not per model.

The rate is set by the tier you chose, whichever model answered. When Normal moves a hard task up to GLM 5.3, the price is still Normal's.

GET /v1/models lists the tiers your plan includes, because tiers are what you call. See the shape.

Price per million tokens by tier
TierModel IDInputOutputCached input
Fastfast$0.15$0.60$0.03
Normalnormal$0.40$1.60$0.08
Smartsmart$1.00$4.00$0.20
Smart+smart-plus$1.00$4.00$0.20

USD per 1M tokens. Plans and the cost calculator

Pick a lane.Ship the work.

Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.