The pool behind the lanes.You never have to name one.
Four models, benchmarked on the same graded tasks. The router picks among them per task, and every response names the one it used.
Model catalog
| Images | Effort | Routed from | ||||||
|---|---|---|---|---|---|---|---|---|
DeepSeek 4.1 FlashDeepSeekdeepseek-4.1-flash | 1,048,576 | 262,144 | Yes | off to max | FastNormalSmartSmart+ | 100% | 100 to 290 ms | 290 to 360 |
GLM 5.3Z.aiglm-5.3 | 524,288 | 131,072 | Small images | off to max | NormalSmartSmart+Fast on failover | 100% (high) | 410 ms | 190 |
GLM 5.3 VisionZ.aiglm-5.3-vision | 524,288 | 131,072 | Yes | off to max | NormalSmartSmart+large images | 85% (high) | 320 ms | 230 |
GLM 5.3 FlashZ.aiglm-5.3-flash | 524,288 | 163,840 | Small images | low to max | custom rules | 69% low, 100% high | 640 to 960 ms | 50 to 80 |
| No model in the pool matches both filters. | ||||||||
Context and max output in tokens. Effort is RouterLane's scale: off, low, medium, high, max. "Small images" means a request whose images pass about 8K tokens moves to GLM 5.3 Vision. GLM 5.3 Flash is available to custom routing rules on dedicated deployments.
The benchmark.Where the defaults come from.
13 graded tasks: code checked by unit tests, reasoning with checked answers, and a tool-use task. Streamed at each effort, 2026-10-02.
| Model | Effort | Pass | First byte p50 | Total p50 | Total p90 | Tokens/s |
|---|---|---|---|---|---|---|
| DeepSeek 4.1 Flash | off, medium, high, max | 100% | 100 to 290 ms | 2.2 to 4.9 s | 8 to 16 s | 290 to 360 |
| GLM 5.3 | high | 100% | 410 ms | 1.7 s | 20.5 s | 190 |
| GLM 5.3 | max | 85% | 920 ms | 26.9 s | 71.6 s | 200 |
| GLM 5.3 Flash | low | 69% | 960 ms | 3.6 s | 7.5 s | 50 to 80 |
| GLM 5.3 Flash | high | 100% | 640 ms | 7.7 s | 15.1 s | 50 to 80 |
| GLM 5.3 Vision | high | 85% | 320 ms | 15.9 s | 61.9 s | 230 |
DeepSeek 4.1 Flash
Passed every task at every effort, with the fastest first byte and 290 to 360 tokens a second. The default for Fast and Normal, and the reader and reviewer in Smart+.
GLM 5.3 at high
Passed every task with the fastest median answer, 1.7 s. Max effort did worse: 85% and a p90 of 71.6 s. So Smart runs high, never max.
Flash and Vision
GLM 5.3 Flash and GLM 5.3 Vision were weaker and slower. Vision takes over only when large images need it.
You pay per tier.Not per model.
The rate is set by the tier you chose, whichever model answered. When Normal moves a hard task up to GLM 5.3, the price is still Normal's.
GET /v1/models lists the tiers your plan includes, because tiers are what you call. See the shape.
| Tier | Model ID | Input | Output | Cached input |
|---|---|---|---|---|
| Fast | fast | $0.15 | $0.60 | $0.03 |
| Normal | normal | $0.40 | $1.60 | $0.08 |
| Smart | smart | $1.00 | $4.00 | $0.20 |
| Smart+ | smart-plus | $1.00 | $4.00 | $0.20 |
USD per 1M tokens. Plans and the cost calculator
Pick a lane.Ship the work.
Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.