Fast
fastLowest latency. Quick answers and small edits, thinking only when needed.
- Default
- DeepSeek 4.1 Flash
- Effort
- off or medium
- Plans
- Starter, Pro, Enterprise
Choose Fast, Normal, Smart or Smart+. A small decision model reads every new task in 200 to 350 ms and sends it to the model and reasoning effort that fit, in the API format your client already speaks.
$ curl -i https://api.routerlane.com/v1/messages \ -H "x-api-key: $ROUTERLANE_API_KEY" \ -H "anthropic-version: 2023-06-01" \ -H "content-type: application/json" \ -d '{"model": "smart", "max_tokens": 4096, "messages": [{"role": "user", "content": "Find the deadlock in cache.py"}]}' HTTP/2 200 x-router-tier: smart x-router-model: glm-5.3 x-router-effort: high x-router-source: rule x-router-request-id: 4f1c9a07d2e86b13
Each tier has its own rules. The router reads the task, then picks the model and the effort inside the tier you chose. Pick a tier to see how it handles three prompts.
A mechanical edit. No thinking time, so the first byte arrives fastest.
One fact, one answer. Fast replies straight away.
Real investigation. Fast stays on its quickest model and turns on medium reasoning.
Chit-chat gets no thinking time.
Everyday feature work: the balanced default.
Hard debugging moves up to the larger model at high effort.
Smart still answers a thank-you quickly.
A small task on the strongest model, with medium reasoning.
Design work gets the strongest model at high effort. When the call is close, Smart rounds up.
Explore and QA on DeepSeek 4.1 Flash, build on GLM 5.3. One upload, one patch back.
The project's own tests run on the copy before the patch is sent.
The client declared no shell tool, so nothing can be uploaded. The turn runs as a Smart turn.
You pay the tier's rate, not the model's. Input served from the prompt cache is billed at the cached rate.
Lowest latency. Quick answers and small edits, thinking only when needed.
Balanced. Everyday coding with enough reasoning to get it right.
The strongest model and deep reasoning for hard, multi-step work.
Long jobs run on RouterLane's machine. One patch comes back.
Cached input: Fast $0.03, Normal $0.08, Smart and Smart+ $0.20 per 1M tokens. Full pricing
The decision happens once per task, not once per call. Tool loops inside a task ride on it for free.
A small decision model labels the new task: how hard it is, from trivial to hard, and what kind of work it is, from debug to plan.
200 to 350 msThe tier's rules turn that, plus images and context size, into a model and a reasoning effort. Close calls round the tier's way: Smart up, Fast down.
model + effortTool results inside the task reuse its decision with no delay, and the conversation stays on one model so the prompt cache stays warm.
5K to 22K cached tokens a turnThe call goes out in your client's format. A model that errors or stalls before the first byte is swapped for the tier's next one.
x-router-modelSmart+ packs your repository, uploads it once, and runs the whole job on RouterLane's machine. You get one patch back, inside the same agent turn.
node_modules.git apply. Nothing is committed.> add a --json flag to the report command, with a test Smart+ is on, so this runs on RouterLane's machine. ● Bash pack and upload the repository 5 files, 435 B setup copy ready, baseline commit, tests found explore 4 calls plan written build 5 calls 2 files changed, +26 -1 test passed qa 3 calls verdict pass deliver one patch ● Bash download the patch and apply it applied cleanly, nothing staged 24 s · 25K tokens · $0.035
Counts from a measured 5-file job, 2026-10-03.
Enterprise deployments come with their own console: the full transcript of every agent session, answers about your traffic, and filters that can stop a request before it leaves.
Every session as user, session, thread, turn and step, with the routing decision on each call. Full-text search and filters.
Ask a question about chosen users and sessions and get an answer that cites the turns. Sessions summarize themselves.
Yes or no rules that read prompts, replies, reasoning and tool calls. Log what matches, or block it on the way in or out.
The defaults come from a benchmark of 13 graded tasks per model and effort, and from real agent sessions.
Tell us what you are building. We email you a key that starts with rl_.
Base URL https://api.routerlane.com, or /v1 for OpenAI clients.
fast, normal, smart or smart-plus. Claude-style names work too.
export ANTHROPIC_BASE_URL=https://api.routerlane.com
export ANTHROPIC_AUTH_TOKEN=$ROUTERLANE_API_KEY
export ANTHROPIC_MODEL='smart[1m]'
export ANTHROPIC_DEFAULT_OPUS_MODEL='smart[1m]'
export ANTHROPIC_DEFAULT_SONNET_MODEL='normal[1m]'
export ANTHROPIC_DEFAULT_HAIKU_MODEL=fast
claudemodel = "normal"
model_provider = "routerlane"
[model_providers.routerlane]
name = "RouterLane"
base_url = "https://api.routerlane.com/v1"
env_key = "ROUTERLANE_API_KEY"
wire_api = "responses"curl https://api.routerlane.com/v1/chat/completions \
-H "Authorization: Bearer $ROUTERLANE_API_KEY" \
-H "content-type: application/json" \
-d '{"model": "fast", "messages": [{"role": "user", "content": "hello"}]}'import os
from openai import OpenAI
client = OpenAI(base_url="https://api.routerlane.com/v1",
api_key=os.environ["ROUTERLANE_API_KEY"])
r = client.chat.completions.create(
model="normal",
messages=[{"role": "user", "content": "Explain this stack trace."}],
)
print(r.choices[0].message.content)One from the pool your tier routes to: DeepSeek 4.1 Flash, GLM 5.3 or GLM 5.3 Vision. Every response names it in the x-router-model header, next to the effort and the reason for the choice.
You choose the tier. The router chooses the model and the effort inside it, per task. If you want the same behavior every time, pick the tier whose rules match: Fast sends every task to DeepSeek 4.1 Flash, and Smart sends anything harder than trivial to GLM 5.3.
It is replaced by the router's choice. For Anthropic-format requests the effort becomes a thinking budget of 1K, 4K, 12K or 24K tokens, or thinking off. For OpenAI formats it becomes the model's own reasoning effort value.
Yes, unchanged. Claude Code needs two environment variables, Codex a provider block in its config file, and pi an entry in its models file. Claude-style model names map by family: haiku to Fast, sonnet to Normal, opus to Smart. See Integrations.
Before the first byte reaches you, the router retries on the tier's next model after a 429, a 5xx, a timeout, a capability error, or a stream that stays silent for 45 seconds. Three failures in a row open a circuit and the model is skipped. When every model is busy, requests wait in the router's queue instead of failing, with Fast served first and tasks in progress ahead of new ones.
No. A conversation stays on one model unless a later task is harder, and then it moves up. Tool-result turns reuse the task's decision. In real agent loops that kept 5K to 22K input tokens per turn in the cache, billed at the cached rate.
Per token, at your tier's rate, metered per request and billed monthly. There is no subscription fee. Identical requests repeated within 10 minutes are served from cache for free, and you never pay for the routing decision. See Pricing.
Request metadata is kept for billing. Transcripts are kept for 30 days by default, then the bodies are deleted. Enterprise deployments choose: transcripts on or off, raw parameters on or off, and the retention window. Smart+ workspaces are deleted after 48 hours. See Security.
Smart+ is for whole jobs, not single answers: a feature with tests, a migration, a refactor across files. Your agent uploads the repository once, the job runs on RouterLane's machine with its own tests and a QA review, and one patch comes back. It is on Enterprise plans.
Keys are issued on request. Request access with a few lines about your use, and the key arrives by email. It is shown once, so store it in your secret manager.
Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.