LLM API for coding agents

Pick a lane.We pick the model.

Choose Fast, Normal, Smart or Smart+. A small decision model reads every new task in 200 to 350 ms and sends it to the model and reasoning effort that fit, in the API format your client already speaks.

Anthropic Messages · OpenAI Chat Completions · OpenAI Responses

Works with
  • Claude Code
  • Codex
  • pi
  • OpenAI SDK
  • Anthropic SDK
  • curl

Four lanes.Watch the router pick.

Each tier has its own rules. The router reads the task, then picks the model and the effort inside the tier you chose. Pick a tier to see how it handles three prompts.

decisions from the default rules
Default DeepSeek 4.1 Flash, effort mediumEffort range off to high$0.40 in, $1.60 out per 1M tokens
Prompt
thanks, that fixed it
Read as
trivial chat
Routed toDeepSeek 4.1 Flashoff

Chit-chat gets no thinking time.

Prompt
add pagination to the /orders endpoint and cover it with tests
Read as
moderate feature
Routed toDeepSeek 4.1 Flashmedium

Everyday feature work: the balanced default.

Prompt
two workers deadlock when they refresh the cache together, find the race
Read as
hard debug
Routed toGLM 5.3high

Hard debugging moves up to the larger model at high effort.

One price per lane.Whichever model answers.

You pay the tier's rate, not the model's. Input served from the prompt cache is billed at the cached rate.

Fast

fast

Lowest latency. Quick answers and small edits, thinking only when needed.

Default
DeepSeek 4.1 Flash
Effort
off or medium
Plans
Starter, Pro, Enterprise
$0.15 / $0.60in / out per 1M tokens

Normal

normal

Balanced. Everyday coding with enough reasoning to get it right.

Default
DeepSeek 4.1 Flash
Effort
off to high
Plans
Starter, Pro, Enterprise
$0.40 / $1.60in / out per 1M tokens

Smart+

smart-plus

Long jobs run on RouterLane's machine. One patch comes back.

Default
GLM 5.3 builds
Effort
high
Plans
Enterprise
Smart ratesfor every token the agents spend

Cached input: Fast $0.03, Normal $0.08, Smart and Smart+ $0.20 per 1M tokens. Full pricing

Every task gets read.Then it gets a lane.

The decision happens once per task, not once per call. Tool loops inside a task ride on it for free.

  1. Read

    A small decision model labels the new task: how hard it is, from trivial to hard, and what kind of work it is, from debug to plan.

    200 to 350 ms
  2. Route

    The tier's rules turn that, plus images and context size, into a model and a reasoning effort. Close calls round the tier's way: Smart up, Fast down.

    model + effort
  3. Stick

    Tool results inside the task reuse its decision with no delay, and the conversation stays on one model so the prompt cache stays warm.

    5K to 22K cached tokens a turn
  4. Deliver

    The call goes out in your client's format. A model that errors or stalls before the first byte is swapped for the tier's next one.

    x-router-model

How routing works, in depth

Smart+

Stop shipping files.Ship the repo once.

Smart+ packs your repository, uploads it once, and runs the whole job on RouterLane's machine. You get one patch back, inside the same agent turn.

  • Tracked and new files only. No history, no ignored files, no node_modules.
  • Setup, explore, build, test and an independent QA review, all local to the copy.
  • One patch, applied with git apply. Nothing is committed.
  • A spend ceiling on every job, $1.50 by default.

Counts from a measured 5-file job, 2026-10-03.

Measured.Not guessed.

The defaults come from a benchmark of 13 graded tasks per model and effort, and from real agent sessions.

200 to 350msfor the decision model to read a new task
0msrouting delay on continuation turns inside a task
100 to 290msfirst byte p50 on DeepSeek 4.1 Flash, the Fast default
5K to 22Kcached input tokens per turn in real agent loops
3API formats: Anthropic Messages, OpenAI Chat Completions and OpenAI Responses
$0.035for a measured Smart+ job: 5 files, tests passed, patch applied in 24 s

Two variables.That is the integration.

  1. Request access

    Tell us what you are building. We email you a key that starts with rl_.

  2. Point your client at RouterLane

    Base URL https://api.routerlane.com, or /v1 for OpenAI clients.

  3. Use a tier as the model

    fast, normal, smart or smart-plus. Claude-style names work too.

shell
export ANTHROPIC_BASE_URL=https://api.routerlane.com
export ANTHROPIC_AUTH_TOKEN=$ROUTERLANE_API_KEY
export ANTHROPIC_MODEL='smart[1m]'
export ANTHROPIC_DEFAULT_OPUS_MODEL='smart[1m]'
export ANTHROPIC_DEFAULT_SONNET_MODEL='normal[1m]'
export ANTHROPIC_DEFAULT_HAIKU_MODEL=fast
claude

Questions,answered.

Which model answers my request?

One from the pool your tier routes to: DeepSeek 4.1 Flash, GLM 5.3 or GLM 5.3 Vision. Every response names it in the x-router-model header, next to the effort and the reason for the choice.

Can I choose the model myself?

You choose the tier. The router chooses the model and the effort inside it, per task. If you want the same behavior every time, pick the tier whose rules match: Fast sends every task to DeepSeek 4.1 Flash, and Smart sends anything harder than trivial to GLM 5.3.

What happens to the thinking or effort setting my client sends?

It is replaced by the router's choice. For Anthropic-format requests the effort becomes a thinking budget of 1K, 4K, 12K or 24K tokens, or thinking off. For OpenAI formats it becomes the model's own reasoning effort value.

Does it work with Claude Code, Codex and pi?

Yes, unchanged. Claude Code needs two environment variables, Codex a provider block in its config file, and pi an entry in its models file. Claude-style model names map by family: haiku to Fast, sonnet to Normal, opus to Smart. See Integrations.

What if a model is slow or down?

Before the first byte reaches you, the router retries on the tier's next model after a 429, a 5xx, a timeout, a capability error, or a stream that stays silent for 45 seconds. Three failures in a row open a circuit and the model is skipped. When every model is busy, requests wait in the router's queue instead of failing, with Fast served first and tasks in progress ahead of new ones.

Will a long agent session lose its prompt cache?

No. A conversation stays on one model unless a later task is harder, and then it moves up. Tool-result turns reuse the task's decision. In real agent loops that kept 5K to 22K input tokens per turn in the cache, billed at the cached rate.

How am I billed?

Per token, at your tier's rate, metered per request and billed monthly. There is no subscription fee. Identical requests repeated within 10 minutes are served from cache for free, and you never pay for the routing decision. See Pricing.

Do you store my prompts and code?

Request metadata is kept for billing. Transcripts are kept for 30 days by default, then the bodies are deleted. Enterprise deployments choose: transcripts on or off, raw parameters on or off, and the retention window. Smart+ workspaces are deleted after 48 hours. See Security.

What is Smart+, and when should I use it?

Smart+ is for whole jobs, not single answers: a feature with tests, a migration, a refactor across files. Your agent uploads the repository once, the job runs on RouterLane's machine with its own tests and a QA review, and one patch comes back. It is on Enterprise plans.

How do I get a key?

Keys are issued on request. Request access with a few lines about your use, and the key arrives by email. It is shown once, so store it in your secret manager.

Pick a lane.Ship the work.

Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.