API reference
Base host https://api.routerlane.com. Requests and responses follow the Anthropic and OpenAI formats; RouterLane adds tier names and decision headers.
| Endpoint | Format | Used by |
|---|---|---|
POST /v1/messages | Anthropic Messages | Claude Code, pi, Anthropic SDKs |
POST /v1/messages/count_tokens | Anthropic token counting | Claude Code, Anthropic SDKs |
POST /v1/chat/completions | OpenAI Chat Completions | OpenAI SDKs, Vercel AI SDK, LangChain |
POST /v1/responses | OpenAI Responses | Codex, OpenAI SDKs |
GET /v1/models | Model list, readable by both | Any client |
POST /v1/handoff/{job}/upload | Smart+ repository upload | The Smart+ pack command |
GET /v1/handoff/{job}/patch | Smart+ patch download | The Smart+ apply command |
Authenticate with x-api-key or Authorization: Bearer. See Authentication.
POST /v1/messages
POSThttps://api.routerlane.com/v1/messages
The Anthropic Messages API. Send a tier as model; tools, images, system prompts, streaming and anthropic-beta headers pass through. The router sets thinking from its chosen effort.
curl https://api.routerlane.com/v1/messages \
-H "x-api-key: $ROUTERLANE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "smart",
"max_tokens": 4096,
"system": "You are a careful reviewer.",
"messages": [{"role": "user", "content": "Two workers deadlock on cache refresh. Find the race."}]
}'{
"id": "msg_9c2e41f07a3b4d58",
"type": "message",
"role": "assistant",
"model": "smart",
"content": [
{"type": "text", "text": "Both workers read the expiry before either one writes it. Take the lock before the read."}
],
"stop_reason": "end_turn",
"usage": {"input_tokens": 1834, "output_tokens": 412, "cache_read_input_tokens": 0}
}When the router turns thinking on, the content starts with a thinking block, as in the Anthropic API.
POST /v1/messages/count_tokens
POSThttps://api.routerlane.com/v1/messages/count_tokens
Counts the input tokens of a Messages request on the tier's default model. If the count cannot be made, the router answers with its own estimate and "estimated": true.
curl https://api.routerlane.com/v1/messages/count_tokens \
-H "x-api-key: $ROUTERLANE_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{"model": "normal", "messages": [{"role": "user", "content": "How long is this?"}]}'{"input_tokens": 14}POST /v1/chat/completions
POSThttps://api.routerlane.com/v1/chat/completions
The OpenAI Chat Completions API. Tools, images and streaming work as usual. reasoning_effort is set by the router.
curl https://api.routerlane.com/v1/chat/completions \
-H "Authorization: Bearer $ROUTERLANE_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "normal",
"messages": [
{"role": "system", "content": "Answer with the regex only."},
{"role": "user", "content": "Match dates like 2026-10-06."}
]
}'{
"id": "chatcmpl-5e1a9f3c27d04b86",
"object": "chat.completion",
"created": 1791331200,
"model": "normal",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "^\\d{4}-\\d{2}-\\d{2}$"},
"finish_reason": "stop"
}
],
"usage": {"prompt_tokens": 31, "completion_tokens": 14, "total_tokens": 45}
}POST /v1/responses
POSThttps://api.routerlane.com/v1/responses
The OpenAI Responses API, the one Codex uses. reasoning.effort is set by the router; other reasoning fields pass through.
curl https://api.routerlane.com/v1/responses \
-H "Authorization: Bearer $ROUTERLANE_API_KEY" \
-H "content-type: application/json" \
-d '{"model": "normal", "input": "In one sentence: what does a cursor-paginated /orders endpoint return on the last page?"}'{
"id": "resp_7d0c2b9e41a85f36",
"object": "response",
"model": "normal",
"status": "completed",
"output": [
{
"type": "message",
"role": "assistant",
"content": [{"type": "output_text", "text": "The last page returns its items with no next cursor."}]
}
],
"usage": {"input_tokens": 240, "output_tokens": 19, "total_tokens": 259}
}GET /v1/models
GEThttps://api.routerlane.com/v1/models
Lists the tiers your key's plan includes, in a shape both OpenAI and Anthropic clients read. GET /v1/models/{id} returns one entry. reasoning is false because the router sets the effort itself.
{
"object": "list",
"data": [
{
"id": "fast",
"object": "model",
"type": "model",
"created": 1790985600,
"created_at": "2026-10-02T00:00:00Z",
"owned_by": "routerlane",
"display_name": "Fast",
"description": "Lowest latency. Quick answers and small edits; thinks only when the task needs it.",
"context_window": 1048576,
"context_length": 1048576,
"max_input_tokens": 1048576,
"max_output_tokens": 131072,
"max_completion_tokens": 131072,
"capabilities": {
"anthropic_messages": true,
"openai_chat": true,
"openai_responses": true,
"tools": true,
"vision": true,
"json_schema": true,
"reasoning": false
},
"client_compat": {"pi": {"maxTokensField": "max_tokens", "supportsReasoningEffort": false}}
},
{
"id": "normal",
"object": "model",
"type": "model",
"created": 1790985600,
"created_at": "2026-10-02T00:00:00Z",
"owned_by": "routerlane",
"display_name": "Normal",
"description": "Balanced. Everyday coding work with enough reasoning to get it right first time.",
"context_window": 1048576,
"context_length": 1048576,
"max_input_tokens": 1048576,
"max_output_tokens": 131072,
"max_completion_tokens": 131072,
"capabilities": {
"anthropic_messages": true,
"openai_chat": true,
"openai_responses": true,
"tools": true,
"vision": true,
"json_schema": true,
"reasoning": false
},
"client_compat": {"pi": {"maxTokensField": "max_tokens", "supportsReasoningEffort": false}}
}
],
"has_more": false,
"first_id": "fast",
"last_id": "normal"
}Streaming
Set "stream": true in any of the three formats and the response is a server-sent event stream in that format. The router commits to a model only once it streams its first event, so failover is invisible to you. The x-router-* headers arrive with the first byte.
event: message_start
data: {"type":"message_start","message":{"id":"msg_9c2e41f07a3b4d58","type":"message","role":"assistant","model":"smart","content":[]}}
event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Both workers read"}}
event: message_stop
data: {"type":"message_stop"}data: {"id":"chatcmpl-5e1a9f3c27d04b86","object":"chat.completion.chunk","model":"normal","choices":[{"index":0,"delta":{"content":"^\\d{4}"}}]}
data: {"id":"chatcmpl-5e1a9f3c27d04b86","object":"chat.completion.chunk","model":"normal","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]event: response.output_text.delta
data: {"type":"response.output_text.delta","delta":"The last page returns"}
event: response.completed
data: {"type":"response.completed","response":{"id":"resp_7d0c2b9e41a85f36","model":"normal","status":"completed"}}- Chat Completions: send
"stream_options": {"include_usage": true}to receive the usage chunk at the end. - A content filter that blocks a streamed reply cuts it and sends an error event. See errors in streams.
- A Smart+ job turn streams its phases as ordinary text deltas.
Response headers
| Header | Values |
|---|---|
x-router-tier | fast, normal, smart, smart-plus |
x-router-model | The model that answered, for example deepseek-4.1-flash or glm-5.3 |
x-router-effort | off, low, medium, high, max |
x-router-source | rule, default, continuation, sticky, smartplus |
x-router-request-id | A 16-character id for this request |
x-router-cache | hit when served from the 10-minute response cache |
Caching
- Response cache. An identical request from the same account within 10 minutes is answered from cache, free, with
x-router-cache: hit. - Prompt cache. Conversations stay on one model so the model's prompt cache stays warm. Cached input is billed at the cached rate and reported in usage as
cache_read_input_tokensorcached_tokens. - Decision cache. The same task text reuses its classification for 24 hours.