Docs

API reference

Base host https://api.routerlane.com. Requests and responses follow the Anthropic and OpenAI formats; RouterLane adds tier names and decision headers.

EndpointFormatUsed by
POST /v1/messagesAnthropic MessagesClaude Code, pi, Anthropic SDKs
POST /v1/messages/count_tokensAnthropic token countingClaude Code, Anthropic SDKs
POST /v1/chat/completionsOpenAI Chat CompletionsOpenAI SDKs, Vercel AI SDK, LangChain
POST /v1/responsesOpenAI ResponsesCodex, OpenAI SDKs
GET /v1/modelsModel list, readable by bothAny client
POST /v1/handoff/{job}/uploadSmart+ repository uploadThe Smart+ pack command
GET /v1/handoff/{job}/patchSmart+ patch downloadThe Smart+ apply command

Authenticate with x-api-key or Authorization: Bearer. See Authentication.

POST /v1/messages

POSThttps://api.routerlane.com/v1/messages

The Anthropic Messages API. Send a tier as model; tools, images, system prompts, streaming and anthropic-beta headers pass through. The router sets thinking from its chosen effort.

request
curl https://api.routerlane.com/v1/messages \
  -H "x-api-key: $ROUTERLANE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{
    "model": "smart",
    "max_tokens": 4096,
    "system": "You are a careful reviewer.",
    "messages": [{"role": "user", "content": "Two workers deadlock on cache refresh. Find the race."}]
  }'
response
{
  "id": "msg_9c2e41f07a3b4d58",
  "type": "message",
  "role": "assistant",
  "model": "smart",
  "content": [
    {"type": "text", "text": "Both workers read the expiry before either one writes it. Take the lock before the read."}
  ],
  "stop_reason": "end_turn",
  "usage": {"input_tokens": 1834, "output_tokens": 412, "cache_read_input_tokens": 0}
}

When the router turns thinking on, the content starts with a thinking block, as in the Anthropic API.

POST /v1/messages/count_tokens

POSThttps://api.routerlane.com/v1/messages/count_tokens

Counts the input tokens of a Messages request on the tier's default model. If the count cannot be made, the router answers with its own estimate and "estimated": true.

request
curl https://api.routerlane.com/v1/messages/count_tokens \
  -H "x-api-key: $ROUTERLANE_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "content-type: application/json" \
  -d '{"model": "normal", "messages": [{"role": "user", "content": "How long is this?"}]}'
response
{"input_tokens": 14}

POST /v1/chat/completions

POSThttps://api.routerlane.com/v1/chat/completions

The OpenAI Chat Completions API. Tools, images and streaming work as usual. reasoning_effort is set by the router.

request
curl https://api.routerlane.com/v1/chat/completions \
  -H "Authorization: Bearer $ROUTERLANE_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "normal",
    "messages": [
      {"role": "system", "content": "Answer with the regex only."},
      {"role": "user", "content": "Match dates like 2026-10-06."}
    ]
  }'
response
{
  "id": "chatcmpl-5e1a9f3c27d04b86",
  "object": "chat.completion",
  "created": 1791331200,
  "model": "normal",
  "choices": [
    {
      "index": 0,
      "message": {"role": "assistant", "content": "^\\d{4}-\\d{2}-\\d{2}$"},
      "finish_reason": "stop"
    }
  ],
  "usage": {"prompt_tokens": 31, "completion_tokens": 14, "total_tokens": 45}
}

POST /v1/responses

POSThttps://api.routerlane.com/v1/responses

The OpenAI Responses API, the one Codex uses. reasoning.effort is set by the router; other reasoning fields pass through.

request
curl https://api.routerlane.com/v1/responses \
  -H "Authorization: Bearer $ROUTERLANE_API_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "normal", "input": "In one sentence: what does a cursor-paginated /orders endpoint return on the last page?"}'
response
{
  "id": "resp_7d0c2b9e41a85f36",
  "object": "response",
  "model": "normal",
  "status": "completed",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "The last page returns its items with no next cursor."}]
    }
  ],
  "usage": {"input_tokens": 240, "output_tokens": 19, "total_tokens": 259}
}

GET /v1/models

GEThttps://api.routerlane.com/v1/models

Lists the tiers your key's plan includes, in a shape both OpenAI and Anthropic clients read. GET /v1/models/{id} returns one entry. reasoning is false because the router sets the effort itself.

response, Starter plan
{
  "object": "list",
  "data": [
    {
      "id": "fast",
      "object": "model",
      "type": "model",
      "created": 1790985600,
      "created_at": "2026-10-02T00:00:00Z",
      "owned_by": "routerlane",
      "display_name": "Fast",
      "description": "Lowest latency. Quick answers and small edits; thinks only when the task needs it.",
      "context_window": 1048576,
      "context_length": 1048576,
      "max_input_tokens": 1048576,
      "max_output_tokens": 131072,
      "max_completion_tokens": 131072,
      "capabilities": {
        "anthropic_messages": true,
        "openai_chat": true,
        "openai_responses": true,
        "tools": true,
        "vision": true,
        "json_schema": true,
        "reasoning": false
      },
      "client_compat": {"pi": {"maxTokensField": "max_tokens", "supportsReasoningEffort": false}}
    },
    {
      "id": "normal",
      "object": "model",
      "type": "model",
      "created": 1790985600,
      "created_at": "2026-10-02T00:00:00Z",
      "owned_by": "routerlane",
      "display_name": "Normal",
      "description": "Balanced. Everyday coding work with enough reasoning to get it right first time.",
      "context_window": 1048576,
      "context_length": 1048576,
      "max_input_tokens": 1048576,
      "max_output_tokens": 131072,
      "max_completion_tokens": 131072,
      "capabilities": {
        "anthropic_messages": true,
        "openai_chat": true,
        "openai_responses": true,
        "tools": true,
        "vision": true,
        "json_schema": true,
        "reasoning": false
      },
      "client_compat": {"pi": {"maxTokensField": "max_tokens", "supportsReasoningEffort": false}}
    }
  ],
  "has_more": false,
  "first_id": "fast",
  "last_id": "normal"
}

Streaming

Set "stream": true in any of the three formats and the response is a server-sent event stream in that format. The router commits to a model only once it streams its first event, so failover is invisible to you. The x-router-* headers arrive with the first byte.

event stream
event: message_start
data: {"type":"message_start","message":{"id":"msg_9c2e41f07a3b4d58","type":"message","role":"assistant","model":"smart","content":[]}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"Both workers read"}}

event: message_stop
data: {"type":"message_stop"}
  • Chat Completions: send "stream_options": {"include_usage": true} to receive the usage chunk at the end.
  • A content filter that blocks a streamed reply cuts it and sends an error event. See errors in streams.
  • A Smart+ job turn streams its phases as ordinary text deltas.

Response headers

HeaderValues
x-router-tierfast, normal, smart, smart-plus
x-router-modelThe model that answered, for example deepseek-4.1-flash or glm-5.3
x-router-effortoff, low, medium, high, max
x-router-sourcerule, default, continuation, sticky, smartplus
x-router-request-idA 16-character id for this request
x-router-cachehit when served from the 10-minute response cache

Caching

  • Response cache. An identical request from the same account within 10 minutes is answered from cache, free, with x-router-cache: hit.
  • Prompt cache. Conversations stay on one model so the model's prompt cache stays warm. Cached input is billed at the cached rate and reported in usage as cache_read_input_tokens or cached_tokens.
  • Decision cache. The same task text reuses its classification for 24 hours.