Docs

Errors and limits

Errors come back in the format you called, with a status code, a type and a plain message. Most are final; 429s and 5xx are worth a retry.

Error bodies

/v1/messages
{
  "type": "error",
  "error": {
    "type": "permission_error",
    "message": "the smart tier is not included in plan Starter; available: fast, normal"
  }
}

Status codes

StatusTypeExample messageCauseRetry
400invalid_request_errorrequest body must be a JSON objectThe body is not a JSON objectNo
400not_found_errorunknown model 'turbo': use fast, normal or smartThe model is not a tier or an aliasNo
401authentication_errorinvalid or missing API keyNo key, an unknown key or a revoked keyNo
403permission_errorthe smart tier is not included in plan Starter; available: fast, normalThe tier is outside your planNo
403permission_errorclient is suspendedThe account is suspendedNo
403permission_errorblocked by the content filter 'Private keys'A content filter blocked the input. It was never forwardedNo
429rate_limit_errorrate limit: 60 requests per minute on plan StarterOver the plan's requests per minuteYes, with backoff
429rate_limit_errorconcurrency limit: 4 requests at once on plan StarterOver the plan's concurrent requestsYes, when one finishes
429rate_limit_errormonthly token limit reached (20,000,000 tokens on plan Starter)The plan's monthly token capNext month, or a bigger plan
429rate_limit_errormonthly spend limit reached, with the capYour monthly spend capNext month, or a higher cap
429overloaded_errorall upstream slots are busy; retry shortlyThe request waited 2 minutes in the router's queueYes, with backoff
502, 503, 504api_errorupstream failed: and the last errorEvery model the router tried failedYes, once or twice
Otheras returnedthe model's own errorThe model rejected the request itself, for example an invalid tool schemaNo

When the last attempt returned an error body of its own, you receive that body with its status. Monthly limits reset at the start of each calendar month, UTC.

Errors in streams

Failover happens before the first byte, so most failures surface as an ordinary error response. Once a stream has started, a content filter that blocks the reply cuts it and sends an error event in your format. A connection to the model that breaks mid-stream ends the stream early, with an error event in the Anthropic format.

event stream
event: error
data: {"type": "error", "error": {"type": "permission_error", "message": "the reply was blocked by the content filter 'Customer data in replies'"}}

A regex filter holds back the last 512 characters of a streamed reply, so a match is cut before it reaches you.

Plan limits

PlanRequests a minuteConcurrent requestsTokens a monthTiers
Starter60420MFast, Normal
Pro30016no capFast, Normal, Smart
EnterprisecustomcustomcustomAll four

A streaming request counts against concurrency until its stream ends. Requests answered from the response cache do not count against the per-minute limit.

Retrying

  • Retry 429 and 5xx with exponential backoff and jitter, starting at about one second. RouterLane does not send Retry-After.
  • Do not retry 400, 401 or 403: the same request will fail the same way.
  • A 5xx means the router already tried up to three models. One or two retries are enough.
  • The OpenAI and Anthropic SDKs retry 429 and 5xx twice by default. Raise max_retries for long agent runs.
retries.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.routerlane.com/v1",
    api_key=os.environ["ROUTERLANE_API_KEY"],
    max_retries=5,
    timeout=900,
)

Getting help

Every answered request carries x-router-request-id. Send it to support@routerlane.com and we can see the decision, the attempts and the timing for that request. For an error, send the time, the status and the message.