Errors and limits
Errors come back in the format you called, with a status code, a type and a plain message. Most are final; 429s and 5xx are worth a retry.
Error bodies
{
"type": "error",
"error": {
"type": "permission_error",
"message": "the smart tier is not included in plan Starter; available: fast, normal"
}
}{
"error": {
"message": "the smart tier is not included in plan Starter; available: fast, normal",
"type": "permission_error",
"code": "permission_error"
}
}Status codes
| Status | Type | Example message | Cause | Retry |
|---|---|---|---|---|
| 400 | invalid_request_error | request body must be a JSON object | The body is not a JSON object | No |
| 400 | not_found_error | unknown model 'turbo': use fast, normal or smart | The model is not a tier or an alias | No |
| 401 | authentication_error | invalid or missing API key | No key, an unknown key or a revoked key | No |
| 403 | permission_error | the smart tier is not included in plan Starter; available: fast, normal | The tier is outside your plan | No |
| 403 | permission_error | client is suspended | The account is suspended | No |
| 403 | permission_error | blocked by the content filter 'Private keys' | A content filter blocked the input. It was never forwarded | No |
| 429 | rate_limit_error | rate limit: 60 requests per minute on plan Starter | Over the plan's requests per minute | Yes, with backoff |
| 429 | rate_limit_error | concurrency limit: 4 requests at once on plan Starter | Over the plan's concurrent requests | Yes, when one finishes |
| 429 | rate_limit_error | monthly token limit reached (20,000,000 tokens on plan Starter) | The plan's monthly token cap | Next month, or a bigger plan |
| 429 | rate_limit_error | monthly spend limit reached, with the cap | Your monthly spend cap | Next month, or a higher cap |
| 429 | overloaded_error | all upstream slots are busy; retry shortly | The request waited 2 minutes in the router's queue | Yes, with backoff |
| 502, 503, 504 | api_error | upstream failed: and the last error | Every model the router tried failed | Yes, once or twice |
| Other | as returned | the model's own error | The model rejected the request itself, for example an invalid tool schema | No |
When the last attempt returned an error body of its own, you receive that body with its status. Monthly limits reset at the start of each calendar month, UTC.
Errors in streams
Failover happens before the first byte, so most failures surface as an ordinary error response. Once a stream has started, a content filter that blocks the reply cuts it and sends an error event in your format. A connection to the model that breaks mid-stream ends the stream early, with an error event in the Anthropic format.
event: error
data: {"type": "error", "error": {"type": "permission_error", "message": "the reply was blocked by the content filter 'Customer data in replies'"}}data: {"error": {"message": "the reply was blocked by the content filter 'Customer data in replies'", "type": "permission_error"}}
data: [DONE]event: error
data: {"type": "error", "code": "permission_error", "message": "the reply was blocked by the content filter 'Customer data in replies'"}A regex filter holds back the last 512 characters of a streamed reply, so a match is cut before it reaches you.
Plan limits
| Plan | Requests a minute | Concurrent requests | Tokens a month | Tiers |
|---|---|---|---|---|
| Starter | 60 | 4 | 20M | Fast, Normal |
| Pro | 300 | 16 | no cap | Fast, Normal, Smart |
| Enterprise | custom | custom | custom | All four |
A streaming request counts against concurrency until its stream ends. Requests answered from the response cache do not count against the per-minute limit.
Retrying
- Retry 429 and 5xx with exponential backoff and jitter, starting at about one second. RouterLane does not send
Retry-After. - Do not retry 400, 401 or 403: the same request will fail the same way.
- A 5xx means the router already tried up to three models. One or two retries are enough.
- The OpenAI and Anthropic SDKs retry 429 and 5xx twice by default. Raise
max_retriesfor long agent runs.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.routerlane.com/v1",
api_key=os.environ["ROUTERLANE_API_KEY"],
max_retries=5,
timeout=900,
)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.routerlane.com/v1",
apiKey: process.env.ROUTERLANE_API_KEY,
maxRetries: 5,
timeout: 900_000,
});Getting help
Every answered request carries x-router-request-id. Send it to support@routerlane.com and we can see the decision, the attempts and the timing for that request. For an error, send the time, the status and the message.