Starter
two tiersFor getting an agent or a product onto RouterLane.
- Fast and Normal
- 60 requests a minute
- 4 concurrent requests
- 20M tokens a month
One price per tier, per million tokens, whichever model answers. Metered per request and billed monthly. No subscription fee.
| Tier | Model ID | Input | Output | Cached input |
|---|---|---|---|---|
| Fast | fast | $0.15 | $0.60 | $0.03 |
| Normal | normal | $0.40 | $1.60 | $0.08 |
| Smart | smart | $1.00 | $4.00 | $0.20 |
| Smart+ | smart-plus | $1.00 | $4.00 | $0.20 |
USD per 1M tokens. Cached input is input the model reads from its prompt cache. Smart+ bills every token its agents spend at Smart rates.
A plan sets which tiers you can call and how fast. The per-token prices above apply on every plan.
For getting an agent or a product onto RouterLane.
For teams that run agents all day.
A dedicated deployment with its own console.
Enter millions of tokens per month per tier. Agent loops resend their history every turn, so a large share of their input is usually served from the cache.
| Tier | Input, M tokens | Output, M tokens | Cost |
|---|---|---|---|
| Fast | $2.76 | ||
| Normal | $14.72 | ||
| Smart | $9.20 | ||
| Smart+ | $0.00 |
Smart is included on Pro and Enterprise.
An identical request repeated within 10 minutes is answered from the response cache, free. The response carries x-router-cache: hit.
Reading the task and choosing the model is on us, on every new task.
When a model fails before the first byte and the router moves to the next one, only the attempt that answered is metered.
If every attempt fails, the error comes back to you and nothing is billed.
A bad key, a tier outside your plan, a limit reached, or a filter that blocks the input: the request is never forwarded and never billed.
The repository goes straight to RouterLane over HTTPS, never through a model, and costs no tokens.
Every request is metered as it happens, with its input, output and cached tokens. You are billed monthly for what you used. There is no subscription fee.
No. You pay your tier's rate whichever model answers. When Normal sends a hard task to GLM 5.3, it is still billed at Normal's rate.
Input the model reads from its prompt cache instead of processing again. It is billed at the cached rate, a fifth of the input rate. The router keeps each conversation on one model so that cache stays warm.
Every token the job's agents spend is billed at Smart rates, on the request for the turn that started the job. Nothing is billed twice. Each job has a spend ceiling, $1.50 by default: a job that reaches it stops, delivers what it has, and says so.
Requests past your plan's rate, concurrency, monthly token or monthly spend limit get a 429 with a message that names the limit. Monthly limits reset at the start of each calendar month, UTC.
Yes. Ask for a monthly spend cap on your account and requests past it are refused with a 429 until the month turns. On Enterprise you set caps per plan and per client in the console.
Every response carries an x-router-request-id. Quote it to support for the request's tokens, model and cost. Enterprise consoles show cost on every request, turn and session.
Keys are issued on request. Tell us what you are building and which clients you use, and we will send a key by email.