Limits
Concurrent requests, per-key spend limits, balance reserve and request body size.
| Limit | Value | When exceeded |
|---|---|---|
| Concurrent requests | Limited per account, across all keys | 429 rate_limit_exceeded, Retry-After: 1 |
| The model provider's rate limit | Depends on the provider | 429 rate_limit_exceeded, Retry-After when the provider named a pause |
| Key spend limit | Set per key in the cabinet | 403 spend_limit_exceeded |
| Balance | Free remainder minus the reserve of requests in progress | 402 insufficient_quota |
| Request body: text | 25 MiB — /v1/chat/completions, /v1/messages, /v1/responses | 413 |
| Request body: media | 70 MiB — /v1/images/generations, /v1/videos | 413 |
| Images per request | n from 1 to 10 | The value is clamped to the range |
Concurrent requests
The number of requests running at the same time for an account is limited —
across all keys. There is no separate requests-per-minute limit. Once exceeded,
the next request gets 429 rate_limit_exceeded with a Retry-After header —
retry after the indicated pause, when one of the running requests finishes.
Video jobs (POST /v1/videos) are not subject to this limit.
If the model's provider hits its rate limit, the gateway also responds 429;
Retry-After is present when the provider named a pause (1 to 60 seconds).
Spend limit and balance reserve
- A key can have a spend limit: once reached, new requests on that key get
403 spend_limit_exceededwhile your other keys keep working. - While a request runs, its expected cost is reserved on your balance and in
the key's limit (video jobs hold no reserve). If the remainder is held by requests in progress, a new one
gets
402or403— retry once they finish.
More on charges in Balance and pricing.
Request body size
25 MiB for text endpoints and 70 MiB for images and video — the ceiling for
base64 media attachments. Anything larger gets 413. Pass large references as
links rather than data: URLs when the model supports it.
Updated September 29, 2026