Limits

Limits

Concurrent requests, per-key spend limits, balance reserve and request body size.

LimitValueWhen exceeded
Concurrent requestsLimited per account, across all keys429 rate_limit_exceeded, Retry-After: 1
The model provider's rate limitDepends on the provider429 rate_limit_exceeded, Retry-After when the provider named a pause
Key spend limitSet per key in the cabinet403 spend_limit_exceeded
BalanceFree remainder minus the reserve of requests in progress402 insufficient_quota
Request body: text25 MiB — /v1/chat/completions, /v1/messages, /v1/responses413
Request body: media70 MiB — /v1/images/generations, /v1/videos413
Images per requestn from 1 to 10The value is clamped to the range

Concurrent requests

The number of requests running at the same time for an account is limited — across all keys. There is no separate requests-per-minute limit. Once exceeded, the next request gets 429 rate_limit_exceeded with a Retry-After header — retry after the indicated pause, when one of the running requests finishes. Video jobs (POST /v1/videos) are not subject to this limit.

If the model's provider hits its rate limit, the gateway also responds 429; Retry-After is present when the provider named a pause (1 to 60 seconds).

Spend limit and balance reserve

  • A key can have a spend limit: once reached, new requests on that key get 403 spend_limit_exceeded while your other keys keep working.
  • While a request runs, its expected cost is reserved on your balance and in the key's limit (video jobs hold no reserve). If the remainder is held by requests in progress, a new one gets 402 or 403 — retry once they finish.

More on charges in Balance and pricing.

Request body size

25 MiB for text endpoints and 70 MiB for images and video — the ceiling for base64 media attachments. Anything larger gets 413. Pass large references as links rather than data: URLs when the model supports it.

Updated September 29, 2026