# Streaming

> Streaming output over SSE — events per format, usage, keep-alive and mid-stream errors.

Text endpoints stream the response when the request has `"stream": true`. The
stream is Server-Sent Events (`text/event-stream`) in the same format as the
original API. Images and video do not stream.

## Example

```bash [curl]
curl -N https://ai-seller.vibe-codes.ru/v1/chat/completions \
  -H "Authorization: Bearer $AISELLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "stream": true,
    "stream_options": {"include_usage": true},
    "messages": [{"role": "user", "content": "Tell me a short story"}]
  }'
```

```python [Python]
from openai import OpenAI

client = OpenAI(base_url="https://ai-seller.vibe-codes.ru/v1", api_key="$AISELLER_API_KEY")
stream = client.chat.completions.create(
    model="gpt-5.6-sol",
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Tell me a short story"}],
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")
```

## Events and usage per format

| Endpoint | Events | Where usage is |
| --- | --- | --- |
| `/v1/chat/completions` | `chat.completion.chunk`, then `data: [DONE]` | The last chunk — only with `"stream_options": {"include_usage": true}` |
| `/v1/responses` | `response.created`, `response.output_text.delta`, …, `response.completed` | `response.completed` |
| `/v1/messages` | `message_start`, `content_block_*`, `message_delta`, `message_stop` | `message_start` and `message_delta` |

## Keep-alive and long silences

- Reasoning models may stay silent for tens of seconds before the first token.
  The gateway waits about a minute for the provider's response to start; give
  your client a generous timeout.
- The stream may contain comment lines starting with `:` (keep-alive). SDKs
  skip them; if you parse SSE yourself, ignore such lines.

## Mid-stream errors

- If the provider fails before the response starts, the gateway can still
  switch to another provider of the model or return a regular JSON error with
  an HTTP status — see [Errors](https://ai-seller.vibe-codes.ru/en/docs/errors).
- Once the response has started (HTTP 200 sent), there is no switching: the
  error arrives as an error event in the endpoint's format, or the stream just
  closes without its final event (`[DONE]`, `response.completed`,
  `message_stop`). Treat such a response as cut off and retry the request.
- If the provider broke off the stream before the first piece of the response,
  nothing is charged. If it broke off mid-response, or you aborted the request
  yourself, the input and the generated output are charged (see
  [Balance and pricing](https://ai-seller.vibe-codes.ru/en/docs/pricing)).
