> ## Documentation Index
> Fetch the complete documentation index at: https://docs.budecosystem.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Realtime Sessions

> Speech-to-speech over a WebSocket, on a realtime deployment, with the OpenAI Realtime protocol.

<RequestExample>
  ```python Python theme={null}
  import asyncio

  from openai import AsyncOpenAI


  async def main():
      client = AsyncOpenAI(
          api_key="YOUR_API_KEY",
          websocket_base_url="wss://gateway.bud.studio/v1",
      )

      # The deployment's defaults (voice, instructions, turn detection) already apply.
      async with client.realtime.connect(model="my-realtime-deployment") as connection:
          await connection.conversation.item.create(
              item={
                  "type": "message",
                  "role": "user",
                  "content": [{"type": "input_text", "text": "Say hello!"}],
              }
          )
          await connection.response.create()

          async for event in connection:
              if event.type in ("response.output_text.delta", "response.output_audio_transcript.delta"):
                  print(event.delta, end="", flush=True)
              elif event.type == "response.done":
                  break


  asyncio.run(main())
  ```

  ```javascript Node theme={null}
  import OpenAI from "openai";
  import { OpenAIRealtimeWS } from "openai/realtime/ws";

  const client = new OpenAI({
    apiKey: "YOUR_API_KEY",
    baseURL: "https://gateway.bud.studio/v1",
  });

  // Connects to wss://gateway.bud.studio/v1/realtime?model=my-realtime-deployment
  const rt = new OpenAIRealtimeWS({ model: "my-realtime-deployment" }, client);

  rt.socket.on("open", () => {
    rt.send({
      type: "conversation.item.create",
      item: {
        type: "message",
        role: "user",
        content: [{ type: "input_text", text: "Say hello!" }],
      },
    });
    rt.send({ type: "response.create" });
  });

  rt.on("response.output_audio_transcript.delta", (event) => process.stdout.write(event.delta));
  rt.on("response.done", () => rt.close());
  ```

  ```javascript Browser theme={null}
  // `ek` is a short-lived client secret your server minted with
  // POST /v1/realtime/client_secrets. Never put an API key in a browser.
  const ws = new WebSocket(
    "wss://gateway.bud.studio/v1/realtime?model=my-realtime-deployment",
    ["realtime", "openai-insecure-api-key." + ek],
  );

  ws.onopen = () => {
    ws.send(JSON.stringify({ type: "response.create" }));
  };

  ws.onmessage = (message) => {
    const event = JSON.parse(message.data);
    if (event.type === "response.output_audio.delta") {
      // base64 PCM16 audio: queue it for playback
    }
  };
  ```
</RequestExample>

<ResponseExample>
  ```json session.created theme={null}
  {
    "type": "session.created",
    "event_id": "event_B1c2...",
    "session": {
      "type": "realtime",
      "id": "sess_C9x...",
      "model": "my-realtime-deployment",
      "output_modalities": ["audio"],
      "audio": {
        "output": { "voice": "marin" }
      }
    }
  }
  ```

  ```json error (refused by policy) theme={null}
  {
    "type": "error",
    "event_id": "evt_bud_7f3a...",
    "error": {
      "type": "invalid_request_error",
      "code": "event_not_allowed",
      "message": "tools.mcp is not allowed on this deployment"
    }
  }
  ```
</ResponseExample>

A realtime deployment holds a speech-to-speech conversation over a WebSocket. The gateway serves
it at `wss://gateway.bud.studio/v1/realtime`, using the
[OpenAI Realtime API](https://platform.openai.com/docs/guides/realtime) (the GA protocol). The
OpenAI Python and Node SDKs, the OpenAI Agents SDK, LiveKit Agents and Pipecat connect to it
unchanged: point them at the gateway and use your Bud API key.

The deployment decides the provider, the model and the provider credential. Your client names the
deployment and never sees the provider's key. Every provider below speaks the same protocol to your
client; see [Providers](#providers) for what each one supports.

## Connect

```text theme={null}
wss://gateway.bud.studio/v1/realtime?model=<deployment>
```

| Query parameter | Required | Description                                                                                                                                 |
| --------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| `model`         | Yes      | The name of a realtime deployment your credential can use, or its endpoint ID. Without it the upgrade is refused with `400 model_required`. |
| `intent`        | No       | Accepted and ignored. The deployment decides the session type.                                                                              |
| `call_id`       | No       | Refused with `400 unsupported_parameter`. WebRTC, SIP and sideband calls are not supported.                                                 |
| `token`         | No       | Refused with `400 use_subprotocol`. Query strings end up in access logs, so credentials are not accepted there.                             |

Any other query parameter is ignored and not passed on to the provider.

The client's own `model` field (in `session.update`, which SDKs send routinely) is removed, not
refused. The provider always receives the deployment's model.

## Authentication

Send one credential, in any of these three ways:

| Where                  | Example                                                                  | Use from                            |
| ---------------------- | ------------------------------------------------------------------------ | ----------------------------------- |
| `Authorization` header | `Authorization: Bearer YOUR_API_KEY`                                     | Servers and SDKs                    |
| `api-key` header       | `api-key: YOUR_API_KEY`                                                  | Clients configured for Azure OpenAI |
| WebSocket subprotocol  | `Sec-WebSocket-Protocol: realtime, openai-insecure-api-key.<credential>` | Browsers, which cannot set headers  |

The credential can be a Bud API key (`bud_…`), a Keycloak access token, or a client secret
(`ek_bud_…`). In a browser, use a client secret: your server mints it with
[`POST /v1/realtime/client_secrets`](/api-sdk/realtime/client-secrets) and the browser never holds a
long-lived key.

With the subprotocol, the server selects `realtime` in its response (the Node `ws` package fails a
handshake with no selected protocol), and never echoes the credential back. The
`openai-organization.*`, `openai-project.*` and `openai-agents-sdk.*` subprotocols are tolerated and
ignored.

The beta protocol is not served: an `OpenAI-Beta` header or an `openai-beta.realtime-v1`
subprotocol is refused with `400 beta_api_shape_disabled`.

### Refusals before the connection opens

These are ordinary HTTP responses with the OpenAI error envelope.

| Status | When                                                                                                    |
| ------ | ------------------------------------------------------------------------------------------------------- |
| `400`  | A bad request: `model_required`, `unsupported_parameter`, `use_subprotocol`, `beta_api_shape_disabled`. |
| `401`  | No credential, or one that is not valid.                                                                |
| `403`  | The credential is valid but cannot use this deployment.                                                 |
| `404`  | No such deployment, or it does not serve `realtime_session`.                                            |
| `429`  | The deployment's rate or concurrency limit. Retry after the `Retry-After` header.                       |
| `503`  | The gateway is not ready, or the provider is failing and calls to it are paused for a moment.           |

## Providers

| Provider       | Deployment source | Models                                                                                     | How the gateway serves it               | Priced                             |
| -------------- | ----------------- | ------------------------------------------------------------------------------------------ | --------------------------------------- | ---------------------------------- |
| OpenAI         | `openai`          | `gpt-realtime`, `gpt-realtime-2.1`, `gpt-realtime-mini`, `gpt-realtime-whisper`, …         | Relays the protocol as is               | Per token, or per minute or second |
| Azure OpenAI   | `azure`           | Your Azure OpenAI realtime deployments                                                     | Relays the protocol as is               | Per token, or per minute or second |
| xAI            | `xai`             | Grok Voice (`grok-voice-think-fast-2.0`, `grok-voice-latest`)                              | Relays the protocol (xAI speaks it too) | Per token, or per minute or second |
| Google Gemini  | `gemini`          | Gemini Live models (`gemini-2.5-flash-native-audio-*`, `gemini-3.1-flash-live-preview`, …) | Translates to Gemini Live               | Per token, or per minute or second |
| Amazon Bedrock | `bedrock`         | Nova 2 Sonic (`amazon.nova-2-sonic-v1:0`)                                                  | Translates to Nova 2 Sonic              | Per token, or per minute or second |
| Deepgram       | `deepgram`        | A Deepgram Voice Agent                                                                     | Translates to the Voice Agent API       | Per minute or second only          |
| ElevenLabs     | `elevenlabs`      | An ElevenLabs agent                                                                        | Translates to ElevenLabs Agents         | Per minute or second only          |
| Hume           | `hume`            | Hume EVI                                                                                   | Translates to EVI                       | Per minute or second only          |

The OpenAI, Azure OpenAI, Gemini and Nova 2 Sonic models are in the model catalog. xAI's voice
model and the Deepgram, ElevenLabs and Hume agents are not: add them with **+ Cloud Model** and the
**Realtime** category. A deployment of the same provider for chat or for speech-to-text and
text-to-speech is unaffected; only a realtime deployment uses these sessions.

Where the gateway **translates**, your client still speaks the OpenAI Realtime protocol, and some of
it has nowhere to go. On every translated provider:

* Image input, MCP tools and stored prompt references are refused with `event_not_allowed`, whatever
  the deployment's policy says, so the operator is not offered those switches.
* Only speech-to-speech sessions are served. A model with no audio output cannot be deployed for
  realtime on these providers.
* The settings form offers only what the provider applies; a default it cannot apply is refused when
  the operator saves it, naming the provider.

### Google Gemini (Gemini Live)

* **Configuration is fixed when the session starts.** The voice, instructions, tools and turn detection
  are sent in Gemini's setup message. A later `session.update` that changes the voice or the tools is
  refused with `event_not_allowed`, and the session continues with the original setup.
* `conversation.item.truncate` has no Gemini equivalent and is refused with `event_not_allowed`.
* Your 24 kHz input audio is resampled to the 16 kHz Gemini expects; its 24 kHz output reaches you as
  `response.output_audio.delta`.
* Responses are audio, with a transcript. Turn detection is server VAD (prefix padding and silence
  duration; Gemini has no numeric threshold) or off.
* A session lasts at most 15 minutes, Gemini's limit for an audio-only session.
* Voices: Gemini's 30 prebuilt voices (`Puck`, `Kore`, `Charon`, …).

### Amazon Bedrock (Nova 2 Sonic)

* The deployment's credential is an AWS access key with a region; the gateway signs every request
  with it and never with its own AWS identity.
* Nova 2 Sonic closes a connection after 8 minutes. The gateway renews it inside your session, so
  you see one continuous session.
* The deployment sets the voice (`tiffany`, `matthew`, `amy`, …) and instructions. Nova 2 Sonic
  decides turn-taking itself, so turn detection is not a setting.

### xAI (Grok Voice)

* Voices: `eve`, `ara`, `rex`, `sal`, `leo`, or the ID of a voice cloned with xAI's Custom Voices API.
* Turn detection is server VAD (threshold 0.1 to 0.9) or off; speed is 0.7 to 1.5; input
  transcription takes a language hint.
* Input transcripts arrive as cumulative events: each carries the transcript so far, not a delta.
* xAI sends no `rate_limits.updated` event.
* xAI bills its voice agent **per minute** of audio. Pricing the deployment per minute matches that;
  per-token rates are accepted too, and apply to whatever usage xAI reports.

### Deepgram, ElevenLabs and Hume (per-minute agents)

* The agent itself (its speech recognition, language model and voice) is configured at the provider.
  The deployment adds instructions and, for Deepgram and Hume, a voice ID; an ElevenLabs agent keeps
  its own voice. For ElevenLabs, the model you add is the agent's ID.
* These providers report no token usage. A deployment is priced per minute or per second, and the
  session is billed in 60-second segments (a 150-second session is billed as 60 + 60 + 30 seconds).
  A per-token price, or per-modality rates, is refused when you publish.
* Turn-taking is the agent's own. A Hume EVI chat lasts at most 30 minutes; an ElevenLabs agent's own
  maximum conversation length can end a session earlier than the deployment's limit.

## Session types

The deployment decides whether a session is a `realtime` (speech-to-speech) session or a
`transcription` session. A model with audio output runs speech-to-speech sessions; a model without
audio output (for example `gpt-realtime-whisper`) runs transcription sessions. A `session.update`
that tries to change `session.type` is refused with `event_not_allowed`. Transcription sessions are
served on OpenAI and Azure OpenAI only.

## Deployment defaults

An operator can set a deployment's session defaults in its audio settings: voice, instructions,
output modality, turn detection, input transcription, noise reduction, maximum output tokens and
speed.

When the provider opens the session, the gateway sends one `session.update` carrying those defaults
before it forwards anything from your client. Your frames are held until the provider confirms
that update (at most 5 seconds; longer closes the session with `upstream_error`). Your own
`session.update` then overrides any default the deployment's policy leaves open.

Which value applies, first match wins:

1. What your client sends, unless the deployment's policy locks that field.
2. The deployment's default.
3. The provider's default.

On a transcription deployment the defaults also set `session.type` to `transcription` and the
transcription model to the deployment's model.

## Events

Every event of the OpenAI Realtime protocol is relayed, in both directions, except for the rules
below. Audio events (`input_audio_buffer.append`, `response.output_audio.delta`) are forwarded
without being inspected.

### From your client

| Event                                                                                                                                                          | What the gateway does                                                                                                                                                                                 |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `session.update`                                                                                                                                               | Removes `session.model` and `session.tracing`. Refuses a changed `session.type`, and anything the deployment's policy does not allow (see below). Caps `max_output_tokens` at the deployment's value. |
| `response.create`                                                                                                                                              | Checks that your credential is still valid, then applies the same `tools` and `instructions` rules to the per-response settings. `conversation: "none"` is allowed.                                   |
| `conversation.item.create`                                                                                                                                     | Refuses an `input_image` part when the deployment does not allow image input.                                                                                                                         |
| `input_audio_buffer.*`, `conversation.item.retrieve`, `conversation.item.truncate`, `conversation.item.delete`, `response.cancel`, `output_audio_buffer.clear` | Forwarded.                                                                                                                                                                                            |
| Any other type                                                                                                                                                 | Forwarded, so new provider events keep working.                                                                                                                                                       |

A refused event is not forwarded. You receive an `error` event with code `event_not_allowed` that
names the field, and the session continues.

### From the provider

| Event                                                   | What the gateway does                                                                                                    |
| ------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| `session.created`, `session.updated`                    | Forwarded, with `session.model` set to the deployment's name.                                                            |
| `response.done`                                         | Forwarded. Its `usage` is recorded and priced (see [Pricing](#pricing)).                                                 |
| `conversation.item.input_audio_transcription.completed` | Forwarded. Its `usage` is recorded and priced.                                                                           |
| `rate_limits.updated`                                   | **Not forwarded.** It reports the provider account's remaining budget, which every project using that credential shares. |
| `error`                                                 | Forwarded.                                                                                                               |
| Everything else                                         | Forwarded unchanged.                                                                                                     |

### Deployment policy

The deployment's client policy decides what a client may set. Fields a stored prompt, an MCP
server or a trace would reach belong to the **provider account**, which every Bud project using
the same credential shares, so those are off unless the operator turns them on.

| Policy                                               | Default | When off, the gateway refuses                                |
| ---------------------------------------------------- | ------- | ------------------------------------------------------------ |
| Client instructions (`allow_client_instructions`)    | On      | `instructions` in `session.update` and in `response.create`. |
| MCP tools (`allow_mcp_tools`)                        | Off     | Any tool with `type: "mcp"`.                                 |
| Stored prompt references (`allow_prompt_references`) | Off     | `session.prompt`.                                            |
| Image input (`allow_image_input`)                    | On      | `input_image` content in `conversation.item.create`.         |
| Transcription models (`input_transcription_models`)  | Any     | An `audio.input.transcription.model` outside the list.       |

`session.model` and `session.tracing` are always removed, silently.

### Frames

* Text frames only: the protocol is JSON. Binary frames are refused.
* At most 10 MiB per message. A larger one closes the session with `1009` (`message_too_large`).
* `permessage-deflate` is never negotiated.

## Limits

| Limit                  | Default                                                                                                          | What happens                                                                                                                                   |
| ---------------------- | ---------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| Keepalive              | A ping every 20 seconds, to your client and to the provider                                                      | Three missed pongs on either side close the session with `1011`. Browsers answer pings automatically.                                          |
| Maximum session length | 60 minutes; 30 minutes for Hume EVI and 15 minutes for Gemini Live, their own limits. A deployment can set less. | An `error` event with code `session_expiring` 60 seconds before the end, then `session_expired` and close `1000`.                              |
| Idle timeout           | 300 seconds                                                                                                      | A session with no traffic is closed as expired.                                                                                                |
| Rate and concurrency   | The deployment's rate limits                                                                                     | Checked when the session opens; an open session holds one concurrency slot until it closes. Over the limit, the upgrade is refused with `429`. |
| Slow client            | 5 seconds                                                                                                        | If your client stops reading while audio is queued for it, the session closes with `1011` rather than dropping audio.                          |

### Revocation during a session

Every 30 seconds, and before forwarding each `response.create`, the gateway checks that the
session may continue. It closes the session with `session_revoked` and `1008` when:

* the API key was deleted or has expired,
* the user was removed from the deployment's project,
* the deployment was unpublished, or removed from the key's allowed deployments.

A credential **expiring** does not end a session that is already open; a credential being
**revoked** does. With an API key the session closes within 30 seconds. With a Keycloak token it
can take up to 30 seconds plus the gateway's authorization cache lifetime (5 minutes by default).

The conversation lives at the provider. There is no resume: after any close, reconnect and start a
new conversation. A gateway or ingress restart also closes open sessions; reconnect when that
happens.

## Close codes

| Code   | Error code                                          | Meaning                                                                                                          | Retry?                    |
| ------ | --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | ------------------------- |
| `1000` | `session_expired`                                   | The session reached its maximum length or sat idle.                                                              | Reconnect if you need to. |
| `1008` | `session_revoked`                                   | The credential was revoked, or the deployment is no longer allowed.                                              | No.                       |
| `1009` | `message_too_large`                                 | A message was over 10 MiB.                                                                                       | Fix the client.           |
| `1011` | `upstream_error`, `client_too_slow`                 | The provider failed, closed, or stopped answering pings; or your client stopped reading for more than 5 seconds. | Yes.                      |
| `1012` | `server_shutdown`                                   | The gateway is restarting.                                                                                       | Yes, reconnect.           |
| `1013` | `rate_limit_exceeded`, `concurrency_limit_exceeded` | Over the deployment's limit.                                                                                     | Yes, with backoff.        |

`1008` means retrying will not help. The OpenAI Python SDK treats `1011`, `1012` and `1013` as
retryable, and does not retry `1008`.

## Pricing

A realtime deployment is priced either **per token**, with one rate per modality, or **per minute**
or **per second** of session time. An operator sets the price when publishing the deployment. The
per-minute agents (Deepgram, ElevenLabs, Hume) report no tokens and are priced per minute or per
second only.

### Per token

Every response is billed when it completes (`response.done`), from the usage the provider reports,
so a connection that drops after 40 minutes has already been billed for the responses it made.
Rates are quoted per 1,000,000 tokens (or the quantity the deployment is priced per):

| Rate                               | Applies to                                                      |
| ---------------------------------- | --------------------------------------------------------------- |
| Text, audio and image input        | Input tokens of each modality                                   |
| Cached text, audio and image input | The part of the input the provider served from its cache        |
| Text and audio output              | Output tokens of each modality                                  |
| Input transcription                | Per minute of audio transcribed, when input transcription is on |

Cached tokens are part of the input count, so they are subtracted from the input and billed at the
cached rate:

```text theme={null}
cost = [ (text_in  − cached_text)  × input_text  + cached_text  × cached_input_text
       + (audio_in − cached_audio) × input_audio + cached_audio × cached_input_audio
       + (image_in − cached_image) × input_image + cached_image × cached_input_image
       + text_out × output_text + audio_out × output_audio ] / 1,000,000
```

**Worked example.** A response reports 119 text input tokens (64 of them cached), 13 audio input
tokens, 30 text output tokens and 91 audio output tokens. At rates of $4.00 text input, $0.40
cached text input, $32.00 audio input, $24.00 text output and \$64.00 audio output per 1M tokens:

```text theme={null}
(55 × 4 + 64 × 0.40 + 13 × 32 + 30 × 24 + 91 × 64) / 1,000,000
  = 7205.6 / 1,000,000
  = $0.0072056
```

Each response re-reads the whole conversation so far, so a long session costs more per response
as it goes on. The deployment's maximum output tokens and maximum session length keep that in
check.

If the deployment has no rate for a modality that appears in a response, that part is recorded as
**unpriced**, never as free, and the rest of the response is priced normally.

Each provider takes only the rates its usage report can fill; publishing any other rate is refused,
naming the provider:

| Provider             | Rates                                                                            |
| -------------------- | -------------------------------------------------------------------------------- |
| OpenAI, Azure OpenAI | All of the above                                                                 |
| Google Gemini, xAI   | Text and audio input, cached text and audio input, text and audio output         |
| Amazon Nova 2 Sonic  | Text and audio (speech) input and output. Nova 2 Sonic reports no cached tokens. |

Gemini reports usage per response (`usageMetadata`) and Nova 2 Sonic per usage event; each report is
priced exactly once.

### Per minute or per second

The session's duration is billed in 60-second segments as the session runs, at one price per
minute or per second.

Realtime spend appears in your usage and request logs with the rest of your voice usage. It is
reported, but does not yet count toward a project's cost quota.

## Not supported

* WebRTC (`/v1/realtime/calls`), SIP, and sideband connections (`call_id`).
* Realtime translation sessions (`gpt-realtime-translate`, Gemini's `*-live-translate` models) and
  GPT-Live (`gpt-live-1`), which are different products.
* Gemini Live through Vertex AI: use the `gemini` (Google AI Studio) provider.
* Resuming a conversation after a disconnect, or failing over to another provider mid-session.
* Guardrails and moderation on audio: realtime sessions are not checked by `/v1/moderations` or
  guardrail profiles.
