wss://gateway.bud.studio/v1/realtime, using the
OpenAI Realtime API (the GA protocol). The
OpenAI Python and Node SDKs, the OpenAI Agents SDK, LiveKit Agents and Pipecat connect to it
unchanged: point them at the gateway and use your Bud API key.
The deployment decides the provider, the model and the provider credential. Your client names the
deployment and never sees the provider’s key. Every provider below speaks the same protocol to your
client; see Providers for what each one supports.
Connect
Any other query parameter is ignored and not passed on to the provider.
The client’s own
model field (in session.update, which SDKs send routinely) is removed, not
refused. The provider always receives the deployment’s model.
Authentication
Send one credential, in any of these three ways:
The credential can be a Bud API key (
bud_…), a Keycloak access token, or a client secret
(ek_bud_…). In a browser, use a client secret: your server mints it with
POST /v1/realtime/client_secrets and the browser never holds a
long-lived key.
With the subprotocol, the server selects realtime in its response (the Node ws package fails a
handshake with no selected protocol), and never echoes the credential back. The
openai-organization.*, openai-project.* and openai-agents-sdk.* subprotocols are tolerated and
ignored.
The beta protocol is not served: an OpenAI-Beta header or an openai-beta.realtime-v1
subprotocol is refused with 400 beta_api_shape_disabled.
Refusals before the connection opens
These are ordinary HTTP responses with the OpenAI error envelope.Providers
The OpenAI, Azure OpenAI, Gemini and Nova 2 Sonic models are in the model catalog. xAI’s voice
model and the Deepgram, ElevenLabs and Hume agents are not: add them with + Cloud Model and the
Realtime category. A deployment of the same provider for chat or for speech-to-text and
text-to-speech is unaffected; only a realtime deployment uses these sessions.
Where the gateway translates, your client still speaks the OpenAI Realtime protocol, and some of
it has nowhere to go. On every translated provider:
- Image input, MCP tools and stored prompt references are refused with
event_not_allowed, whatever the deployment’s policy says, so the operator is not offered those switches. - Only speech-to-speech sessions are served. A model with no audio output cannot be deployed for realtime on these providers.
- The settings form offers only what the provider applies; a default it cannot apply is refused when the operator saves it, naming the provider.
Google Gemini (Gemini Live)
- Configuration is fixed when the session starts. The voice, instructions, tools and turn detection
are sent in Gemini’s setup message. A later
session.updatethat changes the voice or the tools is refused withevent_not_allowed, and the session continues with the original setup. conversation.item.truncatehas no Gemini equivalent and is refused withevent_not_allowed.- Your 24 kHz input audio is resampled to the 16 kHz Gemini expects; its 24 kHz output reaches you as
response.output_audio.delta. - Responses are audio, with a transcript. Turn detection is server VAD (prefix padding and silence duration; Gemini has no numeric threshold) or off.
- A session lasts at most 15 minutes, Gemini’s limit for an audio-only session.
- Voices: Gemini’s 30 prebuilt voices (
Puck,Kore,Charon, …).
Amazon Bedrock (Nova 2 Sonic)
- The deployment’s credential is an AWS access key with a region; the gateway signs every request with it and never with its own AWS identity.
- Nova 2 Sonic closes a connection after 8 minutes. The gateway renews it inside your session, so you see one continuous session.
- The deployment sets the voice (
tiffany,matthew,amy, …) and instructions. Nova 2 Sonic decides turn-taking itself, so turn detection is not a setting.
xAI (Grok Voice)
- Voices:
eve,ara,rex,sal,leo, or the ID of a voice cloned with xAI’s Custom Voices API. - Turn detection is server VAD (threshold 0.1 to 0.9) or off; speed is 0.7 to 1.5; input transcription takes a language hint.
- Input transcripts arrive as cumulative events: each carries the transcript so far, not a delta.
- xAI sends no
rate_limits.updatedevent. - xAI bills its voice agent per minute of audio. Pricing the deployment per minute matches that; per-token rates are accepted too, and apply to whatever usage xAI reports.
Deepgram, ElevenLabs and Hume (per-minute agents)
- The agent itself (its speech recognition, language model and voice) is configured at the provider. The deployment adds instructions and, for Deepgram and Hume, a voice ID; an ElevenLabs agent keeps its own voice. For ElevenLabs, the model you add is the agent’s ID.
- These providers report no token usage. A deployment is priced per minute or per second, and the session is billed in 60-second segments (a 150-second session is billed as 60 + 60 + 30 seconds). A per-token price, or per-modality rates, is refused when you publish.
- Turn-taking is the agent’s own. A Hume EVI chat lasts at most 30 minutes; an ElevenLabs agent’s own maximum conversation length can end a session earlier than the deployment’s limit.
Session types
The deployment decides whether a session is arealtime (speech-to-speech) session or a
transcription session. A model with audio output runs speech-to-speech sessions; a model without
audio output (for example gpt-realtime-whisper) runs transcription sessions. A session.update
that tries to change session.type is refused with event_not_allowed. Transcription sessions are
served on OpenAI and Azure OpenAI only.
Deployment defaults
An operator can set a deployment’s session defaults in its audio settings: voice, instructions, output modality, turn detection, input transcription, noise reduction, maximum output tokens and speed. When the provider opens the session, the gateway sends onesession.update carrying those defaults
before it forwards anything from your client. Your frames are held until the provider confirms
that update (at most 5 seconds; longer closes the session with upstream_error). Your own
session.update then overrides any default the deployment’s policy leaves open.
Which value applies, first match wins:
- What your client sends, unless the deployment’s policy locks that field.
- The deployment’s default.
- The provider’s default.
session.type to transcription and the
transcription model to the deployment’s model.
Events
Every event of the OpenAI Realtime protocol is relayed, in both directions, except for the rules below. Audio events (input_audio_buffer.append, response.output_audio.delta) are forwarded
without being inspected.
From your client
A refused event is not forwarded. You receive an
error event with code event_not_allowed that
names the field, and the session continues.
From the provider
Deployment policy
The deployment’s client policy decides what a client may set. Fields a stored prompt, an MCP server or a trace would reach belong to the provider account, which every Bud project using the same credential shares, so those are off unless the operator turns them on.session.model and session.tracing are always removed, silently.
Frames
- Text frames only: the protocol is JSON. Binary frames are refused.
- At most 10 MiB per message. A larger one closes the session with
1009(message_too_large). permessage-deflateis never negotiated.
Limits
Revocation during a session
Every 30 seconds, and before forwarding eachresponse.create, the gateway checks that the
session may continue. It closes the session with session_revoked and 1008 when:
- the API key was deleted or has expired,
- the user was removed from the deployment’s project,
- the deployment was unpublished, or removed from the key’s allowed deployments.
Close codes
1008 means retrying will not help. The OpenAI Python SDK treats 1011, 1012 and 1013 as
retryable, and does not retry 1008.
Pricing
A realtime deployment is priced either per token, with one rate per modality, or per minute or per second of session time. An operator sets the price when publishing the deployment. The per-minute agents (Deepgram, ElevenLabs, Hume) report no tokens and are priced per minute or per second only.Per token
Every response is billed when it completes (response.done), from the usage the provider reports,
so a connection that drops after 40 minutes has already been billed for the responses it made.
Rates are quoted per 1,000,000 tokens (or the quantity the deployment is priced per):
Cached tokens are part of the input count, so they are subtracted from the input and billed at the
cached rate:
Gemini reports usage per response (
usageMetadata) and Nova 2 Sonic per usage event; each report is
priced exactly once.
Per minute or per second
The session’s duration is billed in 60-second segments as the session runs, at one price per minute or per second. Realtime spend appears in your usage and request logs with the rest of your voice usage. It is reported, but does not yet count toward a project’s cost quota.Not supported
- WebRTC (
/v1/realtime/calls), SIP, and sideband connections (call_id). - Realtime translation sessions (
gpt-realtime-translate, Gemini’s*-live-translatemodels) and GPT-Live (gpt-live-1), which are different products. - Gemini Live through Vertex AI: use the
gemini(Google AI Studio) provider. - Resuming a conversation after a disconnect, or failing over to another provider mid-session.
- Guardrails and moderation on audio: realtime sessions are not checked by
/v1/moderationsor guardrail profiles.