wss://gateway.bud.studio/ws: your microphone audio is transcribed as you speak, text you send is
spoken back, and — optionally — a voice agent answers each utterance through one of your chat
deployments. Every leg is a deployment in your project, used with that deployment’s own
provider credential, settings, limits and price.
Use it when you want to choose the transcription, speech and language models separately. For a
single speech-to-speech model (OpenAI or Azure OpenAI realtime), use a
realtime session.
Connect
Authenticate with your API key (or a Bud access token) in theAuthorization header. A browser
cannot set that header on a WebSocket, so it may instead send the key as its first message:
{"type": "authenticated"}. Any other message before it closes the socket.
Never put a key in the URL.
Then send one config message. The session answers ready with its stream_id, and from then
on binary frames you send are audio for the transcription leg; binary frames you receive are
synthesized speech.
Addressing deployments
The provider, the provider’s model, the credential and any endpoint address come from the
deployment.
provider may be omitted; if you send it, it is ignored. These fields are refused:
stt_config.api_key,tts_config.api_key— the deployment’s credential is used.conversation_config.base_url,api_key,reasoning_base_url,reasoning_api_key— the agent’s LLM is always your chat deployment, reached through the Bud gateway with your own key.dag_config.definition— use a server template (dag_config.template); its provider and LLM nodes are resolved to your deployments the same way.
extras (provider-specific pass-through parameters) are replaced by the deployment’s own.
Settings
The deployment’s saved audio settings are the defaults, and what you send inconfig wins:
the voice, sample rate, pronunciations and synthesis features for speech; the language,
transcription features and translation targets for transcription. The streaming-only
transcription features — interim_results, vad_events, endpointing_ms, utterance_end_ms,
speech_begin_event — are configured on the deployment under Streaming and apply only here.
A setting the provider cannot honour arrives as a config_warning rather than being dropped
silently.
Refusals
A refusedconfig answers {"type": "error", "message": "<code>: <explanation>"}:
Keeping a token fresh
A session keeps the credential it authenticated with, in memory, to reach your chat deployment and to be re-checked while it lives. An access token expires long before a call may end, so send a fresh one before it does:auth_refresh_refused and the session continues on its current credential. If the
token expires anyway, the agent’s next answer fails with an auth_expired error — the session
stays open, and the next utterance after a refresh is answered.
Deployment limits
Each leg’s deployment admits the session once, when it starts, and the session holds that slot until it ends: a concurrency limit of 10 on the transcription deployment means ten open sessions.Revocation
Every 30 seconds the session is re-checked. If your key is deleted, you are removed from the project, or a leg’s deployment is unpublished, the gateway sendssession_revoked and closes the
socket with 1008.
Pricing
Each leg is billed against its own deployment’s price, and every record carries your project, API key and the deployment’s name:- Transcription — one record per final transcript, for the audio streamed since the previous one (priced per second or per minute). Audio streamed after the last final transcript is billed when the session ends.
- Speech — one record per
speak, including each answer the voice agent speaks (priced per character). - The agent’s LLM — billed by the chat deployment, like any other chat completion made with your key.