Skip to main content
A voice pipeline session streams audio in both directions over one WebSocket at wss://gateway.bud.studio/ws: your microphone audio is transcribed as you speak, text you send is spoken back, and — optionally — a voice agent answers each utterance through one of your chat deployments. Every leg is a deployment in your project, used with that deployment’s own provider credential, settings, limits and price. Use it when you want to choose the transcription, speech and language models separately. For a single speech-to-speech model (OpenAI or Azure OpenAI realtime), use a realtime session.

Connect

Authenticate with your API key (or a Bud access token) in the Authorization header. A browser cannot set that header on a WebSocket, so it may instead send the key as its first message:
The gateway answers {"type": "authenticated"}. Any other message before it closes the socket. Never put a key in the URL. Then send one config message. The session answers ready with its stream_id, and from then on binary frames you send are audio for the transcription leg; binary frames you receive are synthesized speech.

Addressing deployments

The provider, the provider’s model, the credential and any endpoint address come from the deployment. provider may be omitted; if you send it, it is ignored. These fields are refused:
  • stt_config.api_key, tts_config.api_key — the deployment’s credential is used.
  • conversation_config.base_url, api_key, reasoning_base_url, reasoning_api_key — the agent’s LLM is always your chat deployment, reached through the Bud gateway with your own key.
  • dag_config.definition — use a server template (dag_config.template); its provider and LLM nodes are resolved to your deployments the same way.
extras (provider-specific pass-through parameters) are replaced by the deployment’s own.

Settings

The deployment’s saved audio settings are the defaults, and what you send in config wins: the voice, sample rate, pronunciations and synthesis features for speech; the language, transcription features and translation targets for transcription. The streaming-only transcription features — interim_results, vad_events, endpointing_ms, utterance_end_ms, speech_begin_event — are configured on the deployment under Streaming and apply only here. A setting the provider cannot honour arrives as a config_warning rather than being dropped silently.

Refusals

A refused config answers {"type": "error", "message": "<code>: <explanation>"}:

Keeping a token fresh

A session keeps the credential it authenticated with, in memory, to reach your chat deployment and to be re-checked while it lives. An access token expires long before a call may end, so send a fresh one before it does:
The refresh must identify the same caller (the same API key, or the same user); anything else is refused with auth_refresh_refused and the session continues on its current credential. If the token expires anyway, the agent’s next answer fails with an auth_expired error — the session stays open, and the next utterance after a refresh is answered.

Deployment limits

Each leg’s deployment admits the session once, when it starts, and the session holds that slot until it ends: a concurrency limit of 10 on the transcription deployment means ten open sessions.

Revocation

Every 30 seconds the session is re-checked. If your key is deleted, you are removed from the project, or a leg’s deployment is unpublished, the gateway sends session_revoked and closes the socket with 1008.

Pricing

Each leg is billed against its own deployment’s price, and every record carries your project, API key and the deployment’s name:
  • Transcription — one record per final transcript, for the audio streamed since the previous one (priced per second or per minute). Audio streamed after the last final transcript is billed when the session ends.
  • Speech — one record per speak, including each answer the voice agent speaks (priced per character).
  • The agent’s LLM — billed by the chat deployment, like any other chat completion made with your key.
A speech deployment priced per second or per minute of audio is recorded with its character count and no price on this transport.