> ## Documentation Index
> Fetch the complete documentation index at: https://docs.budecosystem.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice Pipeline Sessions

> Streaming transcription, speech and a voice agent over one WebSocket, each leg a deployment.

<RequestExample>
  ```python Python theme={null}
  import asyncio
  import json

  import websockets


  async def main():
      async with websockets.connect(
          "wss://gateway.bud.studio/ws",
          additional_headers={"Authorization": "Bearer YOUR_API_KEY"},
      ) as ws:
          await ws.send(json.dumps({
              "type": "config",
              "audio": True,
              # Each leg names a deployment in your project. No provider, no vendor key.
              "stt_config": {"model": "my-transcription-deployment", "language": "en-US",
                             "sample_rate": 16000, "channels": 1, "punctuation": True},
              "tts_config": {"model": "my-speech-deployment", "audio_format": "linear16",
                             "sample_rate": 24000},
              # Optional: a voice agent whose LLM is one of your chat deployments.
              "conversation_config": {"model": "my-chat-deployment",
                                      "system_prompt": "Answer in one short sentence."},
          }))
          async for frame in ws:
              if isinstance(frame, bytes):
                  continue  # synthesized audio (PCM16, 24 kHz)
              event = json.loads(frame)
              if event["type"] == "ready":
                  ...  # start streaming microphone PCM16 as binary frames
              elif event["type"] == "stt_result" and event["is_final"]:
                  print("you:", event["transcript"])


  asyncio.run(main())
  ```
</RequestExample>

A voice pipeline session streams audio in both directions over one WebSocket at
`wss://gateway.bud.studio/ws`: your microphone audio is transcribed as you speak, text you send is
spoken back, and — optionally — a voice agent answers each utterance through one of your chat
deployments. Every leg is a **deployment in your project**, used with that deployment's own
provider credential, settings, limits and price.

Use it when you want to choose the transcription, speech and language models separately. For a
single speech-to-speech model (OpenAI or Azure OpenAI realtime), use a
[realtime session](/api-sdk/realtime/realtime-sessions).

## Connect

Authenticate with your API key (or a Bud access token) in the `Authorization` header. A browser
cannot set that header on a WebSocket, so it may instead send the key as its **first message**:

```json theme={null}
{"type": "auth", "token": "YOUR_API_KEY"}
```

The gateway answers `{"type": "authenticated"}`. Any other message before it closes the socket.
Never put a key in the URL.

Then send one `config` message. The session answers `ready` with its `stream_id`, and from then
on binary frames you send are audio for the transcription leg; binary frames you receive are
synthesized speech.

## Addressing deployments

| Field                                 | Names                               | Capability                     |
| ------------------------------------- | ----------------------------------- | ------------------------------ |
| `stt_config.model`                    | a transcription deployment          | Audio transcription, streaming |
| `tts_config.model`                    | a text-to-speech deployment         | Text to speech                 |
| `conversation_config.model`           | a chat deployment (optional)        | Chat completions               |
| `conversation_config.reasoning_model` | a second chat deployment (optional) | Chat completions               |

The provider, the provider's model, the credential and any endpoint address come from the
deployment. `provider` may be omitted; if you send it, it is ignored. These fields are **refused**:

* `stt_config.api_key`, `tts_config.api_key` — the deployment's credential is used.
* `conversation_config.base_url`, `api_key`, `reasoning_base_url`, `reasoning_api_key` — the agent's
  LLM is always your chat deployment, reached through the Bud gateway with your own key.
* `dag_config.definition` — use a server template (`dag_config.template`); its provider and LLM
  nodes are resolved to your deployments the same way.

`extras` (provider-specific pass-through parameters) are replaced by the deployment's own.

### Settings

The deployment's saved audio settings are the defaults, and what you send in `config` wins:
the voice, sample rate, pronunciations and synthesis features for speech; the language,
transcription features and translation targets for transcription. The streaming-only
transcription features — `interim_results`, `vad_events`, `endpointing_ms`, `utterance_end_ms`,
`speech_begin_event` — are configured on the deployment under **Streaming** and apply only here.

A setting the provider cannot honour arrives as a `config_warning` rather than being dropped
silently.

## Refusals

A refused `config` answers `{"type": "error", "message": "<code>: <explanation>"}`:

| Code                                                | Why                                                                                           | The socket                             |
| --------------------------------------------------- | --------------------------------------------------------------------------------------------- | -------------------------------------- |
| `deployment_required`                               | a leg names no deployment (for example `provider` alone)                                      | stays open — send a corrected `config` |
| `client_key_not_accepted`                           | a leg carries its own `api_key`                                                               | stays open                             |
| `model_not_found`                                   | the name is not a deployment of that capability your key can reach                            | stays open                             |
| `unsupported_deployment`                            | a transcription deployment that only accepts uploaded files (self-hosted, Azure OpenAI)       | stays open                             |
| `deployment_misconfigured`                          | the deployment cannot reach its provider (for example an AWS deployment without its key pair) | stays open                             |
| `rate_limit_exceeded`, `concurrency_limit_exceeded` | a leg's deployment is at its limit                                                            | closed with **1013** — retry later     |
| `invalid_api_key`                                   | the credential is not valid                                                                   | closed with **1008**                   |

## Keeping a token fresh

A session keeps the credential it authenticated with, in memory, to reach your chat deployment
and to be re-checked while it lives. An access token expires long before a call may end, so send
a fresh one before it does:

```json theme={null}
{"type": "auth", "token": "NEW_ACCESS_TOKEN"}
```

The refresh must identify the same caller (the same API key, or the same user); anything else is
refused with `auth_refresh_refused` and the session continues on its current credential. If the
token expires anyway, the agent's next answer fails with an `auth_expired` error — the session
stays open, and the next utterance after a refresh is answered.

## Deployment limits

Each leg's deployment admits the session once, when it starts, and the session holds that slot
until it ends: a concurrency limit of 10 on the transcription deployment means ten open sessions.

## Revocation

Every 30 seconds the session is re-checked. If your key is deleted, you are removed from the
project, or a leg's deployment is unpublished, the gateway sends `session_revoked` and closes the
socket with **1008**.

## Pricing

Each leg is billed against its own deployment's price, and every record carries your project,
API key and the deployment's name:

* **Transcription** — one record per final transcript, for the audio streamed since the previous
  one (priced per second or per minute). Audio streamed after the last final transcript is billed
  when the session ends.
* **Speech** — one record per `speak`, including each answer the voice agent speaks (priced per
  character).
* **The agent's LLM** — billed by the chat deployment, like any other chat completion made with
  your key.

A speech deployment priced per second or per minute of audio is recorded with its character count
and no price on this transport.
