> ## Documentation Index
> Fetch the complete documentation index at: https://docs.budecosystem.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice Agents

> Talk to an agent: its prompt, tools and governance, with a speech-to-text and a text-to-speech deployment as its ears and voice.

<RequestExample>
  ```python Python theme={null}
  import asyncio
  import base64
  import json

  import websockets


  async def main():
      # An OpenAI Realtime (GA) client, pointed at an agent instead of a model.
      async with websockets.connect(
          "wss://gateway.bud.studio/v1/realtime?model=prompt:support",
          additional_headers={"Authorization": "Bearer YOUR_API_KEY"},
      ) as ws:
          async for frame in ws:
              event = json.loads(frame)
              if event["type"] == "session.created":
                  ...  # start sending microphone audio: PCM16, 24 kHz, mono
                  # await ws.send(json.dumps({"type": "input_audio_buffer.append",
                  #                           "audio": base64.b64encode(chunk).decode()}))
              elif event["type"] == "response.output_audio.delta":
                  play(base64.b64decode(event["delta"]))  # PCM16, 24 kHz
              elif event["type"] == "conversation.item.input_audio_transcription.completed":
                  print("you:", event["transcript"])
              elif event["type"] == "response.done":
                  print("agent:", event["response"]["status"])


  asyncio.run(main())
  ```
</RequestExample>

A **voice agent** is an agent you can talk to. The agent stays exactly what it is — its prompt,
model, tools, skills and governance run on every turn — and two deployments in its project give it
ears and a voice: a **speech-to-text** deployment that transcribes the caller, and a
**text-to-speech** deployment that speaks the agent's answers. The gateway takes the turns, streams
the answer into speech sentence by sentence, and stops talking the moment the caller talks over it.

## Make an agent a voice agent

In the agent builder, click the microphone on the **Input** card (it opens **Settings → Voice → Listen**)
or on the **Output** card (**Voice → Speak**), and turn on **Voice agent**. The same **Voice** tab is in
the **Settings** menu beside **Runtime** and **Model**, and one **Save** keeps all three.

| Setting | What it does |
| - | - |
| Listen | The speech-to-text deployment; its language, picked from the languages that deployment's model takes (**Detect automatically** where the model can); and key terms: product names, people or jargon the transcriber should spell correctly, shown only for vendors that use them |
| Speak | The text-to-speech deployment, the voice (listed from the deployment's vendor account) and speed |
| Voice instructions | Added to the system prompt on voice calls only — for example "answer in one or two short sentences" |
| Greeting | Spoken when a call opens, before the caller says anything |
| Turn taking | **Semantic** (waits for a finished thought), **Silence**, or **Manual** (the client commits each turn) |
| Interruptions | Whether the caller can talk over the agent, and how many words count as an interruption (1 means any speech). When it is off, what the caller says while the agent answers is ignored |
| Fillers | Short phrases said while the caller waits for a slow answer or tool, in the order you list them: the first after a delay, then the next every few seconds while the wait goes on. Each turn starts again from the first, so a list like "Hmm.", "One moment.", "Still checking." escalates with the wait |
| Session | The longest call, when silence ends it, the message spoken when the agent cannot answer, and which settings a caller may change |

Voice is set **per agent version**. Both deployments must be in the agent's project and running;
the panel names the field it refuses, and warns (without refusing) when a deployment is not running
yet or its voice list cannot be read.

## Connect

**Use this agent** on the agent's page shows both addresses below for a version with voice on, with
Python and browser snippets.

### OpenAI Realtime clients — `/v1/realtime`

```
wss://gateway.bud.studio/v1/realtime?model=prompt:<agent-name>
wss://gateway.bud.studio/v1/realtime?model=prompt:<agent-name>:v<version>
```

Without a version the agent's default version answers. Authenticate with a project API key (or a Bud
access token) in the `Authorization` header; a browser sends it as the WebSocket subprotocol
`openai-insecure-api-key.<key>` beside `realtime`. Audio is PCM16, 24 kHz, mono, both ways.

The session behaves like an OpenAI Realtime session: send `input_audio_buffer.append`; the gateway
detects the end of each turn and answers. You will see `input_audio_buffer.speech_started` /
`speech_stopped`, `conversation.item.input_audio_transcription.completed`, `response.created`,
`response.output_audio.delta`, `response.output_audio_transcript.delta`, and `response.done`.
A tool the agent calls appears as an `mcp_call` output item (`response.mcp_call.in_progress` →
`completed` / `failed`).

You may also type: `conversation.item.create` with an `input_text` message, then `response.create`.
Typed and spoken turns share one conversation.

### Voice pipeline sessions — `/ws`

The [voice pipeline](/api-sdk/realtime/voice-pipeline-sessions) WebSocket takes an agent in place of
`conversation_config`:

```json theme={null}
{"type": "config", "agent": {"id": "support", "version": 2, "variables": {"customer": "Acme"}}}
```

Send microphone audio as binary frames (PCM16, 16 kHz by default); synthesized speech arrives as
binary frames. The agent's messages are `agent_response_started`, `agent_response_created`,
`assistant_transcript`, `agent_tool`, `agent_output` (an agent with structured output),
`agent_response_done`, `agent_truncated` and `agent_error`. Send `{"type": "agent_input", "text":
"..."}` to type a turn, and `{"type": "truncate", "audio_end_ms": 1234}` when your player stopped
early.

## What the caller can change

The agent owns its prompt, tools, deployments and instructions. A session may change only what the
agent's author allowed under **Callers may change**, when it starts: in the first `session.update`
on `/v1/realtime`, sent as soon as the socket opens (a session that hears nothing starts with the
agent's own settings 600 ms after connecting), or in the `config` message on `/ws`:

| Setting | `/v1/realtime` | `/ws` |
| - | - | - |
| Voice | `session.audio.output.voice` | `tts_config.voice_id` |
| Speaking speed | `session.audio.output.speed` | `tts_config.speaking_rate` |
| Language | `session.audio.input.transcription.language` | `stt_config.language` |
| Turn-taking eagerness | `session.audio.input.turn_detection.eagerness`: `low`, `medium`, `high` or `auto` | `stt_config.turn_detection.threshold`: 0.3 answers sooner, 0.7 lets the caller pause |

A setting the agent does not let callers change keeps the agent's value. `/v1/realtime` says so with
an `event_not_allowed` error naming the field (`invalid_value` for an unknown eagerness); `/ws` with a
`config_warning` (`deployment_setting_not_applied`) whose message names the setting. Eagerness applies to an agent with semantic turn taking: a manual or silence agent
keeps its own, and an agent never starts answering before the turn ends, whatever `eager` asks.
Everything else is refused with `event_not_allowed` and the session continues: `instructions`,
`tools`, `response.instructions`, another prompt id or version, an `api_key` or a `model` for either
leg.

## Interruptions and history

When the caller talks over the agent, the gateway stops the speech at once, cancels the agent's run,
and remembers **only what the caller heard**: the next turn's history holds the spoken part of the
interrupted answer followed by `[interrupted]`, never the unheard rest. A GA client that tracks its
own playback can make this exact with `conversation.item.truncate` (`audio_end_ms`); the gateway
answers `conversation.item.truncated` either way.

## Errors

| Where | Code | Why |
| - | - | - |
| Handshake (HTTP 404) | `model_not_found` | No such agent, or your key's project cannot reach it |
| Handshake (HTTP 404) | `agent_not_voice_enabled` | The agent (version) has no enabled voice |
| Handshake (HTTP 401) | `client_secret_unsupported` | Ephemeral client secrets cannot open an agent session; use a key or an access token |
| `error` event | `voice_leg_unavailable`, `stt_not_streaming` | A voice deployment is not running, or its transcription cannot stream |
| `error` event | `missing_variables`, `invalid_variables` | The agent's prompt needs variables (`session.prompt.variables` / `config.agent.variables`) |
| `error` event | `rate_limit_exceeded`, `upstream_error`, `upstream_unavailable` | A turn failed; the caller hears the agent's "cannot answer" message and the call continues |
| `error` event | `auth_expired` | The access token expired; send a fresh one (`bud.session.auth` / `auth`) |
| Close **1008** | `session_revoked` | The key was revoked, access removed, the agent's voice turned off, or a deployment unpublished (checked every 30 s) |
| Close **1000** | `idle_timeout`, `session_limit` | Silence ended the call, or it reached the agent's longest call |

## Limits

* One agent turn at a time per session; a new turn interrupts the previous one.
* An agent that needs a human approval mid-turn speaks its approval message instead of acting.
* Turns are traced: each is the root of its own trace, continued by the gateway and the agent run,
  and the speech legs are metered and priced as their deployments.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.