Skip to main content
A voice agent is an agent you can talk to. The agent stays exactly what it is — its prompt, model, tools, skills and governance run on every turn — and two deployments in its project give it ears and a voice: a speech-to-text deployment that transcribes the caller, and a text-to-speech deployment that speaks the agent’s answers. The gateway takes the turns, streams the answer into speech sentence by sentence, and stops talking the moment the caller talks over it.

Make an agent a voice agent

In the agent builder, click the microphone on the Input card (it opens Settings → Voice → Listen) or on the Output card (Voice → Speak), and turn on Voice agent. The same Voice tab is in the Settings menu beside Runtime and Model, and one Save keeps all three. Voice is set per agent version. Both deployments must be in the agent’s project and running; the panel names the field it refuses, and warns (without refusing) when a deployment is not running yet or its voice list cannot be read.

Connect

Use this agent on the agent’s page shows both addresses below for a version with voice on, with Python and browser snippets.

OpenAI Realtime clients — /v1/realtime

Without a version the agent’s default version answers. Authenticate with a project API key (or a Bud access token) in the Authorization header; a browser sends it as the WebSocket subprotocol openai-insecure-api-key.<key> beside realtime. Audio is PCM16, 24 kHz, mono, both ways. The session behaves like an OpenAI Realtime session: send input_audio_buffer.append; the gateway detects the end of each turn and answers. You will see input_audio_buffer.speech_started / speech_stopped, conversation.item.input_audio_transcription.completed, response.created, response.output_audio.delta, response.output_audio_transcript.delta, and response.done. A tool the agent calls appears as an mcp_call output item (response.mcp_call.in_progress → completed / failed). You may also type: conversation.item.create with an input_text message, then response.create. Typed and spoken turns share one conversation.

Voice pipeline sessions — /ws

The voice pipeline WebSocket takes an agent in place of conversation_config:
Send microphone audio as binary frames (PCM16, 16 kHz by default); synthesized speech arrives as binary frames. The agent’s messages are agent_response_started, agent_response_created, assistant_transcript, agent_tool, agent_output (an agent with structured output), agent_response_done, agent_truncated and agent_error. Send {"type": "agent_input", "text": "..."} to type a turn, and {"type": "truncate", "audio_end_ms": 1234} when your player stopped early.

What the caller can change

The agent owns its prompt, tools, deployments and instructions. A session may change only what the agent’s author allowed under Callers may change, when it starts: in the first session.update on /v1/realtime, sent as soon as the socket opens (a session that hears nothing starts with the agent’s own settings 600 ms after connecting), or in the config message on /ws: A setting the agent does not let callers change keeps the agent’s value. /v1/realtime says so with an event_not_allowed error naming the field (invalid_value for an unknown eagerness); /ws with a config_warning (deployment_setting_not_applied) whose message names the setting. Eagerness applies to an agent with semantic turn taking: a manual or silence agent keeps its own, and an agent never starts answering before the turn ends, whatever eager asks. Everything else is refused with event_not_allowed and the session continues: instructions, tools, response.instructions, another prompt id or version, an api_key or a model for either leg.

Interruptions and history

When the caller talks over the agent, the gateway stops the speech at once, cancels the agent’s run, and remembers only what the caller heard: the next turn’s history holds the spoken part of the interrupted answer followed by [interrupted], never the unheard rest. A GA client that tracks its own playback can make this exact with conversation.item.truncate (audio_end_ms); the gateway answers conversation.item.truncated either way.

Errors

Limits

  • One agent turn at a time per session; a new turn interrupts the previous one.
  • An agent that needs a human approval mid-turn speaks its approval message instead of acting.
  • Turns are traced: each is the root of its own trace, continued by the gateway and the agent run, and the speech legs are metered and priced as their deployments.