Make an agent a voice agent
In the agent builder, click the microphone on the Input card (it opens Settings → Voice → Listen) or on the Output card (Voice → Speak), and turn on Voice agent. The same Voice tab is in the Settings menu beside Runtime and Model, and one Save keeps all three.
Voice is set per agent version. Both deployments must be in the agent’s project and running;
the panel names the field it refuses, and warns (without refusing) when a deployment is not running
yet or its voice list cannot be read.
Connect
Use this agent on the agent’s page shows both addresses below for a version with voice on, with Python and browser snippets.OpenAI Realtime clients — /v1/realtime
Authorization header; a browser sends it as the WebSocket subprotocol
openai-insecure-api-key.<key> beside realtime. Audio is PCM16, 24 kHz, mono, both ways.
The session behaves like an OpenAI Realtime session: send input_audio_buffer.append; the gateway
detects the end of each turn and answers. You will see input_audio_buffer.speech_started /
speech_stopped, conversation.item.input_audio_transcription.completed, response.created,
response.output_audio.delta, response.output_audio_transcript.delta, and response.done.
A tool the agent calls appears as an mcp_call output item (response.mcp_call.in_progress →
completed / failed).
You may also type: conversation.item.create with an input_text message, then response.create.
Typed and spoken turns share one conversation.
Voice pipeline sessions — /ws
The voice pipeline WebSocket takes an agent in place of
conversation_config:
agent_response_started, agent_response_created,
assistant_transcript, agent_tool, agent_output (an agent with structured output),
agent_response_done, agent_truncated and agent_error. Send {"type": "agent_input", "text": "..."} to type a turn, and {"type": "truncate", "audio_end_ms": 1234} when your player stopped
early.
What the caller can change
The agent owns its prompt, tools, deployments and instructions. A session may change only what the agent’s author allowed under Callers may change, when it starts: in the firstsession.update
on /v1/realtime, sent as soon as the socket opens (a session that hears nothing starts with the
agent’s own settings 600 ms after connecting), or in the config message on /ws:
A setting the agent does not let callers change keeps the agent’s value.
/v1/realtime says so with
an event_not_allowed error naming the field (invalid_value for an unknown eagerness); /ws with a
config_warning (deployment_setting_not_applied) whose message names the setting. Eagerness applies to an agent with semantic turn taking: a manual or silence agent
keeps its own, and an agent never starts answering before the turn ends, whatever eager asks.
Everything else is refused with event_not_allowed and the session continues: instructions,
tools, response.instructions, another prompt id or version, an api_key or a model for either
leg.
Interruptions and history
When the caller talks over the agent, the gateway stops the speech at once, cancels the agent’s run, and remembers only what the caller heard: the next turn’s history holds the spoken part of the interrupted answer followed by[interrupted], never the unheard rest. A GA client that tracks its
own playback can make this exact with conversation.item.truncate (audio_end_ms); the gateway
answers conversation.item.truncated either way.
Errors
Limits
- One agent turn at a time per session; a new turn interrupts the previous one.
- An agent that needs a human approval mid-turn speaks its approval message instead of acting.
- Turns are traced: each is the root of its own trace, continued by the gateway and the agent run, and the speech legs are metered and priced as their deployments.