Skip to main content
POST

Headers

Body

Core request fields

Sampling

Reasoning & truncation

Tools

Persistence & background mode

Streaming

Identity

Conversations and continuation

There are two ways to continue a thread, and they are mutually exclusive — sending both returns 400 (`previous_response_id` cannot be used in conjunction with `conversation`.).

previous_response_id

Continue from one specific turn. The server walks that turn’s chain and prepends the prior turns. Use it to branch deliberately from an older turn.

conversation

Continue the thread. The server resolves the conversation’s latest turn for you, so a client that only keeps one handle never has to track response ids at all.

The response always carries the conversation

Every response echoes the effective conversation as an object:
This is present on every envelope shape — the completed response, the 200 returned by background: true, and the paused/withheld envelopes described in Response Format — and a later GET /v1/responses/{id} reports the same value.

Ids are minted for you

A call that supplies neither conversation nor previous_response_id is given a server-minted conv_* id, stamped on the stored response and echoed back, so the client has a reusable handle without inventing one:
This is a deliberate divergence from OpenAI, which never mints conversation ids (it has a separate POST /v1/conversations resource that Bud does not implement). Operators can turn minting off with RESPONSES_AUTO_CONVERSATION=false, after which no id is returned unless the client sends one.

The echoed parent is the one that was stored

previous_response_id on the response body is the parent the turn was actually stored with, which is not always the one you sent: when two turns of the same conversation arrive together, the server serializes them and chains the second onto the first rather than letting the thread fork. A turn continued by conversation therefore reports the parent the server resolved, and GET /v1/responses/{id} reports the same one.

What store: false guarantees

store: false means the turn’s content is not retained after the response is returned:
  • an ordinary run’s row is removed once the answer is delivered, so GET /v1/responses/{id} returns 404 afterwards;
  • a run that a governance policy touched keeps a content-free skeleton instead of being deleted — its input, output, instructions and prompt_variables are emptied, but the row survives so the record of an approval a human actually gave is not destroyed with it;
  • a crash between the answer and the cleanup is swept afterwards, so content cannot be retained indefinitely by a badly-timed restart.
Deleting a turn with store: false never truncates a thread it was chained into — see Delete Response.

Notes on deployment-level configuration

Several fields above (max_output_tokens, max_tool_calls, reasoning.effort, truncation, store, background, instructions) have per-deployment defaults and caps configured in the deployment’s Settings → Responses API tab. Behaviour summary:
  • Limits clamp — max_output_tokens / max_tool_calls above the deployment cap return 200 OK with incomplete_details.reason set. The chain-depth limit is enforced as a hard 400 chain_too_deep.
  • Capability gates — background: true against a deployment with background mode disabled returns 400 background_not_allowed. Zero-data-retention deployments force store = false regardless of request.
  • Defaults pre-fill — reasoning.effort, truncation, instructions from the deployment apply when the request omits them. Per-request values always override.
  • Server tools auto-attach — web_search and web_fetch enabled at the deployment level attach to every call without the client declaring them in tools[]. The model decides whether to invoke them.
See the Deployment Settings guide for how to configure these.

Supported Providers

/v1/responses is provider-agnostic — it works with any chat-capable model deployed through Bud, including models from providers that don’t offer a Responses API of their own.

OpenAI / Azure OpenAI

GPT-4o, o-series and Azure deployments.

Anthropic

Claude models, including extended-thinking (reasoning) variants.

Google / Mistral

Gemini and Mistral chat models.

Moonshot / DeepSeek

Kimi K2.5 / K2-thinking and DeepSeek-Reasoner, with reasoning-token accounting preserved.

Self-hosted open weights

Any vLLM / SGLang deployment (Llama, Qwen, etc.).

…and others

Any current or future model available on your Bud deployment.
To serve /v1/responses, a deployment needs the Responses API turned on in its Settings tab (this also requires /v1/chat/completions to be enabled on the same deployment).Model-specific behaviour still applies: e.g. reasoning.effort only affects reasoning-capable models.