Skip to main content

Overview

The Responses API provides a next-generation interface for complex AI interactions, supporting:
  • Prompt-based execution: Execute versioned prompt templates with variable substitution
  • MCP tool integration: Access Model Context Protocol tools for extended functionality
  • Structured outputs: JSON schema-validated responses for reliable data extraction
  • Array-based outputs: Multiple output types (messages, tool calls, reasoning, MCP tool lists)
  • Multi-turn conversations with context preservation
  • Parallel tool/function calling
  • Multimodal inputs (text, image, audio)
  • Reasoning model capabilities
  • Streaming responses

Endpoints

Authentication

Create Response

Generate AI responses with advanced conversational features.

Request Format

Endpoint: POST /v1/responses Headers:
  • Authorization: Bearer YOUR_API_KEY (required)
  • Content-Type: application/json (required)
Request Body:

Parameters

Prompt Input Format

Multimodal Input Format

Response Format

The response contains an array-based output field with multiple item types:

Output Item Types

The output array can contain multiple types of items:

Text Messages

MCP Tool Lists

MCP Tool Calls

Function Tool Calls

Reasoning Items

Streaming Response Format

When streaming is enabled, responses are returned as Server-Sent Events (SSE) with the following format:

Event Lifecycle

1. Initial Events
2. MCP Tool List Events (if MCP tools are configured)
3. Text Output Events
4. Reasoning Events (for thinking/reasoning models)
5. MCP Tool Call Events
6. Function Tool Call Events
7. Completion Event
8. Error Event (on failure)
9. Cancelled Event (when the stream is cancelled)
A stream ends on exactly one of these three terminal events — response.completed, response.failed or response.incomplete. The cancelled frame uses the response.incomplete event type because the stream-event union has no response.cancelled member; the response object it carries says cancelled, matching what GET /v1/responses/{id} reports.

Key Event Fields

  • sequence_number: Monotonically increasing counter for event ordering
  • output_index: Position in the output array (0-indexed)
  • item_id: Unique identifier for the specific item being streamed
  • content_index: Position within the content array (for messages)
  • summary_index: Position within the summary array (for reasoning)

Prompt-Based Execution

Execute pre-configured prompt templates using the prompt parameter: Request Example:
Fields:
  • prompt.id (required) - Template identifier
  • prompt.variables (optional) - Variable substitutions
  • prompt.version (optional) - Template version (defaults to default version)
  • input (optional) - Unstructured user input
Prompt Configuration (via UI or API): Users can pre-configure prompts with:
  • Model deployment and settings (temperature, max_tokens, top_p, etc.)
  • System prompt with Jinja2 template support
  • Conversation messages and context with Jinja2 template support
  • MCP tools (filesystem, web access, custom tools)
  • Input/output schemas for structured data
  • Validation rules and retry limits
  • Streaming configuration

Governance outcomes

A prompt or deployment can carry a governance policy. Two of its outcomes are answers a client has to handle, and both are 200 — not errors:

Paused for human approval

The policy asked a person to approve the call before it runs (or before an already-computed result is released). The response is:
The pause is identified by the mcp_approval_request item, never by the status. The wire status is in_progress because requires_action is not a member of OpenAI’s closed ResponseStatus enum and would fail a strict client’s response validation.
The id on that item is the approval-request id. Once a human decides, the client continues the run by sending the decision back as an input item, threaded to the paused response:
previous_response_id is what binds the decision to the escalation, and it must name a response the calling key owns — a foreign id is a 404, the same ownership protection GET uses. Polling the paused id with GET keeps returning the same envelope until the decision is made; a paused run is never collected by the retention window.

Withheld by policy

The policy blocked the content outright. The run ends incomplete, with the policy’s guidance as the assistant text rather than an error body:
Both outcomes are stored like any other turn: they carry their conversation, they report the tokens they spent getting there, and a later GET /v1/responses/{id} returns the same shape.

Retrieve Response

Get details of a specific response. Endpoint: GET /v1/responses/{response_id}

Response Format

Returns the same format as the create response endpoint.

Delete Response

Remove a response from the system. Endpoint: DELETE /v1/responses/{response_id}

Response Format

Cancel Response

Cancel an in-progress response generation. Endpoint: POST /v1/responses/{response_id}/cancel

Response Format

List Input Items

Retrieve the input conversation history for a response. Endpoint: GET /v1/responses/{response_id}/input_items

Response Format

Usage Examples

Basic Response

Prompt Execution

Multi-turn Conversation

With Tool Calling

Python Example

JavaScript Example

Advanced Features

Reasoning Models

Enable step-by-step reasoning:

Parallel Tool Calling

The API supports calling multiple tools in parallel:

Conversation Context

Maintain context across multiple interactions:

Error Responses

400 Bad Request

401 Unauthorized

429 Rate Limit

Best Practices

  • Conversation Management: Use previous_response_id for coherent multi-turn conversations
  • Tool Design: Create focused, single-purpose tools for better reliability
  • Streaming: Use streaming for long responses to improve user experience
  • Error Handling: Implement robust retry logic for transient failures
  • Metadata: Use metadata to track conversations and user sessions
  • Context Window: Be mindful of token limits when building long conversations
  • Parallel Tools: Leverage parallel tool calling for independent operations
  • Prompt Templates: Design reusable prompt templates with clear variable names for maintainability
  • Variable Management: Use descriptive variable names and provide defaults where appropriate
  • Version Control: Use prompt versioning to iterate on prompts without breaking existing integrations

Limitations

  • Some advanced retrieval features may not be fully implemented
  • Response management endpoints have limited functionality
  • Conversation history is maintained only through previous_response_id chaining
  • Maximum context window depends on the model used
  • Prompt templates must be pre-configured before use
  • Maximum context window depends on the model used, including when configured in a prompt template.

Supported providers

OpenAI

Next-generation Responses API with full support for advanced conversational features, multi-turn interactions, and parallel tool calling.