Skip to main content
POST

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
model
string

The model ID used to generate the response, such as gpt-5.6-sol.

input

Text, image, audio, or file input provided to the model. It can be a string directly or an array of input messages.

instructions
string | null

System or developer directive that inserts into the model context. When using previous_response_id, instructions from the previous response are not automatically inherited.

previous_response_id
string | null

The unique ID of the previous Response, used for multi-turn conversations; cannot be used at the same time as conversation.

conversation

The session to which this response belongs. Session items are automatically added to this context, and input and output items are also written to the session.

background
boolean | null

Whether to run model responses in the background.

include
enum<string>[] | null

Specifies additional data that needs to be included in the response.

Available options:
web_search_call.action.sources,
code_interpreter_call.outputs,
computer_call_output.output.image_url,
file_search_call.results,
message.input_image.image_url,
message.output_text.logprobs,
reasoning.encrypted_content
max_output_tokens
integer | null

The maximum number of tokens allowed in the generated response, including visible output and inference tokens.

Required range: x >= 1
max_tool_calls
integer | null

The maximum number of total built-in tool calls that can be processed in a response.

Required range: x >= 0
metadata
object | null

Up to 16 key-value pair metadata; keys can be up to 64 characters and values can be up to 512 characters.

moderation
object | null

Configuration to perform auditing on request input and build output.

parallel_tool_calls
boolean | null

Whether to allow models to call tools in parallel.

prompt
object | null

Prompt word template and its variables.

prompt_cache_key
string | null

Stable key used to improve cache hit rate for similar requests; replaces user field.

prompt_cache_options
object

Prompt word caching options.

prompt_cache_retention
enum<string> | null

Deprecated, use prompt_cache_options.ttl.

Available options:
in_memory,
24h
reasoning
object | null

Inference model configuration.

safety_identifier
string | null

A stable end-user identifier to assist in detecting abuse; it is recommended to use a hash of the username or email address.

Maximum string length: 64
service_tier
enum<string> | null

The requested service level.

Available options:
auto,
default,
flex,
scale,
priority,
fast,
ultrafast
store
boolean | null

Whether to store the generated Response for subsequent retrieval through the API.

stream
boolean | null

Whether to return events as SSE streaming.

stream_options
object | null

Streaming option, only set when stream is true.

temperature
number | null

Sampling temperature; usually adjusted with top_p alternatively.

Required range: 0 <= x <= 2
top_p
number | null

Kernel sampling threshold; usually adjusted alternatively with temperature.

Required range: 0 <= x <= 1
text
object

Text or structured JSON output configuration.

tools
object[]

Tools that the model can call when generating responses, including function, web_search, file_search, code_interpreter, computer, MCP, etc.

tool_choice

Controls how the model selects tools.

Available options:
none,
auto,
required
top_logprobs
integer | null

The maximum number of high-probability tokens returned for each output position.

Required range: 0 <= x <= 20
truncation
enum<string> | null

Deprecated. The truncation strategy for input when it exceeds the context window.

Available options:
auto,
disabled
user
string

Deprecated, use safety_identifier and prompt_cache_key.

context_management
object[] | null

Context management configuration.

access_programs
object

Request a domain access plan to use.

Response

200 - application/json
id
string
required

The unique ID of the Response.

object
enum<string>
required

Object type, fixed to response.

Available options:
response
created_at
integer
required

Unix timestamp of the creation time of the Response, in seconds.

model
string
required

The actual model ID used.

output
object[]
required

A list of outputs generated by the model, the order and number of which depends on the model response.

completed_at
integer

Unix timestamp of response completion time; only appears if status is completed.

status
enum<string>

Response build status.

Available options:
completed,
failed,
in_progress,
cancelled,
queued,
incomplete
output_text
string

SDK convenience field: The concatenation of all text content in the output.

error
object

Error object when generation fails.

incomplete_details
object

Reason for incomplete response.

instructions

The system or developer directive that applies to this response.

conversation
object

The session to which this response belongs.

previous_response_id
string

The ID of the previous Response.

metadata
object

Associated metadata.

usage
object

Word usage statistics.

parallel_tool_calls
boolean

Whether to allow parallel tool calls.

temperature
number

The sampling temperature used by this response.

top_p
number

The kernel sampling threshold used by this response.

max_output_tokens
integer

The maximum number of output tokens requested.

max_tool_calls
integer

The maximum number of tool calls requested to be set.

service_tier
string

The actual service level used.

background
boolean

Whether to run in the background.

store
boolean

Whether to store the response.

safety_identifier
string

The security identifier requested to be used.