Skip to main content
POST

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <token>, where <token> is your auth token.

Body

application/json
model
string
required

The model ID used to generate the reply, such as gpt-5.6-sol or o3.

messages
object[]
required

List of conversation messages as of current date.

Minimum array length: 1
audio
object | null

Audio output parameters. Required when modalities contains audio.

temperature
number | null

Sampling temperature. Higher values ​​make the output more random; usually adjusted with top_p alternatively.

Required range: 0 <= x <= 2
top_p
number | null

Kernel sampling threshold; usually adjusted alternatively with temperature.

Required range: 0 <= x <= 1
n
integer | null

The number of candidate responses generated for each input.

Required range: x >= 1
stream
boolean | null

Whether to return incremental results in SSE streaming.

stream_options
object | null

Additional configuration for streaming output, only takes effect when stream is true.

max_completion_tokens
integer | null

The maximum number of generated tokens, including visible output and inference tokens.

Required range: x >= 1
max_tokens
integer | null

Deprecated, use max_completion_tokens instead.

Required range: x >= 1
frequency_penalty
number | null

Frequency penalty. Positive values ​​reduce duplicate expressions.

Required range: -2 <= x <= 2
presence_penalty
number | null

There are penalties. This is a great time to encourage models to introduce new content.

Required range: -2 <= x <= 2
logit_bias
object | null

Specifies the generation tendency of the word element. The attribute is named token ID and has values ​​from -100 to 100.

logprobs
boolean | null

Whether to return the log probability of the output token.

top_logprobs
integer | null

The number of high-probability candidate tokens returned for each output position; logprobs needs to be enabled at the same time.

Required range: 0 <= x <= 20
stop

Stop sequence, some new models do not support it.

response_format
object

Output format configuration.

tools
object[]

A list of callable tools provided to the model.

Function tools.

tool_choice

Controls whether the model calls the tool, automatically selects the tool, or forces the specified tool to be called.

Available options:
none,
auto,
required
parallel_tool_calls
boolean

Whether the model is allowed to call multiple tools in parallel.

reasoning_effort
enum<string> | null

Inference strength settings for trade-offs between speed, cost, and inference depth.

Available options:
none,
minimal,
low,
medium,
high,
xhigh,
max
verbosity
enum<string> | null

Output verbosity.

Available options:
low,
medium,
high
seed
integer | null

A random seed used to try to make the output reproducible.

service_tier
enum<string> | null

The requested service level.

Available options:
auto,
default,
flex,
scale,
priority,
fast
store
boolean | null

Whether to store the output of this request.

metadata
object | null

Additional key-value pair metadata.

modalities
enum<string>[] | null

Specify the output type.

Available options:
text,
audio
prediction
object | null

Prediction output configuration.

prompt_cache_key
string | null

Stable key for hint word cache.

prompt_cache_retention
enum<string> | null

The length of time the prompt word cache is retained.

Available options:
in_memory,
24h
prompt_cache_options
object

Prompt word caching options.

safety_identifier
string | null

A stable end-user identifier used to assist in detecting abuse and should not be directly personally identifiable information.

user
string

Old version of end user identification, it is recommended to use safety_identifier.

web_search_options
object

Web search tool configuration.

function_call
deprecated

Deprecated, use tool_choice.

Available options:
none,
auto
functions
object[]
deprecated

Deprecated, please use tools.

moderation
object | null

Configuration to perform auditing on request input and build output.

Response

200 - application/json

Chat Completions non-streaming responses (chat.completion). "Generate via JSON etc. → JSON Schema" direct import for Apifox.

id
string
required

The unique ID of this chat completion.

object
enum<string>
required

Object type, fixed to chat.completion.

Available options:
chat.completion
created
integer
required

Unix timestamp of the time the response was created, in seconds.

model
string
required

The actual model ID used to generate the reply.

choices
object[]
required

A list of candidate responses generated by the model.

usage
object

Word usage statistics for this request.

service_tier
string | null

The actual service level used to handle this request.

system_fingerprint
string | null

The backend configuration fingerprint for the model to run on.

metadata
object

Metadata associated with the request.