Chat completions

POST /chat/completions accepts:

modelRequired. A model id from /models, or @preset/<slug>. Suffixes: :online (web search), :nitro (fastest provider), :floor (cheapest), :exacto (best tool calling), :thinking (reasoning on).
presetA preset slug; request fields override it.
messagesRequired. OpenAI format: text, image_url and file parts; tool messages.
modelsUp to 2 fallback models, tried in order when the model fails.
stream, stream_options.include_usageServer-sent events; a final chunk with usage and cost.
max_tokens / max_completion_tokensCap on generated tokens; also bounds the balance a request needs.
temperature, top_p, top_k, min_p, stop, seedSampling controls (support varies by model).
presence_penalty, frequency_penalty, repetition_penalty, logit_biasToken penalties.
tools, tool_choice, parallel_tool_callsFunction calling.
response_formatjson_object or json_schema (structured outputs).
reasoning, reasoning_effort, include_reasoningReasoning models.
pluginsweb (search), response-healing (fix JSON), context-compression (fit long prompts). Each takes "enabled": false.
transforms["middle-out"] — same as context-compression; [] turns it off.
safety_identifier, prompt_cache_keyEnd-user id (also reported in activity) and OpenAI cache routing hint.
providerProvider routing preferences (see below).
logprobs, top_logprobs, modalities, prediction, userPassed through where supported.

Unknown parameters are ignored and listed in the x-ultragpt-ignored-params header; n must be 1. Every response has our own id (also in x-ultragpt-generation-id), the model that answered, and usage.cost — the USD charged to your balance.