Chat completions
POST /chat/completions accepts:
modelRequired. A model id from /models, or @preset/<slug>. Suffixes: :online (web search), :nitro (fastest provider), :floor (cheapest), :exacto (best tool calling), :thinking (reasoning on).presetA preset slug; request fields override it.messagesRequired. OpenAI format: text, image_url and file parts; tool messages.modelsUp to 2 fallback models, tried in order when the model fails.stream, stream_options.include_usageServer-sent events; a final chunk with usage and cost.max_tokens / max_completion_tokensCap on generated tokens; also bounds the balance a request needs.temperature, top_p, top_k, min_p, stop, seedSampling controls (support varies by model).presence_penalty, frequency_penalty, repetition_penalty, logit_biasToken penalties.tools, tool_choice, parallel_tool_callsFunction calling.response_formatjson_object or json_schema (structured outputs).reasoning, reasoning_effort, include_reasoningReasoning models.pluginsweb (search), response-healing (fix JSON), context-compression (fit long prompts). Each takes "enabled": false.transforms["middle-out"] — same as context-compression; [] turns it off.safety_identifier, prompt_cache_keyEnd-user id (also reported in activity) and OpenAI cache routing hint.providerProvider routing preferences (see below).logprobs, top_logprobs, modalities, prediction, userPassed through where supported.Unknown parameters are ignored and listed in the x-ultragpt-ignored-params header; n must be 1. Every response has our own id (also in x-ultragpt-generation-id), the model that answered, and usage.cost — the USD charged to your balance.