Documentation
OpenAI- and Anthropic-compatible. Point any SDK at https://api.ultragpt.pro/v1.
- Create an API key and store it as
ULTRAGPT_API_KEY. - Add credits to your balance (shared with the UltraGPT app).
- Install the SDK (
pip install openai/npm i openai) and call a model:
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"messages": [{ "role": "user", "content": "Hello!" }]
}'Endpoints
/chat/completionsChat completion (OpenAI format), streaming or not./messagesAnthropic Messages API (Anthropic SDKs, Claude Code)./messages/count_tokensEstimate prompt tokens (Anthropic format)./responsesOpenAI Responses API, with stored conversations./responses/{id}A stored response (also DELETE, /input_items)./completionsLegacy text completions./embeddingsText embeddings./images/generationsGenerate images./audio/speechText to speech./audio/transcriptionsSpeech to text (also /audio/translations)./filesUpload a batch input file (also GET, DELETE, /content)./batchesRun up to 5,000 requests asynchronously, outside your rate limits (also GET, /cancel)./rerankRerank documents against a query (Cohere / Voyage / Qwen)./moderationsClassify text and images for harmful content. Free./videosStart a video generation (async; also GET, GET /content, DELETE)./realtime?model=Realtime speech-to-speech over WebSocket (OpenAI Realtime protocol)./search/{engine}Standalone web search (Tavily, Exa, Perplexity…), per query./modelsModels your key can use (?type=, ?supported_parameters=, ?input_modalities= filters; also /models/count)./presetsPresets of your workspace (writes need a management key)./keysManagement keys: list, create, update, revoke keys./activityDaily usage per model and endpoint, last 30 days./generations/lookupCost and stats of up to 1000 requests at once./keyThe calling key: limits, spend./creditsCredit balance and lifetime totals./generation?id=Tokens, cost and latency of one request.Authentication
Send your key as Authorization: Bearer <key> or x-api-key: <key> (Anthropic SDKs). Keys start with ugpt_live_ or ugpt_test_; test keys have low rate limits but bill real credits.
Harden keys on the API keys page: a monthly spend cap, a model allowlist, a lower RPM, an expiry, a read-only scope (no billed calls), an IP / CIDR allowlist, and — for keys used in a browser — allowed domains checked against Origin / Referer. Rotate a key to replace its secret with an optional grace period for the old one. A revoked key stops working immediately.
SDKs & tools
Anything that speaks the OpenAI API works with just a base URL and key:
import { createOpenAI } from '@ai-sdk/openai'
import { generateText } from 'ai'
const ultragpt = createOpenAI({ baseURL: 'https://api.ultragpt.pro/v1', apiKey: process.env.ULTRAGPT_API_KEY })
const { text } = await generateText({ model: ultragpt.chat('grok-4.7'), prompt: 'Hello!' })Anthropic SDK & Claude Code
POST /v1/messages implements the Anthropic Messages API — system prompts, images, documents, tool use, thinking, cache_control and streaming events — for every chat model. The web_search server tool maps to our web search.
curl https://api.ultragpt.pro/v1/messages \
-H "x-api-key: $ULTRAGPT_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"max_tokens": 1024,
"messages": [{ "role": "user", "content": "Hello!" }]
}'Claude Code — set these before running claude:
export ANTHROPIC_BASE_URL="https://api.ultragpt.pro"
export ANTHROPIC_AUTH_TOKEN="$ULTRAGPT_API_KEY"
export ANTHROPIC_MODEL="grok-4.7"
export ANTHROPIC_SMALL_FAST_MODEL="grok-4.7"
claudeChat completions
POST /chat/completions accepts:
modelRequired. A model id from /models, or @preset/<slug>. Suffixes: :online (web search), :nitro (fastest provider), :floor (cheapest), :exacto (best tool calling), :thinking (reasoning on).presetA preset slug; request fields override it.messagesRequired. OpenAI format: text, image_url and file parts; tool messages.modelsUp to 2 fallback models, tried in order when the model fails.stream, stream_options.include_usageServer-sent events; a final chunk with usage and cost.max_tokens / max_completion_tokensCap on generated tokens; also bounds the balance a request needs.temperature, top_p, top_k, min_p, stop, seedSampling controls (support varies by model).presence_penalty, frequency_penalty, repetition_penalty, logit_biasToken penalties.tools, tool_choice, parallel_tool_callsFunction calling.response_formatjson_object or json_schema (structured outputs).reasoning, reasoning_effort, include_reasoningReasoning models.pluginsweb (search), response-healing (fix JSON), context-compression (fit long prompts). Each takes "enabled": false.transforms["middle-out"] — same as context-compression; [] turns it off.safety_identifier, prompt_cache_keyEnd-user id (also reported in activity) and OpenAI cache routing hint.providerProvider routing preferences (see below).logprobs, top_logprobs, modalities, prediction, userPassed through where supported.Unknown parameters are ignored and listed in the x-ultragpt-ignored-params header; n must be 1. Every response has our own id (also in x-ultragpt-generation-id), the model that answered, and usage.cost — the USD charged to your balance.
Responses API
POST /responses supports input items, instructions, function tools, the web_search tool, text.format (JSON schema) and reasoning. With store (default true) the conversation is kept 30 days, encrypted, so previous_response_id works; send store: false to opt out.
curl https://api.ultragpt.pro/v1/responses \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "grok-4.7", "input": "What is an API?" }'Streaming
With stream: true you receive chat.completion.chunk events ending in data: [DONE]. Add stream_options: {"include_usage": true} for a final chunk with usage and cost. If you disconnect mid-stream the generation still completes and is billed.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-N -d '{
"model": "grok-4.7",
"stream": true,
"messages": [{ "role": "user", "content": "Write a haiku about APIs" }]
}'Tools & JSON
Function calling and structured outputs follow the OpenAI format. Requests with tools to a model that cannot call tools fail fast with 400 unsupported_parameter.
response = client.chat.completions.create(
model="grok-4.7",
messages=[{"role": "user", "content": "Extract: Ada Lovelace, born 1815"}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "person",
"strict": True,
"schema": {
"type": "object",
"properties": {"name": {"type": "string"}, "born": {"type": "integer"}},
"required": ["name", "born"],
"additionalProperties": False,
},
},
},
)Reasoning
Use reasoning_effort, or the reasoning object ({ effort, max_tokens, exclude }). Reasoning text comes back as message.reasoning (delta.reasoning when streaming); reasoning tokens are billed as output.
response = client.chat.completions.create(
model="grok-4.7",
messages=[{"role": "user", "content": "How many r's are in strawberry?"}],
reasoning_effort="high", # or extra_body={"reasoning": {"max_tokens": 4000, "exclude": False}}
)
print(response.choices[0].message.reasoning) # the model's reasoning, when returned
print(response.usage.completion_tokens_details.reasoning_tokens)Images & files in
Send image_url parts with an https URL or a base64 data URL (max 20 images, 20MB each) to models with the vision capability, and PDFs as file parts. Other models reject images with 400. The request body is limited to 10MB — send large images by URL.
Web search
Add plugins: [{ "id": "web" }] or append :online to the model id. We search the web for the last user message, give the model the results and return them as url_citation annotations. Each search is billed at the price on the Models page.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7:online",
"messages": [{ "role": "user", "content": "What happened in AI this week?" }]
}'Provider routing
Models served through OpenRouter accept a provider object: order, only, ignore, allow_fallbacks, sort (price / throughput / latency), require_parameters, data_collection, quantizations and zdr. You are always billed the real cost of the provider that answered.
response = client.chat.completions.create(
model="grok-4.7",
messages=messages,
extra_body={
"provider": {
"sort": "throughput", # or "price" / "latency"
"order": ["anthropic", "google-vertex"],
"allow_fallbacks": True,
"data_collection": "deny", # only providers that don't train on data
"zdr": True, # zero data retention
},
},
)Prompt caching
Cached input tokens bill at the (much lower) cached price shown on each model. OpenAI and DeepSeek models cache repeated prefixes automatically; for Anthropic and Gemini models mark the stable part with cache_control. usage.prompt_tokens_details.cached_tokens shows the hits.
# Mark a large, stable prefix as cacheable (Anthropic / Gemini models via OpenRouter;
# OpenAI and DeepSeek models cache automatically). Cached tokens bill at the cached price.
response = client.chat.completions.create(
model="grok-4.7",
messages=[
{"role": "system", "content": [
{"type": "text", "text": LONG_REFERENCE_DOCUMENT, "cache_control": {"type": "ephemeral"}},
]},
{"role": "user", "content": "Summarize section 3."},
],
)
print(response.usage.prompt_tokens_details.cached_tokens)Model fallbacks
Pass up to two backup models in models. If the primary fails on the provider side (outage, overload, timeout) the request is retried on the next — the response's model tells you which answered. Invalid requests (400) are not retried.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"models": ["<fallback-model-1>", "<fallback-model-2>"],
"messages": [{ "role": "user", "content": "Hello!" }]
}'Presets
Save a model, fallbacks, system prompt, parameters, provider routing, plugins and caching on the Presets page, then reference it from any text endpoint as model: "@preset/{slug}" (or "{model}@preset/{slug}", or "preset": "{slug}"). Request fields win over the preset; tools and plugins are merged. Edits create versions you can roll back — requests always use the latest, so you can change prompts without redeploying.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "@preset/support-bot", "messages": [{ "role": "user", "content": "Hi" }] }'
# List the workspace's presets
curl https://api.ultragpt.pro/v1/presets -H "Authorization: Bearer $ULTRAGPT_API_KEY"Response caching
Send X-UltraGPT-Cache: true (X-OpenRouter-Cache works too) and an identical request from the same key within the TTL (X-UltraGPT-Cache-TTL, 1–86400 s, default 300) is answered from cache instantly and for free: usage counters are 0 and nothing is billed. Works for streaming and non-streaming chat, messages, responses and completions. Responses carry x-ultragpt-cache-status (HIT / MISS), -age, -ttl and -source-id; X-UltraGPT-Cache-Clear: true refreshes the entry. Presets can turn caching on for every request.
curl -i https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-H "X-UltraGPT-Cache: true" \
-H "X-UltraGPT-Cache-TTL: 600" \
-d '{ "model": "grok-4.7", "messages": [{ "role": "user", "content": "Capital of France?" }] }'
# Second identical call within 10 minutes:
# x-ultragpt-cache-status: HIT x-ultragpt-cache-age: 12 usage.cost: 0
# Force a refresh: -H "X-UltraGPT-Cache-Clear: true"Plugins, variants & insurance
response-healing repairs malformed JSON (code fences, trailing commas, unquoted keys, missing brackets) on non-streamed json_object / json_schema requests. context-compression drops messages from the middle of the conversation (keeping system prompts and the latest turn, tool calls with their results) when a prompt exceeds the model's context window — on by default for models with 8K context or less; x-ultragpt-compressed-messages reports how many were dropped. Model suffixes pick routing: :nitro fastest provider, :floor cheapest, :exacto most reliable tool calling, :thinking reasoning on. Zero-completion insurance: a request that returns no tokens or ends in an error is not billed.
response = client.chat.completions.create(
model="grok-4.7:nitro", # :nitro fastest, :floor cheapest, :exacto best tool calling,
# :thinking reasoning on, :online web search
messages=long_conversation,
response_format={"type": "json_object"},
extra_body={
"plugins": [
{"id": "response-healing"}, # repair malformed JSON (non-streaming)
{"id": "context-compression"}, # drop middle messages if over the context window
],
"safety_identifier": "user-123", # your end user (shows in /activity?user=)
},
extra_headers={"X-Client-Request-Id": "order-42"}, # echoed back + stored on the generation
)Embeddings
POST /embeddings takes a string or up to 2,048 strings (or token arrays), with optional dimensions and encoding_format. Billed on input tokens.
curl https://api.ultragpt.pro/v1/embeddings \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "qwen3-embedding-8b", "input": "The quick brown fox" }'Image generation
POST /images/generations with prompt, n (1–4), size, quality and response_format (b64_json or url). Multimodal image models also accept image: [urls] for edits. Billed at the provider's reported cost.
curl https://api.ultragpt.pro/v1/images/generations \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "grok-imagine-image-2.0", "prompt": "A lighthouse at dawn, watercolor", "response_format": "url" }'Speech & transcription
POST /audio/speech returns audio (mp3, opus, aac, flac, wav, pcm) in one of the model's voices, priced per character. POST /audio/transcriptions (and /audio/translations) take a multipart file up to 25MB and return json, text, verbose_json, srt or vtt.
curl https://api.ultragpt.pro/v1/audio/speech \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "model": "ai-voice", "voice": "default", "input": "Hello from UltraGPT!" }' \
--output hello.mp3curl https://api.ultragpt.pro/v1/audio/transcriptions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-F model="audio-transcribe" \
-F file="@meeting.mp3"Batch API
Upload a JSONL of up to 5,000 /v1/chat/completions or /v1/embeddings requests, create a batch, and collect the output file within 24 hours — outside your per-minute rate limits, never above the regular price. Cancelled or expired batches still return the lines that finished. A batch.completed / batch.failed webhook fires at the end.
curl https://api.ultragpt.pro/v1/files -H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-F purpose="batch" -F file="@requests.jsonl"
curl https://api.ultragpt.pro/v1/batches -H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "input_file_id": "file-...", "endpoint": "/v1/chat/completions", "completion_window": "24h" }'Rerank, moderation & search
POST /rerank orders up to 1,000 documents by relevance to a query (Cohere-compatible; billed per search unit). POST /moderations classifies text and images with OpenAI's moderation models — free. POST /search/{engine} runs a standalone web search on Tavily, Exa, Perplexity, Parallel or DataForSEO and returns title, url, snippet and page content, billed per query. GET /models?type=rerank|moderation|search lists them with prices.
# Rerank documents for RAG
curl https://api.ultragpt.pro/v1/rerank -H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "model": "cohere-rerank-v3.5", "query": "capital of France", "documents": ["Paris is...", "Berlin is..."], "top_n": 1 }'
# Moderate text or images (free)
curl https://api.ultragpt.pro/v1/moderations -H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "model": "omni-moderation-latest", "input": "text to check" }'
# Standalone web search (billed per query)
curl https://api.ultragpt.pro/v1/search/dataforseo-search -H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "query": "latest AI news", "max_results": 5 }'Video generation
POST /videos starts a text- or image-to-video job in OpenAI's Videos API shape and returns a video object (status queued). Poll GET /videos/{id} until it is completed (or failed), then download GET /videos/{id}/content — a redirect to the MP4 (?variant=thumbnail for the poster). Pass seconds, and aspect_ratio or size ("1280x720"); resolution and image_url (a public https first frame) where the model supports them — GET /models?type=video lists durations, ratios and the price per second. You are billed the provider's actual cost when the clip completes; failed clips are free. A video.completed / video.failed webhook fires at the end.
curl https://api.ultragpt.pro/v1/videos \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "model": "<video-model>", "prompt": "A paper boat in the rain", "seconds": 4, "size": "1280x720" }'
curl https://api.ultragpt.pro/v1/videos/video_... -H "Authorization: Bearer $ULTRAGPT_API_KEY"
curl -L https://api.ultragpt.pro/v1/videos/video_.../content -H "Authorization: Bearer $ULTRAGPT_API_KEY" -o clip.mp4Realtime voice
Open a WebSocket to /realtime?model=<id> (wss://) and speak OpenAI's Realtime protocol: session.update, input_audio_buffer.append, response.create… Authenticate with the Authorization header, or from a browser with the subprotocol openai-insecure-api-key.<key> (only with a key restricted to your domain). Each response is billed from the usage the model reports; input transcription supports gpt-4o-transcribe, gpt-4o-mini-transcribe and whisper-1. A session closes after 60 minutes, after 10 idle minutes, or when your balance or the key's spend limit runs out (an error event, then close code 1008).
import WebSocket from 'ws'
const ws = new WebSocket('wss://api.ultragpt.pro/v1/realtime?model=gpt-realtime', {
headers: { Authorization: `Bearer ${process.env.ULTRAGPT_API_KEY}` },
})
ws.on('open', () => {
ws.send(JSON.stringify({
type: 'session.update',
session: { instructions: 'Be brief.', voice: 'alloy' },
}))
ws.send(JSON.stringify({ type: 'response.create', response: { modalities: ['text'] } }))
})
ws.on('message', (data) => {
const event = JSON.parse(data.toString())
if (event.type === 'response.done') console.log(event.response.usage)
})Idempotent retries
Send an Idempotency-Key header on billed POSTs. A retry with the same key and body replays the first response (header idempotent-replayed: true) instead of running — and billing — again; a retry while the first is still running gets 409. Keys are kept 24 hours per API key.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Idempotency-Key: order-1234-summary" \
-H "Content-Type: application/json" \
-d '{ "model": "grok-4.7", "messages": [{ "role": "user", "content": "Hello!" }] }'App attribution
Optionally send X-Title (your app's name) and HTTP-Referer (its URL). They split your usage by app and, with a public URL, list your app on the apps leaderboard.
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "X-Title: My App" \
-H "HTTP-Referer: https://myapp.com" \
...Key, credits & generations
GET /key returns the calling key's limits and spend; GET /credits the balance; GET /generation?id= the tokens (including cached and reasoning), cost, time to first token and latency of any request.
# The calling key: rate limits, this month's spend
curl https://api.ultragpt.pro/v1/key -H "Authorization: Bearer $ULTRAGPT_API_KEY"
# Credit balance (shared with the UltraGPT app)
curl https://api.ultragpt.pro/v1/credits -H "Authorization: Bearer $ULTRAGPT_API_KEY"
# Tokens, cost and latency of one request, by its response id
curl "https://api.ultragpt.pro/v1/generation?id=chatcmpl-..." -H "Authorization: Bearer $ULTRAGPT_API_KEY"Management keys & activity
Create a key with scope "management" to provision keys from your own backend — one per customer, each with its own spend limit (limit + limit_reset: daily, weekly, monthly or never), model allowlist and rate limits — via /keys, and to manage presets via /presets. Management keys cannot call models. GET /activity returns daily usage per model and endpoint for the last 30 days (whole workspace for management keys, the key itself otherwise), filterable by date, key and end user. POST /generations/lookup returns the exact cost of up to 1,000 request ids at once — ideal for billing your own customers.
# With a management key: provision a key per customer, capped at $20/month
curl https://api.ultragpt.pro/v1/keys -H "Authorization: Bearer $ULTRAGPT_MANAGEMENT_KEY" -H "Content-Type: application/json" \
-d '{ "name": "customer-42", "limit": 20, "allowed_models": ["gpt-5-mini"], "rpm_limit": 60 }'
# -> { "data": { "id": "...", "limit_remaining": 20, ... }, "key": "ugpt_live_..." } (shown once)
curl https://api.ultragpt.pro/v1/keys -H "Authorization: Bearer $ULTRAGPT_MANAGEMENT_KEY" # list
curl -X PATCH https://api.ultragpt.pro/v1/keys/KEY_ID -d '{ "limit": 50 }' ... # update
curl -X DELETE https://api.ultragpt.pro/v1/keys/KEY_ID -H "Authorization: Bearer $ULTRAGPT_MANAGEMENT_KEY" # revoke
# Daily usage per model for the last 30 days (?date=, ?key_id=, ?user=)
curl https://api.ultragpt.pro/v1/activity -H "Authorization: Bearer $ULTRAGPT_MANAGEMENT_KEY"
# Exact cost of up to 1000 requests at once (any key, incl. read-only)
curl https://api.ultragpt.pro/v1/generations/lookup -H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "ids": ["chatcmpl-...", "chatcmpl-..."] }'Webhooks
Configure an HTTPS endpoint in Settings to receive balance.low, credits.added, key.created / rotated / revoked, key.spend_cap.warning / reached, key.expiring, account.daily_limit.reached, batch.completed / failed, video.completed / failed and model.deprecated. Each POST is JSON { id, type, created, data } signed with HMAC-SHA256.
A failed delivery (non-2xx or timeout) is retried after 30 seconds, 5 minutes and 30 minutes with the same event id — dedupe on it. Retries survive our restarts. Every attempt is listed in Settings, where you can redeliver any event. After 20 consecutive failures the webhook is paused.
import crypto from 'node:crypto'
// UltraGPT-Signature: t=<unix seconds>,v1=<hex HMAC-SHA256 of "<t>.<raw body>">
export function verifyWebhook(rawBody, header, secret, toleranceSec = 300) {
const parts = Object.fromEntries(header.split(',').map(p => p.split('=')))
const expected = crypto.createHmac('sha256', secret).update(`${parts.t}.${rawBody}`).digest('hex')
const fresh = Math.abs(Date.now() / 1000 - Number(parts.t)) < toleranceSec
return fresh && crypto.timingSafeEqual(Buffer.from(expected), Buffer.from(parts.v1 ?? ''))
}Prompt logs, audit & security
Prompts and responses are not stored by default. Turn on prompt logging in Settings to keep them, encrypted, for 1–30 days and see them in each request's detail in Activity (inline images are reduced to a placeholder); purge them at any time.
The audit log in Settings records who created, rotated or revoked keys and changed the webhook, provider keys, settings, presets or organization members, with IP and time — organizations show their owner and admins the whole team's log.
Turn on email confirmation (Settings > Security) to require a 6-digit code before creating or rotating keys, granting the management scope, saving a webhook or provider key, or deleting an organization. A confirmation lasts 15 minutes.
Bring your own key
Add your own OpenRouter or OpenAI key in Settings and requests for that provider's models run on your account: you pay the provider, and we bill only a small fee. If the provider rejects your key you get a byok_key_rejected error.
Organizations
Create an organization on the Team page and invite people by email. Keys created in the organization bill the owner's balance; members see the team's usage and manage their own keys; owners and admins manage everything. Switch workspaces from the sidebar.
Billing
Requests are billed per token from your prepaid credit balance — never your subscription's chat quota. Before a request starts, your available balance must cover its worst case (input plus max_tokens, or 4,096 output tokens when unset); you are charged only for actual usage.
Rate limits
Requests per minute and concurrent requests are per key and the same for every account — there are no usage tiers. An optional daily spend limit (account) and monthly cap (key) are yours to set (see Models). Every /v1 response carries x-ratelimit-limit-requests, x-ratelimit-remaining-requests and x-ratelimit-reset-requests; a 429 adds retry-after.
Each key also has a tokens-per-minute limit (input + output, x-ratelimit-*-tokens headers; a key can only lower it), and can carry a spend limit that resets daily, weekly (Monday, UTC), monthly or never — a prepaid budget for one customer. GET /key reports limit, limit_remaining and usage per period.
# A key for one customer: $20 per week, 30k tokens per minute
curl https://api.ultragpt.pro/v1/keys -H "Authorization: Bearer $ULTRAGPT_MANAGEMENT_KEY" \
-H "Content-Type: application/json" \
-d '{ "name": "customer-42", "limit": 20, "limit_reset": "weekly", "tpm_limit": 30000 }'
# What the calling key has left
curl https://api.ultragpt.pro/v1/key -H "Authorization: Bearer $ULTRAGPT_API_KEY"
# -> { "data": { "limit": 20, "limit_reset": "weekly", "limit_remaining": 13.4,
# "usage_daily": 1.2, "usage_weekly": 6.6, "usage_monthly": 18.1, ... } }Errors
Errors use OpenAI's shape { "error": { message, type, code, param } } (Anthropic's on /v1/messages), so the SDKs raise their usual typed exceptions.
invalid_request_errorMalformed request; `param` names the field. unsupported_parameter: the model lacks tools / image input.authentication_errorMissing, invalid, revoked or expired key.insufficient_quotaBalance too low for this request. Top up, or lower max_tokens.permission_errorRead-only key, IP / domain not allowed, account restricted, or management_key_required / a management key used for model calls.not_found_errorUnknown model (or the key may not use it), or preset_not_found.model_retiredThe model passed its sunset date and has no replacement. The message names what to use.idempotency_conflictA request with this Idempotency-Key is still running.rate_limit_errorToo many requests or concurrent requests. Honor retry-after.rate_limit_errortokens_limit_exceeded: the key's tokens-per-minute limit. Honor retry-after.insufficient_quotaThe key's spend limit (daily_ / weekly_ / monthly_spend_cap_reached, key_limit_reached for a lifetime limit) or the account's daily_spend_limit_reached.api_errorEvery model tried failed or timed out. Safe to retry.service_unavailableThe model is temporarily unavailable.Model deprecations
When a model is scheduled for retirement, every response from it carries Deprecation and Sunset headers (and x-ultragpt-model-replacement), the catalog marks it deprecated, and accounts that used it in the last 30 days get an email, an in-app notice and a model.deprecated webhook — at once and again a week before. After the sunset date, requests to it are answered by its replacement (header x-ultragpt-model-redirected-from) or fail with 410 model_retired. Watch the status page.
curl -si https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" -H "Content-Type: application/json" \
-d '{ "model": "grok-4.7", "messages": [{ "role": "user", "content": "Hi" }] }' | grep -i -E "^(deprecation|sunset|link|x-ultragpt-model)"
# deprecation: @1790000000
# sunset: Thu, 15 Oct 2026 00:00:00 GMT
# x-ultragpt-model-replacement: newer-model
# After the sunset: answered by the replacement (x-ultragpt-model-redirected-from),
# or HTTP 410 { "error": { "code": "model_retired" } } when there is none.