Documentation
OpenAI- and Anthropic-compatible. Point any SDK at https://api.ultragpt.pro/v1.
- Create an API key and store it as
ULTRAGPT_API_KEY. - Add credits to your balance (shared with the UltraGPT app).
- Install the SDK (
pip install openai/npm i openai) and call a model:
curl https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.7",
"messages": [{ "role": "user", "content": "Hello!" }]
}'Endpoints
POST
/chat/completionsChat completion (OpenAI format), streaming or not.POST
/messagesAnthropic Messages API (Anthropic SDKs, Claude Code).POST
/messages/count_tokensEstimate prompt tokens (Anthropic format).POST
/responsesOpenAI Responses API, with stored conversations.GET
/responses/{id}A stored response (also DELETE, /input_items).POST
/completionsLegacy text completions.POST
/embeddingsText embeddings.POST
/images/generationsGenerate images.POST
/audio/speechText to speech.POST
/audio/transcriptionsSpeech to text (also /audio/translations).POST
/filesUpload a batch input file (also GET, DELETE, /content).POST
/batchesRun up to 5,000 requests asynchronously, outside your rate limits (also GET, /cancel).POST
/rerankRerank documents against a query (Cohere / Voyage / Qwen).POST
/moderationsClassify text and images for harmful content. Free.POST
/videosStart a video generation (async; also GET, GET /content, DELETE).WS
/realtime?model=Realtime speech-to-speech over WebSocket (OpenAI Realtime protocol).POST
/search/{engine}Standalone web search (Tavily, Exa, Perplexity…), per query.GET
/modelsModels your key can use (?type=, ?supported_parameters=, ?input_modalities= filters; also /models/count).GET
/presetsPresets of your workspace (writes need a management key).GET
/keysManagement keys: list, create, update, revoke keys.GET
/activityDaily usage per model and endpoint, last 30 days.POST
/generations/lookupCost and stats of up to 1000 requests at once.GET
/keyThe calling key: limits, spend.GET
/creditsCredit balance and lifetime totals.GET
/generation?id=Tokens, cost and latency of one request.