Response caching
Send X-UltraGPT-Cache: true (X-OpenRouter-Cache works too) and an identical request from the same key within the TTL (X-UltraGPT-Cache-TTL, 1–86400 s, default 300) is answered from cache instantly and for free: usage counters are 0 and nothing is billed. Works for streaming and non-streaming chat, messages, responses and completions. Responses carry x-ultragpt-cache-status (HIT / MISS), -age, -ttl and -source-id; X-UltraGPT-Cache-Clear: true refreshes the entry. Presets can turn caching on for every request.
curl -i https://api.ultragpt.pro/v1/chat/completions \
-H "Authorization: Bearer $ULTRAGPT_API_KEY" \
-H "Content-Type: application/json" \
-H "X-UltraGPT-Cache: true" \
-H "X-UltraGPT-Cache-TTL: 600" \
-d '{ "model": "grok-4.7", "messages": [{ "role": "user", "content": "Capital of France?" }] }'
# Second identical call within 10 minutes:
# x-ultragpt-cache-status: HIT x-ultragpt-cache-age: 12 usage.cost: 0
# Force a refresh: -H "X-UltraGPT-Cache-Clear: true"