Prompt caching
Cached input tokens bill at the (much lower) cached price shown on each model. OpenAI and DeepSeek models cache repeated prefixes automatically; for Anthropic and Gemini models mark the stable part with cache_control. usage.prompt_tokens_details.cached_tokens shows the hits.
Python
# Mark a large, stable prefix as cacheable (Anthropic / Gemini models via OpenRouter;
# OpenAI and DeepSeek models cache automatically). Cached tokens bill at the cached price.
response = client.chat.completions.create(
model="grok-4.7",
messages=[
{"role": "system", "content": [
{"type": "text", "text": LONG_REFERENCE_DOCUMENT, "cache_control": {"type": "ephemeral"}},
]},
{"role": "user", "content": "Summarize section 3."},
],
)
print(response.usage.prompt_tokens_details.cached_tokens)