Prompt caching

Cached input tokens bill at the (much lower) cached price shown on each model. OpenAI and DeepSeek models cache repeated prefixes automatically; for Anthropic and Gemini models mark the stable part with cache_control. usage.prompt_tokens_details.cached_tokens shows the hits.

Python
# Mark a large, stable prefix as cacheable (Anthropic / Gemini models via OpenRouter;
# OpenAI and DeepSeek models cache automatically). Cached tokens bill at the cached price.
response = client.chat.completions.create(
    model="grok-4.7",
    messages=[
        {"role": "system", "content": [
            {"type": "text", "text": LONG_REFERENCE_DOCUMENT, "cache_control": {"type": "ephemeral"}},
        ]},
        {"role": "user", "content": "Summarize section 3."},
    ],
)
print(response.usage.prompt_tokens_details.cached_tokens)