← All guides AI Release
RU Subscribe to @ai_release1

How to save LLM tokens: caching, routing and prompt caching in practice

Updated: 07.10.2026 · AI Release · @ai_release1
How to save LLM tokens: caching, routing and prompt caching in practice

This guide solves one practical problem: your LLM API bill grows faster than your product. It covers three techniques that work together — prompt caching, semantic caching, and model routing — and is suitable for teams using commercial APIs (Anthropic, OpenAI) or managing a mix of local and hosted models. No special hardware is required; everything runs with a Python script and a few configuration flags.

Requirements and preparation

Step-by-step instructions

1. Split your prompt into a stable part and a variable part.

Put the system prompt, instructions, and few-shot examples at the beginning. Put the user question at the end.

Expected result: your request becomes a fixed prefix plus a short variable tail — this is the prerequisite for prompt caching.

2. Enable prompt caching on the stable prefix.

In the Anthropic SDK, add `cache_control` to the last block of the stable part:

   from anthropic import Anthropic

   client = Anthropic()
   response = client.messages.create(
       model="claude-...",  # pick a cache-capable model
       max_tokens=1024,
       system=[
           {
               "type": "text",
               "text": "You are a support assistant. Always answer in 3 sentences.",
               "cache_control": {"type": "ephemeral"},
           }
       ],
       messages=[{"role": "user", "content": "How do I reset my password?"}],
   )

Expected result: the first request writes the cache; repeated requests with the same prefix read from the cache and are billed at a reduced rate.

3. Add a semantic cache for full user questions.

Before calling the LLM, embed the query and search a vector store. If similarity to a stored question exceeds a threshold (for example, 0.95), return the stored answer. Otherwise call the model and save the new answer.

Expected result: identical and near-identical questions skip the API call entirely. This matters most in agent loops, where the same context repeats every turn — more in AI agents for task automation: a guide.

4. Route requests by complexity.

Classify each request with a small model or with simple heuristics. Send easy requests to a cheap or free model, hard ones to a large model:

   def route(message: str) -> str:
       if len(message) < 80:
           return "small-model"
       if any(tag in message for tag in ("legal", "code", "analysis")):
           return "large-model"
       return "small-model"

Expected result: a large share of traffic goes to the cheaper model. For picking the right pair, check Best AI models and neural networks in 2026.

5. Monitor your usage data.

Inspect the API response fields: `cache_read_input_tokens` and `cache_creation_input_tokens` (Anthropic) or `cached_tokens` (OpenAI). Track the cache hit rate and the average cost per request weekly.

Expected result: you see exactly where tokens go and can tune the routing threshold.

Possible problems and solutions

The prefix changes between requests. Move all dynamic content — dates, usernames, IDs — after the cached block.

Add a fallback: if the response is empty, contains "I don't know", or fails a simple quality check, retry once on the large model.

Set a TTL (for example, 1–24 hours) and clear cached entries when your source data is updated.

FAQ

Does prompt caching work with all models?

No. Anthropic and OpenAI support it only for specific model versions. Check the provider documentation for the model list and the exact `cache_control` syntax.

How much can I save with prompt caching?

Providers bill cache reads at a heavily discounted rate (Anthropic documents roughly 90% cheaper cached reads). Real savings depend on how many tokens in your prompt are stable and reused.

Do I pay for creating the cache?

Yes. Writing a cache is more expensive than a normal input read, so caching pays off only when the same prefix is reused many times.

Can I use routing without extra infrastructure?

Yes. A simple Python function like the one in step 4 is enough. You can also combine routing with a free model tier — see How to use AI models for free in 2026.

🤖 AI summary