# How to save LLM tokens: caching, routing and prompt caching in practice

> AI Release · @ai_release1 · https://ai-release.net/guides/kak-sekonomit-tokeny-llm_en.html

This guide solves one practical problem: your LLM API bill grows faster than your product. It covers three techniques that work together — prompt caching, semantic caching, and model routing — and is suitable for teams using commercial APIs (Anthropic, OpenAI) or managing a mix of local and hosted models. No special hardware is required; everything runs with a Python script and a few configuration flags.

## Requirements and preparation

- Python 3.9+ installed.
- An API key for a provider that supports prompt caching (Anthropic Claude, OpenAI GPT-4 class models).
- The official SDK: `pip install anthropic openai`.
- Optional: an embedding model and a vector store (FAISS, Chroma) for semantic caching.
- Optional: a second cheaper or local model for the routing tier. For free options, see [How to use AI models for free in 2026](https://ai-release.net/guides/kak-polzovatsya-ii-besplatno.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

## Step-by-step instructions

1. Split your prompt into a stable part and a variable part.
   Put the system prompt, instructions, and few-shot examples at the beginning. Put the user question at the end.
   Expected result: your request becomes a fixed prefix plus a short variable tail — this is the prerequisite for prompt caching.

2. Enable prompt caching on the stable prefix.
   In the Anthropic SDK, add `cache_control` to the last block of the stable part:

   ```python
   from anthropic import Anthropic

   client = Anthropic()
   response = client.messages.create(
       model="claude-...",  # pick a cache-capable model
       max_tokens=1024,
       system=[
           {
               "type": "text",
               "text": "You are a support assistant. Always answer in 3 sentences.",
               "cache_control": {"type": "ephemeral"},
           }
       ],
       messages=[{"role": "user", "content": "How do I reset my password?"}],
   )
   ```

   Expected result: the first request writes the cache; repeated requests with the same prefix read from the cache and are billed at a reduced rate.

3. Add a semantic cache for full user questions.
   Before calling the LLM, embed the query and search a vector store. If similarity to a stored question exceeds a threshold (for example, 0.95), return the stored answer. Otherwise call the model and save the new answer.
   Expected result: identical and near-identical questions skip the API call entirely. This matters most in agent loops, where the same context repeats every turn — more in [AI agents for task automation: a guide](https://ai-release.net/guides/ii-agenty-dlya-avtomatizacii.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

4. Route requests by complexity.
   Classify each request with a small model or with simple heuristics. Send easy requests to a cheap or free model, hard ones to a large model:

   ```python
   def route(message: str) -> str:
       if len(message) < 80:
           return "small-model"
       if any(tag in message for tag in ("legal", "code", "analysis")):
           return "large-model"
       return "small-model"
   ```

   Expected result: a large share of traffic goes to the cheaper model. For picking the right pair, check [Best AI models and neural networks in 2026](https://ai-release.net/guides/luchshie-nejroseti-i-modeli-2026.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).

5. Monitor your usage data.
   Inspect the API response fields: `cache_read_input_tokens` and `cache_creation_input_tokens` (Anthropic) or `cached_tokens` (OpenAI). Track the cache hit rate and the average cost per request weekly.
   Expected result: you see exactly where tokens go and can tune the routing threshold.

## Possible problems and solutions

- Caching never activates.
  The prefix changes between requests. Move all dynamic content — dates, usernames, IDs — after the cached block.
- The router sends complex tasks to the small model.
  Add a fallback: if the response is empty, contains "I don't know", or fails a simple quality check, retry once on the large model.
- Semantic cache returns stale answers.
  Set a TTL (for example, 1–24 hours) and clear cached entries when your source data is updated.

## FAQ

**Does prompt caching work with all models?**

No. Anthropic and OpenAI support it only for specific model versions. Check the provider documentation for the model list and the exact `cache_control` syntax.

**How much can I save with prompt caching?**

Providers bill cache reads at a heavily discounted rate (Anthropic documents roughly 90% cheaper cached reads). Real savings depend on how many tokens in your prompt are stable and reused.

**Do I pay for creating the cache?**

Yes. Writing a cache is more expensive than a normal input read, so caching pays off only when the same prefix is reused many times.

**Can I use routing without extra infrastructure?**

Yes. A simple Python function like the one in step 4 is enough. You can also combine routing with a free model tier — see [How to use AI models for free in 2026](https://ai-release.net/guides/kak-polzovatsya-ii-besplatno.html?utm_source=tg&utm_medium=channel&utm_campaign=guide_inline&utm_content=guide_to_guide).
