Skip to main content

Response caching

If you send the same request twice, the second one may be answered from a short-lived cache instead of the model. It is why an occasional call returns in under a second when the model would have taken five.

You can turn it off per request, with the standard HTTP header.

What counts as "the same request"

Everything the model would see, plus who is asking:

  • the same model, the same messages, byte for byte;
  • the same generation parameters — max_tokens, temperature, top_p, seed, stop, tools, response_format, your reasoning settings, and so on;
  • the same account. Entries are never shared between accounts.

Change any of them and it is a different request, answered by the model.

Entries live for one hour. Streaming requests ("stream": true) are never cached, and requests billed to your own provider key (BYOK) are never cached.

Turning it off

Do not store my prompt or its answer
curl https://llm-api.tokligence.ai/v1/chat/completions \
-H "Authorization: Bearer $TOKLIGENCE_API_KEY" \
-H "Cache-Control: no-store" \
-H "Content-Type: application/json" \
-d '{ "model": "glm-5.2", "messages": [{"role":"user","content":"..."}] }'
HeaderEffect
Cache-Control: no-storeThis request is neither answered from the cache nor written to it.
Cache-Control: no-cacheThis request is answered by the model, but the answer may still be cached for later.

These are the standard HTTP directives (RFC 9111) — nothing of ours to learn. Anything else in the header (max-age, public, …) is ignored.

Which one do you want?

no-store is the one to reach for if the concern is your data. no-cache is the one to reach for if the concern is freshness — you want the model to answer again, but you don't mind us keeping the result.

Billing

A cached answer is billed at the normal per-token rate. The tokens are the tokens the model produced; the cache changes how fast you get them, not what they cost.

That is separate from your provider's own prompt caching, which does make repeated input cheaper — for example a cached input token on DeepSeek V4 Pro costs $0.003988 per million against $0.4785 for a fresh one. That discount is applied automatically, needs no parameter, and shows up in your usage records as cache_read_tokens.

Reproducibility

If you are running evaluations, comparing prompts, or measuring latency, send Cache-Control: no-store (or vary a nonce in the prompt). Otherwise your second run may be measuring our memory rather than the model.