Current billing window
…
Rates USD per 1M tokens
| Model | Input · cache hit | Input · cache miss | Output |
|---|
Off-peak is 50% of peak — peak is 2× off-peak. Cache-hit inputs are 50× cheaper than cache-miss on Flash, 30× on Pro — see the playbook below. CNY column shows DeepSeek's official ¥ prices (ZH pricing page), not an FX conversion.
Cache playbook
On Flash a cache hit costs $0.003/1M against $0.15/1M for a miss — 50× cheaper. Prefix discipline is the bigger lever; off-peak then halves whatever still misses.
How DeepSeek decides
A request hits only if it fully matches a persisted cache prefix unit starting at token 0. A match that starts in the middle never hits. Caching is best-effort — automatic, free, no storage fee, and sampling randomness is untouched because only the prefix is reused.
Prefix units get persisted three ways:
- Request boundaries — every request persists two units, one at end of user input and one at end of model output. Append-only multi-turn therefore hits for free.
- Common-prefix detection — a prefix shared across requests becomes its own unit once DeepSeek has detected it. In DeepSeek's own long-text example requests #1 and #2 miss and #3 is the first to hit.
- Fixed token intervals — long inputs and outputs are carved into units at intervals, so a long prefix is never wholly uncacheable.
Cache builds in seconds and is evicted hours to days after last use. KV cache guide ↗
- Freeze the prefix order System → tools → few-shots → documents → user turn last. Everything stable goes in front so only the tail ever changes between requests.
- Evict volatile tokens from the head Timestamps, request UUIDs, “today is…”, randomly sampled few-shots. One changed token at position 12 rebills the entire prefix behind it at miss price.
- Serialize deterministically Stable JSON key order, stable tool order, identical whitespace —
json.dumps(obj, sort_keys=True, separators=(",", ":")). A dict that iterates differently is a different prefix. - Append, never rewrite history Each request already persists units at end-of-input and end-of-output, so appended turns hit for free. Summarizing or trimming old turns rewrites the prefix and forfeits the whole conversation's cache.
- Warm the prefix before fan-out A shared prefix becomes its own unit only after DeepSeek has seen it across requests — in DeepSeek's own example the first two miss and the third is the first to hit. Send two cheap calls with different tails, then parallelize; a cold burst of N requests is N misses.
- Batch by prefix, not by arrival Unused units are evicted within hours. Group the work that shares a document and run it while that prefix is still resident.
- Ship one prompt version at a time Every live variant is a separate prefix that has to be warmed and then kept warm on its own share of traffic. A 50/50 A/B doubles the cold-start misses and halves each prefix's odds of surviving the eviction window; a gradual rollout is worse, because the prefix keeps moving.
- Measure the hit rate
hit / (hit + miss)fromusageon every request, with an alert on drops. A silent prefix change is a 50× price increase that otherwise surfaces on the invoice.
Docs assistant: 12,000 requests/day · 30,000 input tokens of which 28,500 are a stable prefix · 700 output · 365 days.
| Cold | 95% cached | |
|---|---|---|
| Off-peak | — | — |
| Peak | — | — |
Automate off-peak dispatch
…