DeepSeek's V4 release doubles down on long chain-of-thought: the model spends a large, variable budget of internal tokens before emitting a visible answer. That single behaviour invalidates most of the capacity models teams built for chat-era workloads.
Why reasoning is expensive to serve
Reasoning tokens are decoded exactly like visible ones — one at a time, memory-bandwidth-bound, holding a KV cache slot the entire time. A request that thinks for 8,000 tokens before writing 300 occupies a decode slot ~27x longer than its visible output implies.
| Workload | Visible out | Reasoning out | Billed output |
|---|---|---|---|
| Simple Q&A | 200 | 0–400 | 200–600 |
| Code fix | 400 | 1,500–4,000 | 1,900–4,400 |
| Hard math / planning | 300 | 6,000–20,000 | 6,300–20,300 |
The variance matters more than the mean. Reasoning length has a long right tail, so p99 latency and p99 cost can be an order of magnitude above the median. Capacity planned on averages will page you at 3am.
Controls that actually work
- Set an explicit reasoning-effort or max-reasoning-tokens ceiling per route — treat it as a budget, not a hint.
- Route by difficulty: only escalate to a reasoning model when a cheap classifier or a low-confidence first attempt says you need it.
- Cache aggressively — reasoning traces for identical prompts are pure waste on repeat.
- Alert on reasoning-tokens-per-request, not just request count. It is the leading indicator of a cost incident.
Hosting complexities
V4 is another large sparse MoE, so it inherits the same expert-parallel and KV-cache pressures as its peers — but with far longer average sequences. Two consequences on owned hardware:
- Batch scheduling must be preemption-aware; one 20K-token thinker should not starve a queue of short requests.
- KV memory per replica must be sized for the tail, not the median, or you will thrash and evict mid-generation.
effective_cost = $in x in_tok
+ $out x (visible_tok + reasoning_tok)
A 5x cheaper reasoning model that thinks 10x longer
is 2x more expensive. Compare cost-per-solved-task,
never cost-per-token.Our analytics break output into visible and reasoning tokens per model and per route, so you can see cost-per-solved-task instead of guessing from a price sheet.