llmcloud.ai
← Back to blog
Cost· 9 min read

DeepSeek V4 and the hidden cost of reasoning tokens

Long-thinking models bill you for tokens you never read. A look at V4's serving profile, why reasoning breaks conventional capacity planning, and how to budget for invisible output.

By llmcloud engineering

DeepSeek's V4 release doubles down on long chain-of-thought: the model spends a large, variable budget of internal tokens before emitting a visible answer. That single behaviour invalidates most of the capacity models teams built for chat-era workloads.

Why reasoning is expensive to serve

Reasoning tokens are decoded exactly like visible ones — one at a time, memory-bandwidth-bound, holding a KV cache slot the entire time. A request that thinks for 8,000 tokens before writing 300 occupies a decode slot ~27x longer than its visible output implies.

WorkloadVisible outReasoning outBilled output
Simple Q&A2000–400200–600
Code fix4001,500–4,0001,900–4,400
Hard math / planning3006,000–20,0006,300–20,300

The variance matters more than the mean. Reasoning length has a long right tail, so p99 latency and p99 cost can be an order of magnitude above the median. Capacity planned on averages will page you at 3am.

Controls that actually work

  • Set an explicit reasoning-effort or max-reasoning-tokens ceiling per route — treat it as a budget, not a hint.
  • Route by difficulty: only escalate to a reasoning model when a cheap classifier or a low-confidence first attempt says you need it.
  • Cache aggressively — reasoning traces for identical prompts are pure waste on repeat.
  • Alert on reasoning-tokens-per-request, not just request count. It is the leading indicator of a cost incident.

Hosting complexities

V4 is another large sparse MoE, so it inherits the same expert-parallel and KV-cache pressures as its peers — but with far longer average sequences. Two consequences on owned hardware:

  • Batch scheduling must be preemption-aware; one 20K-token thinker should not starve a queue of short requests.
  • KV memory per replica must be sized for the tail, not the median, or you will thrash and evict mid-generation.
effective_cost = $in x in_tok
               + $out x (visible_tok + reasoning_tok)

A 5x cheaper reasoning model that thinks 10x longer
is 2x more expensive. Compare cost-per-solved-task,
never cost-per-token.
Measure the right thing

Our analytics break output into visible and reasoning tokens per model and per route, so you can see cost-per-solved-task instead of guessing from a price sheet.