llmcloud.ai
← Back to blog
Cost· 10 min read

Kimi K3: what it actually costs to host a 1T-parameter sparse MoE

Moonshot's K3 lands with a huge sparse MoE and an agentic tool-use profile. We break down the serving footprint, the KV-cache tax, and the input/output token equation that decides whether you route or host it.

By llmcloud infrastructure

Kimi K3 continues Moonshot's sparse mixture-of-experts line: roughly a trillion total parameters with only tens of billions active per token. That ratio is why the per-token price looks affordable and why the hosting bill surprises teams that budget on FLOPs alone.

The serving footprint

Sparse MoE decouples compute from memory. You only run a handful of experts per token, but every expert must be resident somewhere fast enough to be reachable within the decode budget. At FP8 that is roughly a terabyte of weights before you touch the KV cache.

ResourceDense 70B (FP8)K3-class 1T sparse MoE (FP8)
Weight memory~70 GB~1 TB
Minimum node1x 8-GPU host2–4x 8-GPU hosts, expert-parallel
Active params/token70B~32B
Dominant bottleneckMemory bandwidthInterconnect + expert routing

Expert parallelism means every decode step ships activations across the fabric. On a well-tuned NVLink/InfiniBand domain that costs a few percent of step time; on commodity Ethernet it can double your time-per-output-token. This is the single largest determinant of whether a K3 deployment is economical.

The KV-cache tax

K3's long context is the feature people buy it for, and it is also what fills your GPUs. KV cache scales linearly with context length and concurrency, and it competes with weights for the same HBM. A 256K-token session can consume more memory than a small dense model.

  • Paged attention and prefix sharing are mandatory, not optimisations — repeated system prompts should be stored once per node.
  • KV quantisation to FP8 roughly halves cache pressure with negligible quality loss on long-context retrieval.
  • Cap max concurrent long-context sessions per replica; a single 1M-token outlier can evict dozens of short chats.

The input/output token equation

Prefill and decode are different businesses. Prefill is compute-bound and batches beautifully: thousands of input tokens per second per GPU. Decode is memory-bandwidth-bound and serial: one token at a time per sequence. That asymmetry is why every provider prices output tokens 3–5x above input.

cost_per_request
  = (input_tokens  x  $in_per_token)
  + (output_tokens x  $out_per_token)

where, for a K3-class model on owned metal:
  $in_per_token   ~= gpu_hour_cost / prefill_tokens_per_hour
  $out_per_token  ~= gpu_hour_cost / (decode_tps x utilisation)

utilisation is the number that kills you: a replica idling at 20%
pays full rent on ~1 TB of HBM.

Practical consequence: agentic workloads with long tool-call transcripts and short replies are cheap on K3, while report generation with 4K-token answers is where your margin evaporates. Measure your own input:output ratio before comparing list prices.

Route or host?

  • Under ~50M output tokens/month: route. You cannot beat a provider's utilisation with bursty traffic.
  • Above that, with steady 24/7 load and a residency requirement: hosting starts to win, provided you can keep replicas above ~60% utilisation.
  • Mixed: host the steady baseline, burst to routed capacity. That is exactly what our hosted fleet plus gateway failover is designed for.
On llmcloud

K3 is available through the gateway at pass-through provider pricing with $0 gateway fee, and is on the roadmap for our first-party hosted fleet in sovereign regions.