Kimi K3 continues Moonshot's sparse mixture-of-experts line: roughly a trillion total parameters with only tens of billions active per token. That ratio is why the per-token price looks affordable and why the hosting bill surprises teams that budget on FLOPs alone.
The serving footprint
Sparse MoE decouples compute from memory. You only run a handful of experts per token, but every expert must be resident somewhere fast enough to be reachable within the decode budget. At FP8 that is roughly a terabyte of weights before you touch the KV cache.
| Resource | Dense 70B (FP8) | K3-class 1T sparse MoE (FP8) |
|---|---|---|
| Weight memory | ~70 GB | ~1 TB |
| Minimum node | 1x 8-GPU host | 2–4x 8-GPU hosts, expert-parallel |
| Active params/token | 70B | ~32B |
| Dominant bottleneck | Memory bandwidth | Interconnect + expert routing |
Expert parallelism means every decode step ships activations across the fabric. On a well-tuned NVLink/InfiniBand domain that costs a few percent of step time; on commodity Ethernet it can double your time-per-output-token. This is the single largest determinant of whether a K3 deployment is economical.
The KV-cache tax
K3's long context is the feature people buy it for, and it is also what fills your GPUs. KV cache scales linearly with context length and concurrency, and it competes with weights for the same HBM. A 256K-token session can consume more memory than a small dense model.
- Paged attention and prefix sharing are mandatory, not optimisations — repeated system prompts should be stored once per node.
- KV quantisation to FP8 roughly halves cache pressure with negligible quality loss on long-context retrieval.
- Cap max concurrent long-context sessions per replica; a single 1M-token outlier can evict dozens of short chats.
The input/output token equation
Prefill and decode are different businesses. Prefill is compute-bound and batches beautifully: thousands of input tokens per second per GPU. Decode is memory-bandwidth-bound and serial: one token at a time per sequence. That asymmetry is why every provider prices output tokens 3–5x above input.
cost_per_request
= (input_tokens x $in_per_token)
+ (output_tokens x $out_per_token)
where, for a K3-class model on owned metal:
$in_per_token ~= gpu_hour_cost / prefill_tokens_per_hour
$out_per_token ~= gpu_hour_cost / (decode_tps x utilisation)
utilisation is the number that kills you: a replica idling at 20%
pays full rent on ~1 TB of HBM.Practical consequence: agentic workloads with long tool-call transcripts and short replies are cheap on K3, while report generation with 4K-token answers is where your margin evaporates. Measure your own input:output ratio before comparing list prices.
Route or host?
- Under ~50M output tokens/month: route. You cannot beat a provider's utilisation with bursty traffic.
- Above that, with steady 24/7 load and a residency requirement: hosting starts to win, provided you can keep replicas above ~60% utilisation.
- Mixed: host the steady baseline, burst to routed capacity. That is exactly what our hosted fleet plus gateway failover is designed for.
K3 is available through the gateway at pass-through provider pricing with $0 gateway fee, and is on the roadmap for our first-party hosted fleet in sovereign regions.