llmcloud.ai
← Back to blog
Cost· 11 min read

Qwen's 2-trillion-parameter MoE: serving a model that doesn't fit anywhere

Alibaba's largest open-weight release pushes past the single-node era. What changes when weights outgrow a rack unit — sharding, cold starts, quantisation, and the real per-token math.

By llmcloud infrastructure

Qwen's 2T-parameter mixture-of-experts release is the clearest signal yet that open weights have outgrown the single-node mental model. You do not 'load' this model; you schedule it.

Sharding is the architecture

A 2T sparse model at FP8 is roughly 2 TB of weights. No current single host holds that in HBM, so you combine expert parallelism across nodes with tensor parallelism inside them. Each additional hop adds latency to every decoded token, so topology is a product decision, not an ops detail.

DeploymentWeights in HBMTypical TTFTNotes
4 nodes, FP8, InfiniBandYes0.4–0.8 sBest latency, highest fixed cost
2 nodes, INT4 expertsYes0.6–1.2 s~45% memory saved, small quality cost
1 node + NVMe expert offloadPartial2–6 sOnly viable for batch

Cold starts you can't hide

Streaming 2 TB of weights from object storage into HBM takes minutes even at 10s of GB/s. Autoscaling on request rate simply does not work at this size. Capacity has to be pre-warmed and held, which converts a variable cost into a fixed one — and fixed costs are only recovered by utilisation.

Quantisation is an economic lever

  • FP8 weights: near-lossless, now the default for MoE serving.
  • INT4 on expert FFNs while keeping attention and router at higher precision: large memory win, measurable degradation only on hard reasoning and code.
  • Never quantise the router — mis-routed experts cost far more quality than a quantised FFN.

The token economics

Because only a small fraction of parameters activate per token, throughput per GPU is closer to a mid-sized dense model than the parameter count suggests. Your cost is dominated by how many GPUs must stay resident, divided by tokens actually served.

blended_price = (r x $in + $out) / (r + 1)
  where r = input_tokens / output_tokens

RAG / agent traffic      r ~ 20 -> blended price near $in
Chat                     r ~ 3  -> mid
Long-form generation     r ~ 0.5 -> blended price near $out

Two teams can pay a 4x different effective rate on the identical published price sheet purely because of their input:output ratio. Prompt caching moves that ratio further in your favour: cached input tokens typically bill at 10–25% of fresh input, and RAG-heavy workloads are almost entirely input.

When a 2T model is the wrong answer

For classification, extraction, routing and most tool-calling steps, a 30B-class model matches it at a fraction of the cost and latency. Reserve frontier scale for the steps that genuinely need it, and cascade the rest. In our own traffic, fewer than 8% of requests measurably benefit from the largest available model.

Cascade it

Point your workload at auto:balanced and the router will try a small model first, escalating to the 2T tier only when confidence checks fail. Typical saving: 55–70% at equal end-task accuracy.