Half-price tokens for work that can wait
Our GPUs have idle hours between traffic peaks. Batch jobs fill them. Upload a JSONL file and results arrive within 24 hours, usually much sooner, at 50% of the serverless rate.
| Model | Batch $/1M in | Batch $/1M out |
|---|---|---|
| llmcloud/deepseek-v3.2-exp | $0.110 | $0.425 |
| llmcloud/llama-4-maverick | $0.140 | $0.450 |
| llmcloud/qwen3-235b-a22b | $0.100 | $0.350 |
| llmcloud/qwen3-coder-480b | $0.175 | $0.600 |
| llmcloud/gpt-oss-120b | $0.050 | $0.200 |
50,000 requests per job
Each line is one chat, embedding or vision request. Run as many jobs in parallel as your quota allows.
24-hour completion window
Most jobs finish in under two hours. Requests that don't finish in the window aren't charged.
Same request body
Each line uses the same format as the real-time API, so you can move a pipeline over without rewriting prompts.
curl https://api.llmcloud.ai/v1/batches \ -H "Authorization: Bearer $LLMCLOUD_KEY" \ -F purpose=batch -F file=@requests.jsonl \ -F completion_window=24h
Batch FAQ
What's batch inference good for?+
Offline evals, dataset labelling, document extraction, synthetic data and backfills. Any job where results aren't needed while a user waits.
Is output quality the same as serverless?+
Yes. Batch runs the same weights at the same served precision. The only difference is when your request gets scheduled.
Can batch use my fine-tuned adapter?+
Yes. Point the model field at your adapter's model ID. It's billed at the base model's batch rate.