Products · Multimodal
More than text, on the same factory
Pipelines that read PDFs, generate images, transcribe calls and embed documents shouldn't need four vendors. Every modality runs on our fleet with one key and one bill.
| Modality | Endpoint | Models | Price | Status |
|---|---|---|---|---|
| Vision & documents | /v1/chat/completions | Qwen3-VL 235B, Llama 4 Maverick | $0.40 / $1.40 per 1M | Beta |
| Image generation | /v1/images/generations | FLUX.2, Qwen-Image | from $0.008 / image | Beta |
| Speech-to-text | /v1/audio/transcriptions | Whisper v4, Parakeet | $0.0015 / audio-min | Coming soon |
| Text-to-speech | /v1/audio/speech | Kokoro, Orpheus | $8 / 1M characters | Coming soon |
| Embeddings & rerank | /v1/embeddings | llmcloud embed-3, Qwen3-Embedding | $0.02 per 1M | Coming soon |
Vision request
client.chat.completions.create(
model="llmcloud/qwen3-vl-235b",
messages=[{"role": "user", "content": [
{"type": "text", "text": "Extract the invoice total"},
{"type": "image_url", "image_url": {"url": invoice_url}},
]}],
)Multimodal FAQ
Can I batch multimodal jobs?+
Yes. Vision, transcription and embedding requests all work with the Batch API at 50% off.
Do images and audio leave the region?+
No. Media files are processed in the region the request is pinned to, and they're deleted after inference unless you turn on retention.