llmcloud.ai
Products · Multimodal

More than text, on the same factory

Pipelines that read PDFs, generate images, transcribe calls and embed documents shouldn't need four vendors. Every modality runs on our fleet with one key and one bill.

ModalityEndpointModelsPriceStatus
Vision & documents/v1/chat/completionsQwen3-VL 235B, Llama 4 Maverick$0.40 / $1.40 per 1MBeta
Image generation/v1/images/generationsFLUX.2, Qwen-Imagefrom $0.008 / imageBeta
Speech-to-text/v1/audio/transcriptionsWhisper v4, Parakeet$0.0015 / audio-minComing soon
Text-to-speech/v1/audio/speechKokoro, Orpheus$8 / 1M charactersComing soon
Embeddings & rerank/v1/embeddingsllmcloud embed-3, Qwen3-Embedding$0.02 per 1MComing soon
Vision request
client.chat.completions.create(
    model="llmcloud/qwen3-vl-235b",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "Extract the invoice total"},
        {"type": "image_url", "image_url": {"url": invoice_url}},
    ]}],
)

Multimodal FAQ

Can I batch multimodal jobs?+

Yes. Vision, transcription and embedding requests all work with the Batch API at 50% off.

Do images and audio leave the region?+

No. Media files are processed in the region the request is pinned to, and they're deleted after inference unless you turn on retention.