llmcloud.ai
← All providers
open source-host

Meta (Llama)

Llama 4 open weights; served across many partner clouds.

live
Uptime (30d)
99.9%
Median tps
165
Median TTFT
150 ms
Price index
0.5×
vs. market median
Overview
HQ
Menlo Park, US
Founded
2004
Models
7
Modalities
text, vision
Regions
us-east, us-west, eu-west
Strengths
open-weightsthroughput
Route via llmcloud
curl https://api.llmcloud.ai/v1/chat/completions \
  -H "Authorization: Bearer $LLMCLOUD_KEY" \
  -H "X-Provider: meta" \
  -d '{
    "model": "auto:quality",
    "messages": [{"role":"user","content":"hi"}]
  }'
Models served
llama-4-405b
q 87.6165 tps$2.4/1M
llama-4-maverick
q 86.9178 tps$1.8/1M
llama-3.1-405b-instruct
q 85.2130 tps$2.7/1M
llama-4-scout
q 82.1240 tps$0.6/1M
llama-4-70b
q 81.1220 tps$0.9/1M
llama-3.3-70b-instruct
q 80.4225 tps$0.7/1M
llama-3.2-90b-vision
q 79.4150 tps$0.9/1M
llama-3.1-70b-instruct
q 78.5240 tps$0.6/1M
llama-3.2-11b-vision
q 72.1260 tps$0.2/1M
llama-3.1-8b-instruct
q 68.2380 tps$0.05/1M
llama-guard-3-8b
q 65.0340 tps$0.05/1M