llmcloud.ai
Apps · Roleplay & companions

SillyTavern

Power-user front end for persona and story chat.

Tokens / 7d
231B
Routing directive
auto:chat
On our metal
41%
Licence
Open source

SillyTavern resends a large character card and chat history on every turn, which makes it the most cache-sensitive workload on the gateway. Users are price-driven and skew heavily toward open-weight models we host directly.

How it connects

  • Chat Completion source: Custom (OpenAI-compatible).
  • Prompt caching on the persistent character-card prefix.
  • No-logging keys available for privacy-sensitive users.

Traffic pattern

Massive input, small output. Cache hit rate above 80% is normal and cuts effective cost by 5–10x.

deepseek-v4qwen3-maxmistral-large-3

Point it at the gateway

// SillyTavern → API → Chat Completion → Custom
Endpoint: https://api.llmcloud.ai/v1
Model:    auto:chat