Apps · Roleplay & companions
SillyTavern
Power-user front end for persona and story chat.
Tokens / 7d
231B
Routing directive
auto:chat
On our metal
41%
Licence
Open source
SillyTavern resends a large character card and chat history on every turn, which makes it the most cache-sensitive workload on the gateway. Users are price-driven and skew heavily toward open-weight models we host directly.
How it connects
- Chat Completion source: Custom (OpenAI-compatible).
- Prompt caching on the persistent character-card prefix.
- No-logging keys available for privacy-sensitive users.
Traffic pattern
Massive input, small output. Cache hit rate above 80% is normal and cuts effective cost by 5–10x.
deepseek-v4qwen3-maxmistral-large-3
Point it at the gateway
// SillyTavern → API → Chat Completion → Custom Endpoint: https://api.llmcloud.ai/v1 Model: auto:chat