Performance estimate
See time to first token, throughput, and memory estimates in seconds — refine as needed.
Tested: Nemotron, DeepSeek V4, Gemma 4, Kimi, ... — type to autocomplete
Gated model? Add your HF token in Settings →
Serving mode:
Based on your configuration — ISL 2048, OSL 128, FP16 KV cache, 1 concurrent users.
Want to change assumptions?