Time to First Token (TTFT)
lower is better · ms · p50
Throughput (TPS)
higher is better · tok/s · p50
The performance figures above are collected by periodically probing each model with a prompt of roughly 1,000 tokens and measuring the response. Actual latency and throughput may differ from your real-world experience depending on prompt size, concurrency, and workload.