What is Latency?
The delay before the model starts answering. Low latency = the first token arrives fast; throughput = how many tokens per second after that.
The delay before the model starts answering. Low latency = the first token arrives fast; throughput = how many tokens per second after that.