Term: Latency
~4 min read
Estimated time: ~4 min read — for the in-app brief plus opening the primary source.
What this is
Latency is how long the user waits for the model’s response.
Everyday example
A support agent waits eight seconds for a suggested reply while the customer is on the line. That delay is latency — and it will decide whether the tool is used.
Latency is how long you wait for an answer — a customer and cost issue, not a tech footnote.
- Reasoning models and long prompts take longer.
- A 20-second wait is fine for a brief; it is fatal in a live call.
- Vendors can look “smarter” by taking more time — or cheaper by being faster and shallower.
- Measure latency on your real tasks, not the demo.
Next action: Set a wait-time target for each AI-facing channel before you scale it.
What changes in how you lead
How decision rights, process, and ownership should change.
- CX and Operations own latency the way they own queue time.
- Do not buy the slowest model for a real-time desk.
Compare related ideas
Latency vs Token pricing
Pricing is what you pay. Latency is what the user feels. A cheap model that is too slow still fails the process.
Open Token pricingDeep dive
Frontier models can be slower and dearer. Smaller models are often faster for simple tasks.
Sales and Operations: set a maximum wait for anything used in a live conversation.
Batch work (overnight contract review) can tolerate higher latency than a call-center copilot.
Ask for measured latency on your workload, not a best-case demo.