TerminologytermStep 3: How you talk to modelsAllOperationsSales

Term: Latency

~4 min read

Estimated time: ~4 min read — for the in-app brief plus opening the primary source.

What this is

Latency is how long the user waits for the model’s response.

Everyday example

A support agent waits eight seconds for a suggested reply while the customer is on the line. That delay is latency — and it will decide whether the tool is used.

Latency is how long you wait for an answer — a customer and cost issue, not a tech footnote.

  • Reasoning models and long prompts take longer.
  • A 20-second wait is fine for a brief; it is fatal in a live call.
  • Vendors can look “smarter” by taking more time — or cheaper by being faster and shallower.
  • Measure latency on your real tasks, not the demo.

Next action: Set a wait-time target for each AI-facing channel before you scale it.

What changes in how you lead

How decision rights, process, and ownership should change.

  • CX and Operations own latency the way they own queue time.
  • Do not buy the slowest model for a real-time desk.

Compare related ideas

Latency vs Token pricing

Pricing is what you pay. Latency is what the user feels. A cheap model that is too slow still fails the process.

Open Token pricing

Deep dive

Frontier models can be slower and dearer. Smaller models are often faster for simple tasks.

Sales and Operations: set a maximum wait for anything used in a live conversation.

Batch work (overnight contract review) can tolerate higher latency than a call-center copilot.

Ask for measured latency on your workload, not a best-case demo.

Related terms

Related weekly lessons

terminologylatencycx