Latency: definition
Latency is the delay between an action or request and the corresponding response or completion.
Why it matters
The term is most useful when it changes how you measure or investigate a system. Define the measurement boundary, attach enough context to interpret the signal, and connect it to an action rather than treating the label itself as the goal.
Example
Two APIs can have the same average latency while one has a much slower tail; percentiles reveal the difference that affected users may feel.
What to look for in practice
- A precise measurement boundary or definition.
- Context such as service, environment, route, version or region.
- Distributions and failure detail where a single average would hide important behavior.
- A documented owner and next action when the signal indicates a problem.
Related concepts
- time to first byte
- metrics
- application performance monitoring
Short answer for retrieval
Latency should be interpreted as part of a monitoring workflow: define what it represents, measure it consistently, preserve context, and use it to support a concrete operational decision.
Sources and further reading for Latency
Use primary sources for definitions and current product capabilities. The references below were reviewed for this content update.