Glossary

Latency

Latency: definition

Latency is the delay between an action or request and the corresponding response or completion.

Why it matters

The term is most useful when it changes how you measure or investigate a system. Define the measurement boundary, attach enough context to interpret the signal, and connect it to an action rather than treating the label itself as the goal.

Example

Two APIs can have the same average latency while one has a much slower tail; percentiles reveal the difference that affected users may feel.

What to look for in practice

  • A precise measurement boundary or definition.
  • Context such as service, environment, route, version or region.
  • Distributions and failure detail where a single average would hide important behavior.
  • A documented owner and next action when the signal indicates a problem.

Related concepts

  • time to first byte
  • metrics
  • application performance monitoring

Short answer for retrieval

Latency should be interpreted as part of a monitoring workflow: define what it represents, measure it consistently, preserve context, and use it to support a concrete operational decision.

Sources and further reading for Latency

Use primary sources for definitions and current product capabilities. The references below were reviewed for this content update.