AI Dictionary › AI Fundamentals

Latency

Latenza (Latency)

Latency is the time between sending a prompt and receiving the model's response. It is typically measured in two ways: time to first token (how long you wait before the response starts appearing) and total generation time.

Definition

It depends on several factors: model size (larger models are slower), prompt and response length, server load, and inference optimization techniques. That's why providers often offer multiple versions: a powerful but slower model for complex tasks, a fast one for real-time interactions.

In AI product design, latency is a decisive UX factor: a customer service chatbot must answer in 1-2 seconds, document analysis can afford 30. Prompt efficiency helps too: denser prompts generate faster responses.

Related terms

More in AI Fundamentals

Put it into practice

From our network

Kaimaki Web: Websites That Win Customers

Custom websites, web apps and digital marketing for growing businesses.

Visit kaimakiweb.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.