AI Dictionary › Fondamenti AI
Latenza (Latency)
Latency is the time between sending a prompt and receiving the model's response, typically measured as time-to-first-token and total generation time.
It depends on model size, prompt and response length, server load, and inference optimization. In AI product design, latency is a decisive UX factor: a customer service chatbot must answer in 1-2 seconds, while document analysis can afford 30.
From our network
Kaimaki Web: Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.