AI Dictionary › Fondamenti AI
Throughput is the amount of work a system completes per unit of time: requests per second, tokens per second, transactions per minute. It measures capacity, to be distinguished from latency, which measures the speed of a single operation. A ferry has high throughput and high latency; a speedboat the opposite: it carries little, yet arrives fast.
Throughput is the amount of work a system completes per unit of time: requests per second, tokens per second, transactions per minute. It measures capacity, to be distinguished from latency, which measures the speed of a single operation. A ferry has high throughput and high latency; a speedboat the opposite: it carries little, yet arrives fast.
Throughput and latency are often in tension: techniques like batching raise overall capacity at the price of making individual requests wait, while optimizing for instant response can waste capacity. Real throughput is set by the bottleneck: CPU, memory, network bandwidth or disk I/O. Finding and widening it is the craft of performance optimization.
In AI, throughput is measured in tokens per second: inference systems like vLLM use continuous batching to serve many requests together and multiply tokens generated per GPU, slashing unit cost. In training, what matters is examples processed per second across the whole cluster. Beyond AI it drives the sizing of databases, networks and payment systems.
The English word combines through and put: what passes through the system. It was already used in manufacturing and railways for productive or transport capacity; computing and telecommunications adopted it for the capacity of channels and processors, and Erlang's queueing theory has provided the mathematical tools for analyzing it since 1909.
Il throughput è la quantità di lavoro che un sistema completa nell'unità di tempo: richieste al secondo, token al secondo, transazioni al minuto. È la misura della portata, da distinguere dalla latenza, che misura invece la velocità della singola operazione. Un traghetto ha throughput alto e latenza alta; un motoscafo l'opposto: trasporta poco, ma arriva subito.
Throughput e latenza sono spesso in tensione: tecniche come il batching aumentano la portata complessiva al prezzo di far attendere le singole richieste, mentre ottimizzare la risposta immediata può sprecare capacità. Il throughput reale dipende dal collo di bottiglia: CPU, memoria, banda di rete o I/O su disco. Individuarlo e allargarlo è il mestiere dell'ottimizzazione delle prestazioni.
Nell'AI il throughput si misura in token al secondo: i sistemi di inferenza come vLLM usano il continuous batching per servire molte richieste insieme e moltiplicare i token generati per GPU, abbattendo il costo unitario. Nel training conta il throughput di esempi processati al secondo sull'intero cluster. Fuori dall'AI governa il dimensionamento di database, reti e sistemi di pagamento.
La parola inglese nasce dalla composizione di through (attraverso) e put (mettere): ciò che passa attraverso il sistema. Era già in uso nell'industria manifatturiera e ferroviaria per la capacità produttiva o di trasporto; l'informatica e le telecomunicazioni la adottarono per la capacità dei canali e degli elaboratori, e la teoria delle code di Erlang ne fornì fin dal 1909 gli strumenti matematici di analisi.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.