AI Dictionary › Fondamenti AI

Quantization

Quantizzazione

Quantization is the technique that reduces the numerical precision used to store a model's parameters, so it takes less memory and runs faster. Weights, normally saved as 16- or 32-bit floating-point numbers, are represented in more compact formats like 8-bit or even 4-bit integers. It's like redrawing a photo using a palette of a few colors instead of millions: the image stays recognizable but takes far less space. Similarly the quantized model keeps most of its abilities while losing a little numerical finesse. Modern methods minimize this loss by carefully choosing how to map original values.

Definition

Quantization is what lets powerful models run on modest hardware like laptops or phones, not just servers with costly GPUs. It's a key ingredient for making AI accessible, affordable, and more energy-sustainable.

Related terms

More in Fondamenti AI

Put it into practice

From our network

INDACO TMS: Transport Management for European Logistics

Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.

Visit indacotms.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.