AI Dictionary › Fondamenti AI

LoRA (Low-Rank Adaptation)

LoRA is an efficient fine-tuning technique that specializes a pre-trained model without modifying all of its original parameters. Instead of updating the entire weight matrix, it freezes the base model and adds small additional matrices that are trained in its place. The result is targeted adaptation that requires a fraction of the memory and compute of full fine-tuning.

Definition

What it is

LoRA is an efficient fine-tuning technique that specializes a pre-trained model without modifying all of its original parameters. Instead of updating the entire weight matrix, it freezes the base model and adds small additional matrices that are trained in its place. The result is targeted adaptation that requires a fraction of the memory and compute of full fine-tuning.

How it works

The method relies on the idea that the changes needed to specialize a model have a low-rank structure, meaning they can be represented with matrices much smaller than the original ones. During training, only these additional matrices are updated, while the base model's weights stay unchanged. At inference time, the two contributions are combined to produce the final output.

Applications

LoRA is now widely used to customize large language models and image-generation models for specific tasks, with hardware requirements far lower than traditional fine-tuning. It also allows multiple lightweight adapters to be kept for different uses, loadable on the same base model without duplicating it entirely.

History & etymology

It was proposed as a parameter-efficient fine-tuning method to reduce the computational costs of fine-tuning large-scale models.

Definizione (italiano)

LoRA è una tecnica di fine-tuning efficiente che permette di specializzare un modello pre-addestrato senza modificare tutti i suoi parametri originali. Invece di aggiornare l'intera matrice dei pesi, congela il modello di base e aggiunge piccole matrici addizionali che vengono addestrate al posto suo. Il risultato è un adattamento mirato che richiede una frazione della memoria e del tempo di calcolo di un fine-tuning completo.

Il funzionamento si basa sull'idea che le variazioni necessarie per specializzare un modello abbiano una struttura a basso rango, cioè possano essere rappresentate con matrici molto più piccole di quelle originali. Durante l'addestramento si aggiornano solo queste matrici aggiuntive, mentre i pesi del modello base restano invariati. A inferenza, i due contributi vengono combinati per produrre l'output finale.

LoRA è oggi molto diffusa per personalizzare grandi modelli linguistici e modelli di generazione immagini su compiti specifici, con hardware limitato rispetto al fine-tuning tradizionale. Consente inoltre di mantenere più adattatori leggeri per usi diversi, caricabili sullo stesso modello base senza duplicarlo interamente.

È stata proposta come metodo di adattamento efficiente dei parametri (parameter-efficient fine-tuning) per ridurre i costi computazionali del fine-tuning su modelli di grandi dimensioni.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.