AI Dictionary › AI Fundamentals

LoRA (Low-Rank Adaptation)

LoRA is an efficient fine-tuning technique that specializes a pre-trained model without modifying all of its original parameters. Instead of updating the entire weight matrix, it freezes the base model and adds small additional matrices that are trained in its place. The result is targeted adaptation that requires a fraction of the memory and compute of full fine-tuning.

Definition

How it works

The method relies on the idea that the changes needed to specialize a model have a low-rank structure, meaning they can be represented with matrices much smaller than the original ones. During training, only these additional matrices are updated, while the base model's weights stay unchanged. At inference time, the two contributions are combined to produce the final output.

Applications

LoRA is now widely used to customize large language models and image-generation models for specific tasks, with hardware requirements far lower than traditional fine-tuning. It also allows multiple lightweight adapters to be kept for different uses, loadable on the same base model without duplicating it entirely.

History & etymology

It was proposed as a parameter-efficient fine-tuning method to reduce the computational costs of fine-tuning large-scale models.

Related terms

More in AI Fundamentals

Put it into practice

From our network

AGORÀ Intelligence: Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.