AI Dictionary › Fondamenti AI

Instruction Tuning

Instruction tuning is a training phase in which a language model is refined on a set of examples made of instructions and correct responses, teaching it to follow commands expressed in natural language. Unlike pre-training, which teaches the model to predict the next piece of text, this phase teaches it to behave like an assistant that carries out specific requests. It is one of the steps that turns a base model into something usable in a conversational product.

Definition

What it is

Instruction tuning is a training phase in which a language model is refined on a set of examples made of instructions and correct responses, teaching it to follow commands expressed in natural language. Unlike pre-training, which teaches the model to predict the next piece of text, this phase teaches it to behave like an assistant that carries out specific requests. It is one of the steps that turns a base model into something usable in a conversational product.

How it works

The model is exposed to many instruction-response pairs covering different tasks: answering questions, summarizing text, writing code, translating. Through this supervised training, the model generalizes the behavior of following instructions even to tasks not seen during tuning, as long as they are structurally similar to the ones observed.

Applications

It is a common step in building conversational assistants based on LLMs, often combined with later techniques such as RLHF to further refine response style and safety. It is applied both by those who build base models and by those who customize them for a specific domain.

History & etymology

It is associated with the shift of language models from simple text-completion systems to assistants capable of executing instructions, a key change in the recent evolution of generative models.

Definizione (italiano)

L'instruction tuning è una fase di addestramento in cui un modello linguistico viene affinato su un insieme di esempi composti da istruzioni e risposte corrette, per insegnargli a seguire comandi espressi in linguaggio naturale. A differenza del pre-addestramento, che insegna a prevedere il testo successivo, questa fase insegna al modello a comportarsi come un assistente che esegue richieste specifiche. È uno dei passaggi che rende un modello base utilizzabile in un prodotto conversazionale.

Il modello viene esposto a molte coppie istruzione-risposta che coprono compiti diversi: rispondere a domande, riassumere testi, scrivere codice, tradurre. Attraverso questo addestramento supervisionato, il modello generalizza il comportamento di seguire istruzioni anche su compiti non visti durante il tuning, purché simili nella struttura a quelli osservati.

È un passaggio comune nella creazione di assistenti conversazionali basati su LLM, spesso combinato con tecniche successive come il RLHF per rifinire ulteriormente lo stile e la sicurezza delle risposte. Viene applicato sia da chi crea i modelli di base sia da chi li personalizza per un dominio specifico.

È associato alla transizione dei modelli linguistici da semplici sistemi di completamento del testo ad assistenti capaci di eseguire istruzioni, un cambiamento chiave nell'evoluzione recente dei modelli generativi.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.