AI Dictionary › Modelli AI

Context Length

Lunghezza del contesto

Context length is the maximum number of tokens a language model can receive and process in a single inference session, including both the input text provided and, often, the text the model generates in response. It is an architectural limit intrinsic to the model, decided and fixed during training.

Definition

What it is

Context length is the maximum number of tokens a language model can receive and process in a single inference session, including both the input text provided and, often, the text the model generates in response. It is an architectural limit intrinsic to the model, decided and fixed during training.

How it works

This limit largely depends on the attention mechanism, whose computational and memory cost grows rapidly with the length of the processed sequence, since every token must potentially be compared against every other token in the sequence. Extending the context length beyond the values used in training requires specific techniques, such as positional encoding schemes designed to generalize beyond the lengths observed during training, or subsequent adaptations aimed precisely at expanding this limit.

Applications

It is a relevant parameter for every practical application of language models: it determines how much text, how much conversation history or how many documents can be included in a single request without resorting to retrieval or summarization techniques. Models with larger contexts can directly process entire documents, long conversations or large portions of source code in a single pass.

History & etymology

The term is closely linked to the more general concept of the context window, and its importance has grown alongside the spread of language models capable of handling ever-larger amounts of text without significant loss of coherence or accuracy.

Definizione (italiano)

La lunghezza del contesto è il numero massimo di token che un modello linguistico può ricevere ed elaborare in un'unica sessione di inferenza, comprendendo sia il testo fornito in input sia, spesso, il testo che il modello genera in risposta. È un limite architetturale intrinseco al modello, deciso e fissato durante l'addestramento.

Questo limite dipende in larga parte dal meccanismo di attenzione, il cui costo computazionale e di memoria cresce rapidamente con la lunghezza della sequenza elaborata, poiché ogni token deve potenzialmente confrontarsi con ogni altro token della sequenza. Estendere la lunghezza del contesto oltre i valori usati in addestramento richiede tecniche specifiche, come schemi di codifica posizionale progettati per generalizzare oltre le lunghezze osservate durante l'addestramento, o adattamenti successivi mirati proprio ad ampliare questo limite.

È un parametro rilevante per ogni applicazione pratica dei modelli linguistici: determina quanto testo, quanta cronologia di conversazione o quanti documenti possono essere inclusi in un'unica richiesta senza dover ricorrere a tecniche di recupero o riassunto. Modelli con contesti più ampi possono elaborare direttamente documenti interi, lunghe conversazioni o grandi porzioni di codice sorgente in un solo passaggio.

Il termine è strettamente collegato al concetto più generale di finestra di contesto, e la sua importanza è cresciuta di pari passo con la diffusione di modelli linguistici capaci di gestire quantità di testo sempre maggiori senza perdita significativa di coerenza o accuratezza.

Related terms

More in Modelli AI

Put it into practice

From our network

Kaimaki Web — Websites That Win Customers

Custom websites, web apps and digital marketing for growing businesses.

Visit kaimakiweb.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.