AI Dictionary › Modelli AI
Multi-head attention is an enhanced version of the attention mechanism in which the computation is not performed once, but in parallel across multiple independent "heads", each with its own parameters. Each head can learn to focus on a different type of relationship between tokens: one head may specialize in syntactic links, another in long-distance references, another in local patterns.
Multi-head attention is an enhanced version of the attention mechanism in which the computation is not performed once, but in parallel across multiple independent "heads", each with its own parameters. Each head can learn to focus on a different type of relationship between tokens: one head may specialize in syntactic links, another in long-distance references, another in local patterns.
The representation vector of each token is split into several subspaces, one per head; each head performs its own attention computation within that reduced subspace, and the outputs of the different heads are then concatenated and projected back into the original space through a linear transformation. The overall computational cost remains comparable to that of a single full-size head, but the model's representational capacity increases.
It is a component present in virtually all modern attention-based architectures, from language models to computer vision systems that process images split into patches. Having multiple heads lets the model capture several types of dependencies simultaneously within the same layer, instead of having to learn them sequentially across separate layers.
The concept was introduced together with the self-attention mechanism as part of sequence-to-sequence architectures based solely on attention, proposed in the second half of the 2010s as an alternative to recurrent and convolutional models.
Il multi-head attention è una versione potenziata del meccanismo di attenzione in cui il calcolo non viene eseguito una sola volta, ma in parallelo su più "teste" indipendenti, ciascuna con i propri parametri. Ogni testa può imparare a concentrarsi su un tipo diverso di relazione tra i token: una testa può specializzarsi su legami sintattici, un'altra su riferimenti a lunga distanza, un'altra ancora su pattern locali.
Il vettore di rappresentazione di ogni token viene suddiviso in più sottospazi, uno per testa; ciascuna testa esegue il proprio calcolo di attenzione in quel sottospazio ridotto, e i risultati delle diverse teste vengono poi concatenati e proiettati nuovamente nello spazio originale tramite una trasformazione lineare. Il costo computazionale complessivo resta paragonabile a quello di una singola testa a piena dimensione, ma la capacità rappresentativa del modello aumenta.
È una componente presente in praticamente tutte le architetture moderne basate su attenzione, dai modelli linguistici ai sistemi di visione artificiale che elaborano immagini suddivise in patch. Avere più teste consente al modello di catturare simultaneamente diversi tipi di dipendenze nello stesso strato, invece di doverle apprendere in sequenza su strati separati.
Il concetto è stato introdotto insieme al meccanismo di self-attention nell'ambito delle architetture sequence-to-sequence basate solo su attenzione, proposte nella seconda metà degli anni 2010 come alternativa ai modelli ricorrenti e convoluzionali.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.