AI Dictionary › AI Models
Multi-head attention is an enhanced version of the attention mechanism in which the computation is not performed once, but in parallel across multiple independent "heads", each with its own parameters. Each head can learn to focus on a different type of relationship between tokens: one head may specialize in syntactic links, another in long-distance references, another in local patterns.
The representation vector of each token is split into several subspaces, one per head; each head performs its own attention computation within that reduced subspace, and the outputs of the different heads are then concatenated and projected back into the original space through a linear transformation. The overall computational cost remains comparable to that of a single full-size head, but the model's representational capacity increases.
It is a component present in virtually all modern attention-based architectures, from language models to computer vision systems that process images split into patches. Having multiple heads lets the model capture several types of dependencies simultaneously within the same layer, instead of having to learn them sequentially across separate layers.
The concept was introduced together with the self-attention mechanism as part of sequence-to-sequence architectures based solely on attention, proposed in the second half of the 2010s as an alternative to recurrent and convolutional models.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.