AI Dictionary › AI Models
Rete Feed-forward
The feed-forward network is the sublayer present in every block of modern attention-based architectures that processes each token representation independently of the others, applying non-linear transformations point by point. More generally, the term also refers to any neural network in which information flows in a single direction, from input to output, without loops or feedback.
Inside a transformer block, after the attention layer that relates tokens to one another, the feed-forward network applies the same transformation to each token: a linear projection that expands the vector's dimension, a non-linear activation function, and a second linear projection that brings the vector back to its original size. This step lets the model process and recombine the information gathered by attention in a more expressive way.
It is present in every block of modern language models, where it accounts for a significant share of the model's total parameter count. Pure feed-forward networks, without attention mechanisms, also underlie the earliest multilayer perceptrons used for classification and regression tasks.
The term "feed-forward" simply describes the direction of the computation flow, as opposed to recurrent networks where the output of a layer can be fed back as input to the same layer at a later step.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.