AI Dictionary › Modelli AI
Rete Feed-forward
The feed-forward network is the sublayer present in every block of modern attention-based architectures that processes each token representation independently of the others, applying non-linear transformations point by point. More generally, the term also refers to any neural network in which information flows in a single direction, from input to output, without loops or feedback.
The feed-forward network is the sublayer present in every block of modern attention-based architectures that processes each token representation independently of the others, applying non-linear transformations point by point. More generally, the term also refers to any neural network in which information flows in a single direction, from input to output, without loops or feedback.
Inside a transformer block, after the attention layer that relates tokens to one another, the feed-forward network applies the same transformation to each token: a linear projection that expands the vector's dimension, a non-linear activation function, and a second linear projection that brings the vector back to its original size. This step lets the model process and recombine the information gathered by attention in a more expressive way.
It is present in every block of modern language models, where it accounts for a significant share of the model's total parameter count. Pure feed-forward networks, without attention mechanisms, also underlie the earliest multilayer perceptrons used for classification and regression tasks.
The term "feed-forward" simply describes the direction of the computation flow, as opposed to recurrent networks where the output of a layer can be fed back as input to the same layer at a later step.
La rete feed-forward è il sottostrato presente in ogni blocco delle architetture moderne basate su attenzione che elabora ciascuna rappresentazione di token in modo indipendente dalle altre, applicando trasformazioni non lineari punto per punto. Il termine, più in generale, indica anche qualunque rete neurale in cui l'informazione fluisce in un'unica direzione, dall'input all'output, senza cicli o retroazioni.
All'interno di un blocco transformer, dopo lo strato di attenzione che mette in relazione i token tra loro, la rete feed-forward applica a ciascun token la stessa trasformazione: una proiezione lineare che espande la dimensione del vettore, una funzione di attivazione non lineare, e una seconda proiezione lineare che riporta il vettore alla dimensione originale. Questo passaggio permette al modello di elaborare e ricombinare le informazioni raccolte dall'attenzione in modo più espressivo.
È presente in ogni blocco dei modelli linguistici moderni, dove costituisce una parte significativa del numero totale di parametri del modello. Reti feed-forward pure, senza meccanismi di attenzione, sono inoltre alla base dei primi percettroni multistrato usati per compiti di classificazione e regressione.
Il termine "feed-forward" descrive semplicemente la direzione del flusso di calcolo, in contrapposizione alle reti ricorrenti dove l'output di uno strato può essere reimmesso come input nello stesso strato in un passo successivo.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.