AI Dictionary › Fondamenti AI

PCA (Principal Component Analysis)

PCA (Analisi delle componenti principali)

Principal Component Analysis, or PCA, is a dimensionality reduction technique that transforms a set of possibly correlated numeric variables into a smaller number of new variables, called principal components, which summarize most of the original information using as few dimensions as possible. It is one of the most widely used tools for simplifying complex datasets without losing what truly matters.

Definition

What it is

Principal Component Analysis, or PCA, is a dimensionality reduction technique that transforms a set of possibly correlated numeric variables into a smaller number of new variables, called principal components, which summarize most of the original information using as few dimensions as possible. It is one of the most widely used tools for simplifying complex datasets without losing what truly matters.

How it works

The technique identifies the directions along which the data varies the most: the first principal component is the direction of maximum variability in the data, the second is the direction of maximum remaining variability, perpendicular to the first, and so on. Mathematically this is achieved by computing the eigenvectors and eigenvalues of the covariance matrix of the original variables, standardized beforehand to prevent those with larger scale from dominating the result. By ranking components by how much variability they explain, it is possible to keep only the first few and discard the rest, drastically reducing the number of variables with a contained and quantifiable loss of information.

Applications

It is used to compress high-dimensional datasets before training other models, to visualize complex data in two or three dimensions understandable to the human eye, to eliminate redundancy between strongly correlated variables, and as a preprocessing step in fields like genomics, image processing and financial portfolio analysis. It also reduces the risk of overfitting in subsequent models by simplifying the feature space they must work with.

History & etymology

The mathematical foundations of PCA trace back to a 1901 paper by the British mathematician and biostatistician Karl Pearson, who introduced it as a method for finding the lines and planes that best approximate a set of points in space. The technique was later reformulated and made practically applicable in statistics by the American statistician Harold Hotelling in 1933, who is also credited with the name principal components, and it has remained one of the most solid and widely used tools of multivariate analysis ever since.

Definizione (italiano)

L'analisi delle componenti principali, o PCA, è una tecnica di riduzione della dimensionalità che trasforma un insieme di variabili numeriche possibilmente correlate in un numero minore di nuove variabili, dette componenti principali, che riassumono la maggior parte dell'informazione originale con il minor numero possibile di dimensioni. È uno degli strumenti più usati per semplificare dataset complessi senza perdere ciò che conta davvero.

La tecnica individua le direzioni lungo le quali i dati variano maggiormente: la prima componente principale è la direzione di massima variabilità nei dati, la seconda è la direzione di massima variabilità rimasta, perpendicolare alla prima, e così via. Matematicamente questo si ottiene calcolando gli autovettori e gli autovalori della matrice di covarianza delle variabili originali, standardizzate in anticipo per evitare che quelle con scala maggiore dominino il risultato. Ordinando le componenti per quanta variabilità spiegano, è possibile tenere solo le prime poche e scartare le altre, riducendo drasticamente il numero di variabili con una perdita di informazione contenuta e quantificabile.

È usata per comprimere dataset ad alta dimensionalità prima di addestrare altri modelli, per visualizzare dati complessi in due o tre dimensioni comprensibili all'occhio umano, per eliminare la ridondanza tra variabili fortemente correlate, e come passo di pre-elaborazione in campi come la genomica, l'elaborazione di immagini e l'analisi finanziaria di portafogli. Riduce anche il rischio di overfitting nei modelli successivi, semplificando lo spazio delle caratteristiche su cui devono lavorare.

Le fondamenta matematiche della PCA risalgono a un articolo del 1901 del matematico e biostatistico britannico Karl Pearson, che la introdusse come metodo per trovare le linee e i piani che meglio approssimano un insieme di punti nello spazio. La tecnica fu poi riformulata e resa praticamente applicabile in ambito statistico dallo statistico americano Harold Hotelling nel 1933, a cui si deve anche il nome componenti principali, e da allora è rimasta uno degli strumenti più solidi e usati dell'analisi multivariata.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.