AI Dictionary › Fondamenti AI
Outlier (Valore anomalo)
An outlier is an observation that deviates markedly from the typical behavior of the rest of the dataset: a transaction a thousand times larger than average, an age recorded as 150 years, a sensor that for an instant reports a physically impossible value. It can be a measurement or data entry error, or a rare but real and potentially very informative event.
An outlier is an observation that deviates markedly from the typical behavior of the rest of the dataset: a transaction a thousand times larger than average, an age recorded as 150 years, a sensor that for an instant reports a physically impossible value. It can be a measurement or data entry error, or a rare but real and potentially very informative event.
There are several methods for detecting outliers, and the choice depends on how well the data's distribution is understood. Classical statistical methods use thresholds based on standard deviation or interquartile range, flagging anything too far from the center of the distribution as anomalous. Density- or distance-based methods, such as isolation forest or local outlier factor, identify points that lie in sparsely populated regions of the feature space relative to their neighbors. The decision of what to do once an outlier is found, whether to remove it, correct it, or study it separately, always depends on context: in a dataset of real estate prices, an extreme value may be an error; in a fraud detection system, it is often exactly what is being sought.
In applied AI, handling outliers matters for two opposite reasons: on one hand, many algorithms, like linear regression or k-means, are sensitive to extreme values and can be significantly distorted by a few untreated outliers; on the other hand, in fields like fraud detection, cybersecurity and industrial quality control, the system's primary goal is precisely to identify outliers, which represent the rare event to be caught.
The term outlier, literally something that lies outside, entered the statistical vocabulary in the nineteenth century, when mathematicians and astronomers began systematically addressing how to treat experimental observations that clearly deviated from other measurements. Twentieth-century statistics then developed increasingly refined formal criteria for distinguishing a genuine measurement error from a rare but real observation, a distinction that remains at the heart of machine learning practice today.
Un outlier, o valore anomalo, è un'osservazione che si discosta in modo marcato dal comportamento tipico del resto del dataset: una transazione di importo mille volte superiore alla media, un'età registrata come 150 anni, un sensore che per un istante segnala un valore fisicamente impossibile. Può essere un errore di misurazione o inserimento dati, oppure un evento raro ma reale e potenzialmente molto informativo.
Esistono diversi metodi per individuare gli outlier, e la scelta dipende da quanto si conosce della distribuzione dei dati. I metodi statistici classici usano soglie basate sulla deviazione standard o sull'intervallo interquartile, segnalando come anomalo tutto ciò che si trova troppo lontano dal centro della distribuzione. I metodi basati sulla densità o sulla distanza, come l'isolation forest o il local outlier factor, individuano i punti che si trovano in regioni scarsamente popolate dello spazio delle caratteristiche rispetto ai loro vicini. La decisione su cosa fare una volta individuato un outlier, se rimuoverlo, correggerlo o studiarlo a parte, dipende sempre dal contesto: in un dataset di prezzi immobiliari un valore estremo può essere un errore, in un sistema antifrode è spesso proprio ciò che si sta cercando.
Nell'AI applicata la gestione degli outlier è cruciale per due ragioni opposte: da un lato molti algoritmi, come la regressione lineare o il k-means, sono sensibili ai valori estremi e possono essere distorti significativamente da pochi outlier non trattati; dall'altro, in ambiti come il rilevamento frodi, la sicurezza informatica e il controllo qualità industriale, l'obiettivo primario del sistema è proprio identificare gli outlier, che rappresentano l'evento raro da catturare.
Il termine outlier, letteralmente ciò che giace fuori, out-lier, entra nel vocabolario statistico nell'Ottocento, quando matematici e astronomi iniziarono a occuparsi sistematicamente di come trattare osservazioni sperimentali che si discostavano nettamente dalle altre misurazioni. La statistica del Novecento ha poi sviluppato criteri formali sempre più raffinati per distinguere un vero errore di misurazione da un'osservazione rara ma genuina, una distinzione che resta al cuore della pratica del machine learning ancora oggi.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.