AI Dictionary › Fondamenti AI
Clustering is the unsupervised learning technique that groups a set of observations into subsets, called clusters, so that items within the same group are more similar to each other than to items in other groups. There are no labels to learn: the algorithm must discover a hidden similarity structure in the data on its own.
Clustering is the unsupervised learning technique that groups a set of observations into subsets, called clusters, so that items within the same group are more similar to each other than to items in other groups. There are no labels to learn: the algorithm must discover a hidden similarity structure in the data on its own.
Similarity is usually measured with a mathematical distance, Euclidean distance being the most common, computed over the numeric features of each observation. Different algorithm families exist: centroid-based ones like k-means, hierarchical ones that build a tree of successive groupings, and density-based ones like DBSCAN, capable of finding irregularly shaped clusters and isolating points that belong to no group at all.
In applied AI, clustering segments customers for targeted marketing campaigns, groups similar documents or articles to organize large archives, detects communities in a social network, compresses datasets by grouping redundant observations, and, in computer vision, helps organize similar images without manual labels. It is often the first exploratory step when tackling a new dataset that is not yet well understood.
The term cluster entered the statistical vocabulary in the mid-twentieth century as researchers began formalizing methods for grouping similar objects in taxonomy, psychology and biology. The expression cluster analysis appears in the scientific literature as early as the 1930s and spread widely in the 1960s with the growth of computational multivariate statistics, well before it became a standard machine learning tool.
Il clustering è la tecnica di apprendimento non supervisionato che raggruppa un insieme di osservazioni in sottoinsiemi, detti cluster, in modo che gli elementi dello stesso gruppo siano più simili tra loro rispetto a quelli di gruppi diversi. Non esistono etichette da imparare: l'algoritmo deve scoprire da solo una struttura di somiglianza nascosta nei dati.
La somiglianza si misura di solito con una distanza matematica, la distanza euclidea è la più comune, calcolata sulle caratteristiche numeriche di ciascuna osservazione. Esistono famiglie diverse di algoritmi: quelli basati su centroidi come k-means, quelli gerarchici che costruiscono un albero di raggruppamenti successivi, e quelli basati sulla densità come DBSCAN, capaci di trovare cluster di forma irregolare e di isolare i punti che non appartengono a nessun gruppo.
Nell'AI applicata il clustering segmenta la clientela per campagne di marketing mirate, raggruppa documenti o articoli simili per organizzare grandi archivi, individua comunità in una rete sociale, comprime dataset raggruppando osservazioni ridondanti e, nella computer vision, aiuta a organizzare immagini simili senza etichette manuali. È spesso il primo passo esplorativo quando si affronta un dataset nuovo e non ancora compreso a fondo.
Il termine cluster, grappolo o ammasso in inglese, entra nel vocabolario statistico a metà Novecento quando i ricercatori iniziarono a formalizzare metodi per raggruppare oggetti simili in tassonomia, psicologia e biologia. L'espressione cluster analysis compare nella letteratura scientifica già negli anni '30 e si diffonde ampiamente negli anni '60 con la crescita della statistica multivariata computazionale, ben prima che diventasse uno strumento standard del machine learning.
From our network
Kaimaki Web — Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.