AI Dictionary › Fondamenti AI
Data labeling (Etichettatura dei dati)
Data labeling is the process of assigning to each example in a dataset the information a supervised learning model must learn to predict: the correct category of an image, the sentiment expressed in a review, the boundaries of an object in a photograph, the correct transcription of an audio clip. It is the work that makes supervised learning possible: without reliable labels, an algorithm has nothing to learn from in a targeted way.
Data labeling is the process of assigning to each example in a dataset the information a supervised learning model must learn to predict: the correct category of an image, the sentiment expressed in a review, the boundaries of an object in a photograph, the correct transcription of an audio clip. It is the work that makes supervised learning possible: without reliable labels, an algorithm has nothing to learn from in a targeted way.
Labeling can be done by people, often through dedicated platforms that distribute small tasks to large groups of annotators, by domain experts when specialized knowledge is needed, as in medical diagnostics, or partly automated with pre-labeling techniques assisted by existing models, which a human then corrects instead of annotating from scratch. Label quality matters as much as quantity: ambiguous instructions or poorly trained annotators produce inconsistent labels, which translate directly into a less accurate model, no matter how sophisticated the chosen algorithm is.
It is an essential step in computer vision, where millions of images need labeling to train recognition systems, in natural language processing, where texts need classifying by sentiment, intent or category, and in training autonomous driving systems, where every video frame requires precise identification of vehicles, pedestrians and road signs. It is often the most expensive and slowest bottleneck of an entire supervised machine learning project, so much so that it is estimated to represent a significant share of the overall budget of many applied AI projects.
The term literally describes the operation: attaching a label to each data point. The practice of manually annotating data for training recognition systems dates back to the earliest pattern recognition experiments of the second half of the twentieth century, but it became a genuine industry with the explosion of deep learning starting in the 2010s, when the availability of huge, carefully labeled datasets, like ImageNet, proved as decisive as computing power for the progress of computer vision.
Il data labeling, o etichettatura dei dati, è il processo di assegnare a ciascun esempio di un dataset l'informazione che un modello di apprendimento supervisionato dovrà imparare a prevedere: la categoria corretta di un'immagine, il sentimento espresso in una recensione, i confini di un oggetto in una fotografia, la trascrizione corretta di un audio. È il lavoro che rende possibile l'apprendimento supervisionato: senza etichette affidabili, un algoritmo non ha nulla da cui imparare in modo mirato.
L'etichettatura può essere svolta da persone, spesso attraverso piattaforme dedicate che distribuiscono piccoli compiti a grandi gruppi di annotatori, da esperti di dominio quando servono competenze specialistiche, come nella diagnostica medica, o parzialmente automatizzata con tecniche di pre-etichettatura assistita da modelli già esistenti, che un umano poi corregge invece di annotare da zero. La qualità delle etichette è tanto importante quanto la quantità: istruzioni ambigue o annotatori poco formati producono etichette incoerenti, che si traducono direttamente in un modello meno accurato, per quanto sofisticato sia l'algoritmo scelto.
È un passaggio essenziale in computer vision, dove serve etichettare milioni di immagini per addestrare sistemi di riconoscimento, in elaborazione del linguaggio naturale, dove servono testi classificati per sentiment, intento o categoria, e nell'addestramento di sistemi di guida autonoma, dove ogni fotogramma video richiede l'identificazione precisa di veicoli, pedoni e segnaletica. È spesso il collo di bottiglia più costoso e lento di un intero progetto di machine learning supervisionato, tanto che si stima rappresenti una quota significativa del budget complessivo di molti progetti di AI applicata.
Il termine descrive letteralmente l'operazione: apporre un'etichetta, label, a ciascun dato. La pratica di annotare manualmente dati per l'addestramento di sistemi di riconoscimento risale ai primi esperimenti di pattern recognition della seconda metà del Novecento, ma diventa un'industria vera e propria con l'esplosione del deep learning a partire dagli anni 2010, quando la disponibilità di dataset enormi e accuratamente etichettati, come ImageNet, si rivelò determinante quanto la potenza di calcolo per i progressi della visione artificiale.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.