AI Dictionary › AI Fundamentals
Data labeling (Etichettatura dei dati)
Data labeling is the process of assigning to each example in a dataset the information a supervised learning model must learn to predict: the correct category of an image, the sentiment expressed in a review, the boundaries of an object in a photograph, the correct transcription of an audio clip. It is the work that makes supervised learning possible: without reliable labels, an algorithm has nothing to learn from in a targeted way.
Labeling can be done by people, often through dedicated platforms that distribute small tasks to large groups of annotators, by domain experts when specialized knowledge is needed, as in medical diagnostics, or partly automated with pre-labeling techniques assisted by existing models, which a human then corrects instead of annotating from scratch. Label quality matters as much as quantity: ambiguous instructions or poorly trained annotators produce inconsistent labels, which translate directly into a less accurate model, no matter how sophisticated the chosen algorithm is.
It is an essential step in computer vision, where millions of images need labeling to train recognition systems, in natural language processing, where texts need classifying by sentiment, intent or category, and in training autonomous driving systems, where every video frame requires precise identification of vehicles, pedestrians and road signs. It is often the most expensive and slowest bottleneck of an entire supervised machine learning project, so much so that it is estimated to represent a significant share of the overall budget of many applied AI projects.
The term literally describes the operation: attaching a label to each data point. The practice of manually annotating data for training recognition systems dates back to the earliest pattern recognition experiments of the second half of the twentieth century, but it became a genuine industry with the explosion of deep learning starting in the 2010s, when the availability of huge, carefully labeled datasets, like ImageNet, proved as decisive as computing power for the progress of computer vision.
Every Grace professional scenario is labeled with detail, industry, difficulty and the skills involved, editorial work that closely resembles data labeling: only with accurate labels can the system offer you scenarios that are truly relevant to your path.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.