AI Dictionary › Fondamenti AI

Data cleaning

Data cleaning is the set of operations used to identify and correct errors, inconsistencies, duplicates and missing values in a dataset before using it to train a machine learning model. It is often the least glamorous but most decisive phase of a data project: a sophisticated model trained on dirty data produces unreliable results, regardless of how advanced the chosen algorithm is.

Definition

What it is

Data cleaning is the set of operations used to identify and correct errors, inconsistencies, duplicates and missing values in a dataset before using it to train a machine learning model. It is often the least glamorous but most decisive phase of a data project: a sophisticated model trained on dirty data produces unreliable results, regardless of how advanced the chosen algorithm is.

How it works

The process includes several typical activities: identifying and handling missing values, removing or correcting duplicate records, standardizing inconsistent formats such as dates written in different ways, spotting clearly wrong or out-of-range values, such as an age of 200 years, and checking consistency between related fields. It often also requires standardizing units of measurement, correcting typos in text fields, and resolving ambiguities in category encoding, for example when the same city appears written in several different ways within the same dataset.

Applications

In applied AI, data cleaning almost always precedes feature engineering and is responsible, according to various industry estimates, for most of the time spent on a typical data science project, often more than the time devoted to choosing and training the actual model. It is especially critical in regulated sectors like healthcare and finance, where dirty data can translate into flawed automated decisions with direct consequences on real people.

History & etymology

The concept of checking and correcting data quality belongs to applied statistics and enterprise data processing well before machine learning, with roots in the quality control practices of censuses and statistical surveys throughout the twentieth century. With the exponential growth of automatically collected data volumes from the 1990s onward, data cleaning established itself as a distinct discipline within the broader field of data governance, with dedicated tools and methodologies.

Definizione (italiano)

Il data cleaning, o pulizia dei dati, è l'insieme delle operazioni con cui si individuano e si correggono errori, incoerenze, duplicati e valori mancanti in un dataset prima di usarlo per addestrare un modello di machine learning. È spesso la fase meno appariscente ma più determinante di un progetto di dati: un modello sofisticato addestrato su dati sporchi produce risultati inaffidabili, indipendentemente da quanto sia avanzato l'algoritmo scelto.

Il processo comprende diverse attività tipiche: individuare e gestire i valori mancanti, rimuovere o correggere i record duplicati, uniformare formati incoerenti come date scritte in modi diversi, individuare valori palesemente errati o fuori scala, come un'età di 200 anni, e verificare la coerenza tra campi collegati tra loro. Spesso richiede anche di uniformare le unità di misura, correggere errori di battitura in campi testuali e risolvere ambiguità nella codifica delle categorie, per esempio quando la stessa città compare scritta in più modi diversi nello stesso dataset.

Nell'AI applicata il data cleaning precede quasi sempre il feature engineering ed è responsabile, secondo diverse stime di settore, della maggior parte del tempo speso in un tipico progetto di data science, spesso più del tempo dedicato alla scelta e all'addestramento del modello vero e proprio. È particolarmente critico in settori regolamentati come sanità e finanza, dove dati sporchi possono tradursi in decisioni automatizzate scorrette con conseguenze dirette su persone reali.

Il concetto di verificare e correggere la qualità dei dati appartiene alla statistica applicata e all'elaborazione dati aziendale ben prima del machine learning, con radici nelle pratiche di controllo qualità dei censimenti e delle rilevazioni statistiche del Novecento. Con la crescita esponenziale dei volumi di dati raccolti automaticamente dagli anni '90 in poi, il data cleaning si è affermato come disciplina distinta all'interno della più ampia data governance, con strumenti e metodologie dedicate.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.