AI Dictionary › Fondamenti AI

Data augmentation

Data augmentation is the technique of generating new training examples from those already available, by applying transformations that change their appearance without altering their essential content. It serves to make a training dataset larger and more varied without having to collect and label new data from scratch, an operation that is often expensive or simply impractical at scale.

Definition

What it is

Data augmentation is the technique of generating new training examples from those already available, by applying transformations that change their appearance without altering their essential content. It serves to make a training dataset larger and more varied without having to collect and label new data from scratch, an operation that is often expensive or simply impractical at scale.

How it works

In computer vision, typical transformations include rotating, flipping, cropping, changing brightness and contrast, or adding noise to an image, while keeping the original label unchanged, a cat rotated by ten degrees is still a cat. In language processing, techniques include replacing words with synonyms, back-translation between two languages to obtain a differently worded version of the same text, or the controlled insertion of small variations. In audio, pitch, speed changes or the addition of background noise are applied. The idea common to all variants is to teach the model to recognize what truly matters, ignoring surface-level variations that should not affect the prediction.

Applications

It is a fundamental practice in deep learning, where models have millions or billions of parameters and require large amounts of data to avoid overfitting: nearly every image training pipeline includes some form of augmentation. It is especially valuable in fields where real data is scarce or expensive to obtain, such as medical imaging diagnostics, where it is not always possible to collect thousands of rare cases, or robotics, where simulating environmental variations reduces the need for physical experimentation.

History & etymology

The earliest systematic applications of data augmentation in image recognition date back to the 1990s and early 2000s, but the technique became an established standard with the deep learning renaissance starting in 2012, when the AlexNet architecture, which dominantly won that year's ImageNet competition, extensively used image cropping and reflection transformations to artificially multiply the amount of available training data and reduce overfitting on a network with millions of parameters.

Definizione (italiano)

Il data augmentation, o aumento dei dati, è la tecnica che genera nuovi esempi di addestramento a partire da quelli già disponibili, applicando trasformazioni che ne cambiano l'aspetto senza alterarne il contenuto essenziale. Serve a rendere un dataset di addestramento più grande e più vario senza dover raccogliere ed etichettare nuovi dati da zero, un'operazione spesso costosa o semplicemente non praticabile su larga scala.

Nella visione artificiale le trasformazioni tipiche includono ruotare, ribaltare, ritagliare, cambiare luminosità e contrasto o aggiungere rumore a un'immagine, mantenendo però l'etichetta originale invariata, un gatto ruotato di dieci gradi resta un gatto. Nell'elaborazione del linguaggio si usano tecniche come la sostituzione di parole con sinonimi, la traduzione andata e ritorno tra due lingue per ottenere una formulazione diversa dello stesso testo, o l'inserimento controllato di piccole variazioni. In ambito audio si applicano cambi di tono, velocità o l'aggiunta di rumore di fondo. L'idea comune a tutte le varianti è insegnare al modello a riconoscere ciò che conta davvero, ignorando le variazioni superficiali che non dovrebbero influenzare la previsione.

È una pratica fondamentale nel deep learning, dove i modelli hanno milioni o miliardi di parametri e richiedono grandi quantità di dati per evitare l'overfitting: quasi ogni pipeline di addestramento per immagini include qualche forma di augmentation. È particolarmente preziosa in ambiti dove i dati reali sono scarsi o costosi da ottenere, come la diagnostica medica per immagini, dove non è sempre possibile raccogliere migliaia di casi rari, o la robotica, dove simulare variazioni dell'ambiente riduce la necessità di sperimentazione fisica.

Le prime applicazioni sistematiche del data augmentation nel riconoscimento di immagini risalgono agli anni '90 e primi 2000, ma la tecnica diventa uno standard consolidato con la rinascita del deep learning a partire dal 2012, quando l'architettura AlexNet, che vinse in modo dominante la competizione ImageNet di quell'anno, utilizzò estensivamente trasformazioni di ritaglio e riflessione delle immagini per moltiplicare artificialmente la quantità di dati di addestramento disponibili e ridurre l'overfitting su una rete con milioni di parametri.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.