AI Dictionary › Fondamenti AI
A data pipeline is an automated sequence of steps that moves and transforms data from source to destination: collection, cleaning, enrichment, loading. Like an aqueduct for information, it makes raw data flow from its points of origin to the systems where it becomes useful, in a repeatable, controlled way.
A data pipeline is an automated sequence of steps that moves and transforms data from source to destination: collection, cleaning, enrichment, loading. Like an aqueduct for information, it makes raw data flow from its points of origin to the systems where it becomes useful, in a repeatable, controlled way.
Each stage receives the previous stage's output: a connector extracts data from a database or API, successive transformations validate, normalize and aggregate it, and the result is loaded into a data warehouse, data lake or index. Orchestrators like Apache Airflow, Prefect or Dagster handle dependencies, scheduling, retries on failure and monitoring.
In AI, pipelines are the circulatory system: they prepare training datasets (deduplication, filtering, tokenization), feed RAG vector indexes with fresh documents and carry production logs to model evaluation systems. Beyond AI, they move sales, sensor and transaction data into business reports and dashboards.
The metaphor comes from the oil pipeline, the conduit carrying crude through successive stages. In computing the concept has a noble father: Douglas McIlroy of Bell Labs, who in 1973 introduced Unix pipes, the mechanism for chaining programs by streaming one's output into the next one's input. From that idea, composing simple steps into complex flows, today's data pipelines descend.
Una data pipeline è una sequenza automatizzata di passaggi che sposta e trasforma i dati da una sorgente a una destinazione: raccolta, pulizia, arricchimento, caricamento. Come un acquedotto per le informazioni, fa fluire i dati grezzi dai punti di origine fino ai sistemi dove diventano utili, in modo ripetibile e controllato.
Ogni stadio della pipeline riceve l'output del precedente: un connettore estrae i dati da un database o un'API, trasformazioni successive li validano, li normalizzano e li aggregano, infine il risultato viene caricato in un data warehouse, un data lake o un indice. Orchestratori come Apache Airflow, Prefect o Dagster gestiscono dipendenze, pianificazione, retry sui fallimenti e monitoraggio.
Nell'AI le pipeline sono il sistema circolatorio: preparano i dataset di addestramento (deduplicazione, filtraggio, tokenizzazione), alimentano gli indici vettoriali del RAG con documenti aggiornati e trasportano i log di produzione verso i sistemi di valutazione dei modelli. Fuori dall'AI muovono i dati di vendite, sensori e transazioni verso report e dashboard aziendali.
La metafora viene dalla pipeline petrolifera, la conduttura che trasporta greggio attraverso stadi successivi. In informatica il concetto ha un padre nobile: Douglas McIlroy dei Bell Labs, che nel 1973 introdusse le pipe di Unix, il meccanismo per concatenare programmi facendo scorrere l'output dell'uno nell'input del successivo. Da quell'idea, comporre passaggi semplici in flussi complessi, discendono le moderne pipeline di dati.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.