AI Dictionary › Fondamenti AI
A data lake is a centralized repository that stores huge volumes of data in its original, raw format: structured files, documents, images, logs, audio. Unlike a data warehouse, which requires defining structure before loading, the lake accepts everything immediately and defers interpretation to the moment of use (schema-on-read).
A data lake is a centralized repository that stores huge volumes of data in its original, raw format: structured files, documents, images, logs, audio. Unlike a data warehouse, which requires defining structure before loading, the lake accepts everything immediately and defers interpretation to the moment of use (schema-on-read).
Technically it sits on cheap, scalable object storage such as Amazon S3 or Azure Data Lake Storage, where data lands from ingestion pipelines without heavy transformation. A data catalog keeps track of what is there and where. The known risk is the data swamp: without governance and metadata, the lake turns into an unusable dump. Lakehouse architectures (Delta Lake, Iceberg) add warehouse-grade transactions and quality on top of lake storage.
For AI the data lake is the mine: multimodal corpora for training models, text, images, audio, live in lakes, and data preparation pipelines draw from them to filter and refine. Documents to be indexed for RAG often start there too. Beyond AI it serves exploratory analytics, long-term archiving and compliance.
The term was coined by James Dixon, then CTO of Pentaho, in a 2010 post: he compared the data mart to a bottle of water, cleaned, packaged and ready to drink, and the data lake to a natural lake where water flows in raw and everyone draws what they need. The metaphor took off with the big data explosion of the following decade.
Un data lake è un repository centralizzato che conserva enormi volumi di dati nel loro formato originale, grezzo: file strutturati, documenti, immagini, log, audio. A differenza del data warehouse, che richiede di definire la struttura prima di caricare, il lake accoglie tutto subito e rimanda l'interpretazione al momento dell'uso (schema-on-read).
Tecnicamente si appoggia a storage a oggetti economico e scalabile, come Amazon S3 o Azure Data Lake Storage, dove i dati arrivano da pipeline di ingestione senza trasformazioni pesanti. Un catalogo dati tiene traccia di cosa c'è e dove. Il rischio noto è il data swamp, la palude: senza governance e metadati, il lago si trasforma in un deposito inservibile. Le architetture lakehouse (Delta Lake, Iceberg) aggiungono transazioni e qualità da warehouse sopra lo storage da lake.
Per l'AI il data lake è la miniera: i corpora multimodali per addestrare i modelli, testo, immagini, audio, vivono in lake, e le pipeline di preparazione dati vi attingono per filtrare e raffinare. Anche i documenti da indicizzare per il RAG partono spesso da lì. Fuori dall'AI serve analytics esplorativa, archiviazione a lungo termine e conformità.
Il termine fu coniato da James Dixon, allora CTO di Pentaho, in un post del 2010: paragonò il data mart a una bottiglia d'acqua, pulita, confezionata e pronta al consumo, e il data lake a un lago naturale in cui l'acqua scorre grezza e ognuno attinge ciò che gli serve. La metafora fece fortuna con l'esplosione dei big data nel decennio successivo.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.