AI Dictionary › AI Fundamentals
Interpretabilità
Interpretability is the extent to which it is possible to understand, at a mechanical and internal level, why an AI model produces a given output, by observing what happens inside the neural network during computation. It should be distinguished from explainability, which focuses instead on providing an end user with an understandable account of the result: interpretability looks inside the model, explainability looks at the result the model exposes outward.
Interpretability techniques try to map the behavior of individual neurons, layers, or groups of parameters to recognizable concepts: which internal circuits activate for a certain type of input, which components contribute most to a specific decision, whether internal representations exist that correspond to ideas or externally verifiable facts. It is a research area technically far more complex than explainability, because modern models have billions of parameters whose interactions were not designed to be human-readable.
In modern AI, interpretability is central to frontier model safety research: understanding internal mechanisms allows latent behaviors to be spotted before they surface in production, verifying whether a model is actually reasoning or simply memorizing patterns, and surgically intervening on specific capabilities without retraining the entire system. It is also a diagnostic tool for debugging anomalous behavior in large language models.
The term has a longer history than explainability in classic machine learning, but it acquired a more specific and technical meaning with the rise of the research strand known as mechanistic interpretability, developed intensively starting in the early 2020s around large transformer models.
From our network
INDACO TMS: Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.