AI Dictionary › Fondamenti AI

Failover

Failover is the automatic switch to a backup system when the primary one stops working. It is the mechanism that turns a potentially catastrophic outage into a brief transition invisible to users. Together with redundancy, it is the pillar of high availability: spare components are of little use if switching to them requires manual intervention.

Definition

What it is

Failover is the automatic switch to a backup system when the primary one stops working. It is the mechanism that turns a potentially catastrophic outage into a brief transition invisible to users. Together with redundancy, it is the pillar of high availability: spare components are of little use if switching to them requires manual intervention.

How it works

The primary and the standby monitor each other through periodic signals called heartbeats. When monitoring detects that the primary has been silent beyond a threshold, it promotes the standby to new primary, redirects traffic and alerts operators. Typical setups are active-passive, where the backup waits on standby, and active-active, where both nodes work and either absorbs the other's load on failure.

Applications

In AI platforms failover applies at several layers: model replicas across cloud regions, redundant vector databases, and LLM gateways that reroute calls to an alternative provider when the main one returns errors. Beyond AI it protects banking databases, phone exchanges, air traffic control systems and any service where minutes of downtime are expensive.

History & etymology

The term combines fail and over: literally, moving past the failure. The practice is rooted in fault-tolerant systems engineering: a historical reference is Tandem Computers, which from 1976 built fully redundant NonStop systems for banks and stock exchanges, where every component had a twin ready to take over.

Definizione (italiano)

Il failover è il passaggio automatico a un sistema di riserva quando quello principale smette di funzionare. È il meccanismo che trasforma un guasto potenzialmente catastrofico in una breve transizione invisibile all'utente. Insieme alla ridondanza, è il pilastro dell'alta disponibilità: avere componenti di scorta serve a poco se il passaggio richiede un intervento manuale.

Il sistema primario e quello di riserva si controllano a vicenda tramite segnali periodici detti heartbeat. Quando il monitoraggio rileva che il primario tace oltre una soglia, promuove la riserva a nuovo primario, ridirige il traffico e notifica gli operatori. Le configurazioni tipiche sono attivo-passivo, dove la riserva attende in standby, e attivo-attivo, dove entrambi i nodi lavorano e uno assorbe il carico dell'altro in caso di guasto.

Nelle piattaforme AI il failover si applica a più livelli: repliche dei modelli su regioni cloud diverse, database vettoriali ridondati e gateway LLM che dirottano le chiamate su un provider alternativo quando il principale restituisce errori. Fuori dall'AI protegge database bancari, centralini telefonici, sistemi di controllo del traffico aereo e qualunque servizio dove i minuti di fermo costano cari.

Il termine combina fail (guasto) e over (passaggio): letteralmente, passare oltre il guasto. La pratica affonda le radici nell'ingegneria dei sistemi fault-tolerant: un riferimento storico è Tandem Computers, che dal 1976 con i sistemi NonStop costruì computer interamente ridondati per banche e borse, dove ogni componente aveva un gemello pronto a subentrare.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.