AI Dictionary › Fondamenti AI
Uptime e SLA
Uptime is the percentage of time a service is operational and reachable; an SLA (Service Level Agreement) is the contract that fixes the promised service level, uptime included, and the compensation owed if the promise is missed. Together they turn reliability from a vague aspiration into a measurable, binding commitment.
Uptime is the percentage of time a service is operational and reachable; an SLA (Service Level Agreement) is the contract that fixes the promised service level, uptime included, and the compensation owed if the promise is missed. Together they turn reliability from a vague aspiration into a measurable, binding commitment.
Uptime is expressed in nines: 99.9% (three nines) allows about 8 hours 46 minutes of downtime per year, 99.99% about 53 minutes, 99.999% barely 5 minutes. Each extra nine costs exponentially more in redundancy and engineering. Alongside the SLA live SLOs (internal targets, typically stricter than the contract) and SLIs (the metrics actually measured). The error budget, the downtime still spendable, guides the balance between fast releases and stability.
In AI, SLAs are now central: LLM providers publish API availability commitments for enterprise plans, and anyone building products on those models must compose vendor SLAs with their own, planning failover and graceful degradation. Beyond AI, SLAs have governed hosting, cloud and telecommunications for decades.
Uptime comes from mainframe jargon, the time the machine is up, as opposed to downtime. Formal service level agreements took hold in the 1980s and 90s with IT outsourcing and telecommunications, when a contractual way was needed to define how reliable a purchased service had to be. The SLO and error budget culture was codified by Google's Site Reliability Engineering, popularized in the 2016 book of the same name.
L'uptime è la percentuale di tempo in cui un servizio è operativo e raggiungibile; lo SLA (Service Level Agreement) è il contratto che fissa il livello di servizio promesso, uptime incluso, e le compensazioni dovute se la promessa viene mancata. Insieme trasformano l'affidabilità da aspirazione vaga a impegno misurabile e vincolante.
L'uptime si esprime in nove: 99,9% (tre nove) ammette circa 8 ore e 46 minuti di fermo l'anno, 99,99% circa 53 minuti, 99,999% appena 5 minuti. Ogni nove in più costa esponenzialmente in ridondanza e ingegneria. Accanto allo SLA vivono gli SLO (obiettivi interni, tipicamente più severi del contratto) e gli SLI (le metriche effettivamente misurate). L'error budget, la quota di fermo ancora spendibile, guida il bilanciamento tra rilasci rapidi e stabilità.
Nell'AI gli SLA sono ormai centrali: i provider LLM pubblicano impegni di disponibilità delle API per i piani enterprise, e chi costruisce prodotti su quei modelli deve comporre gli SLA dei fornitori con i propri, prevedendo failover e degradazione controllata. Fuori dall'AI, gli SLA regolano hosting, cloud e telecomunicazioni da decenni.
Uptime nasce dal gergo dei mainframe, il tempo in cui la macchina è su, contrapposto al downtime. Gli accordi formali sul livello di servizio si affermarono negli anni '80 e '90 con l'outsourcing IT e le telecomunicazioni, quando serviva un modo contrattuale per definire quanto affidabile dovesse essere un servizio comprato da terzi. La cultura degli SLO ed error budget è stata codificata dal Site Reliability Engineering di Google, divulgato nel libro omonimo del 2016.
From our network
INDACO TMS — Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.