Production Site Reliability Engineering (Paperback)
Lingua: inglese
Editore: Independently Published, 2026
- Brossura
- Nuovo

Da: Grand Eagle Retail, Bensenville, IL, U.S.A.Grand Eagle Retail
Venditore AbeBooks dal 12 ottobre 2005
Condizione: Nuovo
EUR 23,50
Quantità: 1 disponibili
Aggiungi al carrelloDescrizione dell’articolo da parte del venditore
Paperback. Your system is running. But is it actually reliable?A service can pass every test, deploy successfully, and look healthy on a dashboard-then collapse when traffic spikes, a dependency slows down, Kubernetes reschedules workloads, a bad release reaches production, or retries turn a small failure into a cascading outage.If you work with modern production systems, you already know the challenge. Keeping software running is not the same as engineering reliability. Whether you're exploring SRE for DevOps engineers, platform engineering, cloud operations, or backend development, you need to know what to measure, when to alert, how to diagnose failures, how to contain their blast radius, and how to recover without relying on guesswork.Production Site Reliability Engineering provides a practical, systems-focused approach to site reliability engineering and production reliability engineering, showing you how to design, operate, troubleshoot, and continuously improve dependable production environments. Instead of treating SRE as a collection of disconnected tools, this hands-on guide brings together cloud native reliability, observability engineering, Kubernetes production operations, distributed systems reliability, production incident management, and safe automation as parts of one coherent reliability system.Inside, you'll learn how to: Design meaningful SLIs, SLOs, error budgets, and burn-rate alerts, and apply SLO error budget monitoring to real user journeys and engineering decisions.Build production observability with Prometheus metrics, structured logs, distributed tracing, OpenTelemetry, RED and USE signals, and evidence-driven dashboards.Diagnose and manage production incidents systematically using timelines, telemetry correlation, failure-domain isolation, hypotheses, mitigation, recovery verification, runbooks, and postmortems.Engineer reliable Kubernetes workloads with health probes, resource controls, graceful termination, workload QoS, autoscaling, disruption budgets, stateful workloads, and resilient placement.Prevent cascading distributed-system failures using deadlines, timeouts, retry budgets, exponential backoff, jitter, circuit breakers, bulkheads, load shedding, and graceful degradation.Design reliable asynchronous systems with queues, backpressure, idempotency, dead-letter handling, replay, and controlled recovery.Plan capacity and engineer performance using throughput, concurrency, saturation, tail latency, load testing, forecasting, autoscaling, and production headroom.Validate resilience before disaster strikes through failure injection, chaos engineering, disaster recovery, game days, backup and restore verification, RTO, RPO, and recovery drills.Reduce deployment and operational risk with CI/CD reliability gates, rolling and blue-green deployments, canary releases, progressive delivery, automated remediation, operational guardrails, and AI-assisted SRE.Apply these practices to an evolving Kubernetes-based cloud-native commerce system, connecting reliability requirements to realistic production failures and evidence-driven engineering decisions.This isn't a book about achieving reliability by adding more dashboards, replicas, alerts, or automation. Reliability is demonstrated by how your system behaves when production stops behaving as expected-and by your ability to prove, diagnose, and improve that behavior.Ready to build production systems that survive failure, recover safely, and improve through evidence? Get your copy of Production Site R Shipping may be from multiple locations in the US or from the UK, depending on stock availability.…
Codice articolo 9798170223053
- Titolo
- Production Site Reliability Engineering (Paperback)
- Autore
- Nolan Veyne
- Editore
- Independently Published
- Anno di pubblicazione
- 2026
- Condizione
- new
- Rilegatura
- Paperback
- Lingua
- inglese
- ISBN 13
- 9798170223053
Your system is running. But is it actually reliable?
A service can pass every test, deploy successfully, and look healthy on a dashboard—then collapse when traffic spikes, a dependency slows down, Kubernetes reschedules workloads, a bad release reaches production, or retries turn a small failure into a cascading outage.
If you work with modern production systems, you already know the challenge. Keeping software running is not the same as engineering reliability. Whether you're exploring SRE for DevOps engineers, platform engineering, cloud operations, or backend development, you need to know what to measure, when to alert, how to diagnose failures, how to contain their blast radius, and how to recover without relying on guesswork.
Production Site Reliability Engineering provides a practical, systems-focused approach to site reliability engineering and production reliability engineering, showing you how to design, operate, troubleshoot, and continuously improve dependable production environments. Instead of treating SRE as a collection of disconnected tools, this hands-on guide brings together cloud native reliability, observability engineering, Kubernetes production operations, distributed systems reliability, production incident management, and safe automation as parts of one coherent reliability system.
Inside, you'll learn how to:
-
Design meaningful SLIs, SLOs, error budgets, and burn-rate alerts, and apply SLO error budget monitoring to real user journeys and engineering decisions.
-
Build production observability with Prometheus metrics, structured logs, distributed tracing, OpenTelemetry, RED and USE signals, and evidence-driven dashboards.
-
Diagnose and manage production incidents systematically using timelines, telemetry correlation, failure-domain isolation, hypotheses, mitigation, recovery verification, runbooks, and postmortems.
-
Engineer reliable Kubernetes workloads with health probes, resource controls, graceful termination, workload QoS, autoscaling, disruption budgets, stateful workloads, and resilient placement.
-
Prevent cascading distributed-system failures using deadlines, timeouts, retry budgets, exponential backoff, jitter, circuit breakers, bulkheads, load shedding, and graceful degradation.
-
Design reliable asynchronous systems with queues, backpressure, idempotency, dead-letter handling, replay, and controlled recovery.
-
Plan capacity and engineer performance using throughput, concurrency, saturation, tail latency, load testing, forecasting, autoscaling, and production headroom.
-
Validate resilience before disaster strikes through failure injection, chaos engineering, disaster recovery, game days, backup and restore verification, RTO, RPO, and recovery drills.
-
Reduce deployment and operational risk with CI/CD reliability gates, rolling and blue-green deployments, canary releases, progressive delivery, automated remediation, operational guardrails, and AI-assisted SRE.
-
Apply these practices to an evolving Kubernetes-based cloud-native commerce system, connecting reliability requirements to realistic production failures and evidence-driven engineering decisions.
This isn't a book about achieving reliability by adding more dashboards, replicas, alerts, or automation. Reliability is demonstrated by how your system behaves when production stops behaving as expected—and by your ability to prove, diagnose, and improve that behavior.
Ready to build production systems that survive failure, recover safely, and improve through evidence? Get your copy of Production Site Reliability Engineering today.
"Riassunto" può appartenere a un’altra edizione di questo titolo.
Grand Eagle Retail
Bensenville, IL, U.S.A.
Venditore AbeBooks dal 12 ottobre 2005
Tariffe di spedizione nazionale per U.S.A.
| Articolo | Da 6 a 14 giorni lavorativi | Da 6 a 16 giorni lavorativi |
|---|---|---|
| Primo articolo | EUR 0,00 | EUR 0,00 |
Metodi di pagamento
Informazioni sull’azienda del venditore
APOLLO ONLINE CORP.
605 Geddes Street
Wilmington, DE U.S.A. 19805
Condizioni di vendita
We guarantee the condition of every book as it¿s described on the Abebooks web sites. If you¿ve changed
your mind about a book that you¿ve ordered, please use the Ask bookseller a question link to contact us
and we¿ll respond within 2 business days.
Books ship from California and Michigan.
Diritto di recesso
Se sei un consumatore puoi recedere dal contratto in conformità con quanto segue. Per Consumatore si intende qualsiasi persona fisica che agisce per scopi estranei alla propria attività commerciale, imprenditoriale, artigianale o professionale.
Informazioni sul diritto di recesso
Diritto legale di recesso
Hai il diritto di recedere dal presente contratto entro 14 giorni senza fornire alcuna motivazione.
Il periodo di recesso scade dopo 14 giorni dal giorno in cui tu o una terza parte, diversa dal vettore e da te indicata, acquisisce il possesso fisico dell'ultimo bene o dell'ultimo lotto o pezzo.
Per esercitare il diritto di recesso, compila e invia elettronicamente una dichiarazione esplicita sul nostro sito Web, alla voce “I miei acquisti” nella sezione “Mio account”. Ti comunicheremo senza indugio una conferma di ricezione di tale recesso su un supporto durevole (ad es. via e-mail).
Per rispettare il termine di recesso, è sufficiente inviare la comunicazione relativa all'esercizio del diritto di recesso prima della scadenza del periodo di recesso stesso.
Effetti del recesso
In caso di recesso dal presente contratto, ti rimborseremo tutti i pagamenti ricevuti, compresi i costi di spedizione (ad eccezione dei costi supplementari derivanti dalla tua eventuale scelta di un tipo di spedizione diverso dal tipo meno costoso di consegna standard da noi offerto).
Potremo effettuare una detrazione dal rimborso per la perdita di valore dei beni forniti, qualora tale perdita sia il risultato di una manipolazione non necessaria da parte tua.
Eseguiremo il rimborso senza indebito ritardo e non oltre 14 giorni dal giorno in cui saremo informati della tua decisione di recedere dal presente contratto.
Il rimborso sarà effettuato utilizzando lo stesso mezzo di pagamento da te usato per la transazione iniziale, salvo che tu non abbia espressamente concordato altrimenti; in ogni caso, non dovrai sostenere alcun costo quale conseguenza di tale rimborso.
Possiamo trattenere il rimborso finché non avremo ricevuto i beni oppure finché non avrai fornito la prova di averli rispediti, a seconda di quale condizione si verifichi per prima.
Dovrai rispedire i beni o consegnarli a Grand Eagle Retail, Bensenville, Illinois, U.S.A., senza indebito ritardo e, in ogni caso, entro 14 giorni dal giorno in cui ci hai comunicato la tua volontà di recedere dal presente contratto. Il termine è rispettato se rispedisci i beni prima della scadenza del periodo di 14 giorni. I costi diretti della restituzione dei beni saranno a tuo carico. Sei responsabile solo della diminuzione del valore dei beni risultante da una manipolazione diversa da quella necessaria per stabilire la natura, le caratteristiche e il funzionamento dei beni stessi.
Eccezioni al diritto di recesso
Il diritto di recesso non si applica a:
- La fornitura di giornali, periodici o riviste ad eccezione dei contratti di abbonamento; e
- La fornitura di contenuto digitale non fornito su un supporto materiale (ad es. su un CD o DVD), se al momento dell'invio dell'ordine hai accettato l'inizio dell'esecuzione e hai riconosciuto che non avresti potuto recedere una volta iniziata l'esecuzione.
Condizioni di spedizione
Orders usually ship within 2 business days. All books within the US ship free of charge. Delivery is 4-14 business days anywhere in the United States.
Books ship from California and Michigan.
If your book order is heavy or oversized, we may contact you to let you know extra shipping is required.