Metrics-Driven Reliability
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Create account & apply
An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.
Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.
Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.
Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.
Modules stay collapsed for quick scanning. Open any module to inspect its topics.
What is SRE? Evolution from Operations to DevOps to SRE
DevOps vs SRE and Reliability Engineering Principles
Production Mindset and Service Lifecycle
Shared Responsibility Model, Toil, and Automation
SRE Roles and Responsibilities
Lab: Calculate Availability and Identify Toil
Why Reliability Needs Metrics: Availability vs Reliability
Service Level Indicators (SLI) and Objectives (SLO)
Service Level Agreements (SLA) and Error Budgets
MTTD, MTTR, and MTBF
DORA Metrics for Delivery Performance
Lab: Define SLOs and Calculate Error Budgets
Monitoring vs Observability and the Three Pillars
Metrics, Logs, and Traces Fundamentals
Golden Signals, RED Method, and USE Method
Metrics Collection with Prometheus
Visualizing System Health with Grafana and Distributed Tracing
Lab: Analyze Dashboards and Detect Bottlenecks
Alerting Principles and Prometheus Alertmanager
Alert Fatigue, Severity, and Prioritization
Incident Lifecycle and Response Process
On-call Best Practices and Blameless Postmortems
Root Cause Analysis and Writing Effective Runbooks
Lab: Simulate Incidents, Perform RCA, and Write a Postmortem
Self-Healing, Liveness, Readiness, and Startup Probes
Resource Requests, Limits, and Horizontal/Vertical Autoscaling
Pod Disruption Budgets and Scheduling for Reliability
Node Affinity, Anti-Affinity, and Topology Spread Constraints
Reliable Deployment Strategies: Rolling, Blue-Green, Canary
Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA
CI/CD Reliability and GitOps Principles
Progressive Delivery and Safe Deployment Practices
Rollback vs Roll-forward and Release Management
Configuration and Secret Management
Change Management, Deployment Verification, and Supply Chain Security
Lab: GitOps Sync and Safe Production Release Simulation
Disaster Recovery Fundamentals and High Availability Architecture
Backup & Restore, RTO, and RPO
Single Point of Failure and Capacity Planning Review
Load Testing Concepts and Chaos Engineering
Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light
Lab: Restore from Backup and Run Chaos Testing Scenarios
End-to-End Troubleshooting Methodology
Common Production Failures and Case Studies
Cost vs Reliability Trade-offs
AI for SRE: AIOps, AI-assisted Troubleshooting and RCA
SRE Career Roadmap and Continuous Improvement
Capstone Lab: Full Production Incident Simulation
Only current, open groups are shown.
“The program helped me connect individual skills into the way real teams design, build and deliver software.”
See real graduate stories from the Ingress community.
Explore graduate resultsWe use one Portal account for applications, assessments and future learning progress. You will not need to email your details or select the training again.
Questions first? Talk to an advisorCreate or sign in to your Portal accountYour contact details stay connected to one student profile.
Confirm your application detailsSite Reliability Engineering (SRE) Bootcamp is preselected.
Submit your applicationThe admissions team receives it immediately and can follow up from the Portal.
SELECTED TRAININGSite Reliability Engineering (SRE) BootcampOnline
Continue in Ingress Portal Already registered? The Portal will let you sign in instead.