Production Reliability Engineering (SRE) Course | Ingress Ac
Hesab yarat və müraciət et
Site Reliability Engineer DEVOPS & LINUX ENGINEERING · İRƏLI SƏVİYYƏ

Site Reliability Engineering (SRE) Bootcamp

An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.

İrəli8 həftə48 saatOnlayn
Seçdiyiniz kurs Ingress Portala köçürüləcək.
KURSUN ÜSTÜNLÜKLƏRİ
📊

Metrics-Driven Reliability

Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.

🚨

Real Incident Simulations

Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.

☸️

Kubernetes-Native Reliability

Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.

Hazır olduğunuza əmin deyilsiniz? Bacarıq qiymətləndirməsi keçin
SƏVİYYƏNİZ
BaşlanğıcOrtaİrəliEkspert
NƏ EDƏ BİLƏCƏKSİNİZ

Bu kursun sonunda sahib olacağınız bacarıqlar.

  • Define and calculate SLIs, SLOs, SLAs, and error budgets for production services
  • Build and interpret observability stacks using Prometheus, Grafana, logs, and traces
  • Design effective alerting strategies and manage the full incident lifecycle
  • Apply Kubernetes reliability patterns including probes, autoscaling, PDBs, and deployment strategies
  • Implement GitOps-based, safe software delivery with rollback and progressive delivery techniques
  • Plan disaster recovery strategies and execute chaos engineering and load testing exercises
  • Lead root cause analysis and produce blameless postmortems and runbooks
  • Leverage AI-assisted tools for troubleshooting, RCA, and incident response
QISALDILMIŞ KURİKULUM

Dərslik oxumadan strukturu görün.

Modullar sürətli baxış üçün yığcam qalır. Mövzularına baxmaq üçün istənilən modulu açın.

01Introduction to Site Reliability Engineering6 dərs+

What is SRE? Evolution from Operations to DevOps to SRE

DevOps vs SRE and Reliability Engineering Principles

Production Mindset and Service Lifecycle

Shared Responsibility Model, Toil, and Automation

SRE Roles and Responsibilities

Lab: Calculate Availability and Identify Toil

02Measuring Service Reliability6 dərs+

Why Reliability Needs Metrics: Availability vs Reliability

Service Level Indicators (SLI) and Objectives (SLO)

Service Level Agreements (SLA) and Error Budgets

MTTD, MTTR, and MTBF

DORA Metrics for Delivery Performance

Lab: Define SLOs and Calculate Error Budgets

03Monitoring & Observability6 dərs+

Monitoring vs Observability and the Three Pillars

Metrics, Logs, and Traces Fundamentals

Golden Signals, RED Method, and USE Method

Metrics Collection with Prometheus

Visualizing System Health with Grafana and Distributed Tracing

Lab: Analyze Dashboards and Detect Bottlenecks

04Alerting & Incident Management6 dərs+

Alerting Principles and Prometheus Alertmanager

Alert Fatigue, Severity, and Prioritization

Incident Lifecycle and Response Process

On-call Best Practices and Blameless Postmortems

Root Cause Analysis and Writing Effective Runbooks

Lab: Simulate Incidents, Perform RCA, and Write a Postmortem

05Kubernetes Reliability Engineering6 dərs+

Self-Healing, Liveness, Readiness, and Startup Probes

Resource Requests, Limits, and Horizontal/Vertical Autoscaling

Pod Disruption Budgets and Scheduling for Reliability

Node Affinity, Anti-Affinity, and Topology Spread Constraints

Reliable Deployment Strategies: Rolling, Blue-Green, Canary

Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA

06Reliable Software Delivery6 dərs+

CI/CD Reliability and GitOps Principles

Progressive Delivery and Safe Deployment Practices

Rollback vs Roll-forward and Release Management

Configuration and Secret Management

Change Management, Deployment Verification, and Supply Chain Security

Lab: GitOps Sync and Safe Production Release Simulation

07Production Resilience & Disaster Recovery6 dərs+

Disaster Recovery Fundamentals and High Availability Architecture

Backup & Restore, RTO, and RPO

Single Point of Failure and Capacity Planning Review

Load Testing Concepts and Chaos Engineering

Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light

Lab: Restore from Backup and Run Chaos Testing Scenarios

08Modern SRE Operations & Production Scenarios6 dərs+

End-to-End Troubleshooting Methodology

Common Production Failures and Case Studies

Cost vs Reliability Trade-offs

AI for SRE: AIOps, AI-assisted Troubleshooting and RCA

SRE Career Roadmap and Continuous Improvement

Capstone Lab: Full Production Incident Simulation

YAXINLAŞAN QRUPLAR

Həqiqətən qatıla biləcəyiniz qrupu seçin.

Yalnız cari, açıq qruplar göstərilir.

KONTEKSTUAL SÜBUT
“The program helped me connect individual skills into the way real teams design, build and deliver software.”

Ingress icmasından real məzun hekayələrinə baxın.

Məzun nəticələrini kəşf et
INGRESS PORTAL VASİTƏSİLƏ MÜRACİƏT

Hesabınızı yaratdığınız müddətdə seçdiyiniz kurs qalır.

We use one Portal account for applications, assessments and future learning progress. You will not need to email your details or select the training again.

Əvvəlcə sualınız var? Məsləhətçi ilə danışın
  1. 01

    Portal hesabınızı yaradın və ya daxil olunƏlaqə məlumatlarınız vahid tələbə profilinə bağlı qalır.

  2. 02

    Müraciət məlumatlarınızı təsdiqləyinSite Reliability Engineering (SRE) Bootcamp əvvəlcədən seçilib.

  3. 03

    Müraciətinizi göndərinQəbul komandası müraciəti dərhal alır və Portal vasitəsilə əlaqə saxlaya bilər.

SEÇİLMİŞ TƏLİMSite Reliability Engineering (SRE) BootcampOnlayn

Ingress Portalda davam et Artıq qeydiyyatdan keçmisiniz? Portal daxil olmağınıza imkan verəcək.