Production Reliability Engineering (SRE) Course | Ingress Ac
Создать аккаунт и подать заявку
Site Reliability Engineer DEVOPS & LINUX ENGINEERING · ПРОДВИНУТЫЙ УРОВЕНЬ

Site Reliability Engineering (SRE) Bootcamp

An intensive 2-month, hands-on SRE program that teaches you how to measure, monitor, and defend production reliability in modern cloud-native and Kubernetes environments.

Продвинутый8 недели48 часовОнлайн
Выбранный курс будет перенесён в Ingress Portal.
ПРЕИМУЩЕСТВА КУРСА
📊

Metrics-Driven Reliability

Learn to define SLIs, SLOs, and error budgets that turn reliability into a measurable, actionable engineering practice.

🚨

Real Incident Simulations

Practice alerting, RCA, and blameless postmortems through realistic production incident scenarios, not just theory.

☸️

Kubernetes-Native Reliability

Apply advanced Kubernetes patterns like PDBs, affinity rules, and progressive rollouts to build resilient production systems.

Не уверены, что готовы? Пройдите бесплатную оценку навыков
ВАШ УРОВЕНЬ
НачальныйСреднийПродвинутыйЭксперт
ЧТО ВЫ СМОЖЕТЕ ДЕЛАТЬ

Навыки, которыми вы овладеете к концу этого курса.

  • Define and calculate SLIs, SLOs, SLAs, and error budgets for production services
  • Build and interpret observability stacks using Prometheus, Grafana, logs, and traces
  • Design effective alerting strategies and manage the full incident lifecycle
  • Apply Kubernetes reliability patterns including probes, autoscaling, PDBs, and deployment strategies
  • Implement GitOps-based, safe software delivery with rollback and progressive delivery techniques
  • Plan disaster recovery strategies and execute chaos engineering and load testing exercises
  • Lead root cause analysis and produce blameless postmortems and runbooks
  • Leverage AI-assisted tools for troubleshooting, RCA, and incident response
КРАТКАЯ ПРОГРАММА

Ознакомьтесь со структурой, не читая учебник.

Модули свёрнуты для быстрого просмотра. Откройте любой модуль, чтобы увидеть его темы.

01Introduction to Site Reliability Engineering6 уроков+

What is SRE? Evolution from Operations to DevOps to SRE

DevOps vs SRE and Reliability Engineering Principles

Production Mindset and Service Lifecycle

Shared Responsibility Model, Toil, and Automation

SRE Roles and Responsibilities

Lab: Calculate Availability and Identify Toil

02Measuring Service Reliability6 уроков+

Why Reliability Needs Metrics: Availability vs Reliability

Service Level Indicators (SLI) and Objectives (SLO)

Service Level Agreements (SLA) and Error Budgets

MTTD, MTTR, and MTBF

DORA Metrics for Delivery Performance

Lab: Define SLOs and Calculate Error Budgets

03Monitoring & Observability6 уроков+

Monitoring vs Observability and the Three Pillars

Metrics, Logs, and Traces Fundamentals

Golden Signals, RED Method, and USE Method

Metrics Collection with Prometheus

Visualizing System Health with Grafana and Distributed Tracing

Lab: Analyze Dashboards and Detect Bottlenecks

04Alerting & Incident Management6 уроков+

Alerting Principles and Prometheus Alertmanager

Alert Fatigue, Severity, and Prioritization

Incident Lifecycle and Response Process

On-call Best Practices and Blameless Postmortems

Root Cause Analysis and Writing Effective Runbooks

Lab: Simulate Incidents, Perform RCA, and Write a Postmortem

05Kubernetes Reliability Engineering6 уроков+

Self-Healing, Liveness, Readiness, and Startup Probes

Resource Requests, Limits, and Horizontal/Vertical Autoscaling

Pod Disruption Budgets and Scheduling for Reliability

Node Affinity, Anti-Affinity, and Topology Spread Constraints

Reliable Deployment Strategies: Rolling, Blue-Green, Canary

Lab: Troubleshoot Pods, Simulate Node Failure, and Test HPA

06Reliable Software Delivery6 уроков+

CI/CD Reliability and GitOps Principles

Progressive Delivery and Safe Deployment Practices

Rollback vs Roll-forward and Release Management

Configuration and Secret Management

Change Management, Deployment Verification, and Supply Chain Security

Lab: GitOps Sync and Safe Production Release Simulation

07Production Resilience & Disaster Recovery6 уроков+

Disaster Recovery Fundamentals and High Availability Architecture

Backup & Restore, RTO, and RPO

Single Point of Failure and Capacity Planning Review

Load Testing Concepts and Chaos Engineering

Recovery Strategies: Active-Active, Active-Passive, Warm Standby, Pilot Light

Lab: Restore from Backup and Run Chaos Testing Scenarios

08Modern SRE Operations & Production Scenarios6 уроков+

End-to-End Troubleshooting Methodology

Common Production Failures and Case Studies

Cost vs Reliability Trade-offs

AI for SRE: AIOps, AI-assisted Troubleshooting and RCA

SRE Career Roadmap and Continuous Improvement

Capstone Lab: Full Production Incident Simulation

БЛИЖАЙШИЕ ГРУППЫ

Выберите группу, которую вы действительно сможете посещать.

Показаны только текущие открытые группы.

НАЧАЛО

Будет объявлено

Онлайн
Расписание
Будет объявлено
Формат
Онлайн
Длительность
8 недели · 48 часов
Язык
Уточните у консультанта
Записаться в лист ожидания
КОНТЕКСТНОЕ ПОДТВЕРЖДЕНИЕ
“Программа помогла мне связать отдельные навыки с тем, как реальные команды проектируют, создают и поставляют ПО.”

Смотрите реальные истории выпускников из сообщества Ingress.

Изучить результаты выпускников
ЗАЯВКА ЧЕРЕЗ INGRESS PORTAL

Выбранный курс сохраняется, пока вы создаёте аккаунт.

Мы используем один аккаунт Portal для заявок, оценок и будущего прогресса обучения. Вам не нужно отправлять данные по почте или снова выбирать курс.

Сначала вопросы? Поговорите с консультантом
  1. 01

    Создайте аккаунт Portal или войдитеВаши контактные данные привязаны к одному профилю студента.

  2. 02

    Подтвердите данные заявкиSite Reliability Engineering (SRE) Bootcamp уже выбран.

  3. 03

    Отправьте заявкуПриёмная команда получает её сразу и может связаться с вами через Portal.

ВЫБРАННЫЙ КУРСSite Reliability Engineering (SRE) BootcampОнлайн

Продолжить в Ingress Portal Уже зарегистрированы? Portal позволит вам войти.