Skip to content

IBM Cloud Site Reliability Engineer Certification - Fact Sheet

Quick Reference

  • Certification: IBM Certified Site Reliability Engineer (SRE) - Cloud v2
  • Exam Code: C1000-174
  • Duration: 90 minutes
  • Questions: 60 multiple choice
  • Passing Score: 70%
  • Cost: $200 USD
  • Language: English
  • Delivery: Pearson VUE (online or test center)
  • Prerequisites: None (recommended 2+ years SRE/DevOps experience)
  • Recertification: Every 3 years
  • Target Audience: Site reliability engineers, DevOps engineers, operations engineers, platform engineers

Official Resources

Exam Domains

1. Monitoring and Observability (25%)

  • Monitoring strategy and implementation
  • Metrics collection and analysis
  • Log aggregation and analysis
  • Distributed tracing
  • Service Level Indicators (SLIs)
  • Service Level Objectives (SLOs)
  • Service Level Agreements (SLAs)
  • Error budgets
  • Alerting and notification strategies
  • Dashboard design and visualization
  • Synthetic monitoring
  • Real User Monitoring (RUM)

Monitoring Resources: - πŸ“– IBM Cloud Monitoring - πŸ“– Monitoring with Sysdig - πŸ“– Log Analysis - πŸ“– Log Analysis with LogDNA - πŸ“– Activity Tracker - πŸ“– Flow Logs for VPC - πŸ“– Prometheus on Kubernetes - πŸ“– SLI and SLO Guide - πŸ“– Error Budget Implementation - πŸ“– Alerting Best Practices

2. Incident Management and Response (20%)

  • Incident detection and triage
  • Incident response procedures
  • On-call management
  • Incident communication
  • Post-incident reviews (PIRs)
  • Root cause analysis (RCA)
  • Blameless postmortems
  • Incident tracking and metrics
  • Escalation procedures
  • Disaster recovery procedures
  • Chaos engineering principles

Incident Management Resources: - πŸ“– Incident Management Best Practices - πŸ“– PagerDuty Integration - πŸ“– Slack Integration - πŸ“– Event Notifications - πŸ“– Activity Tracker for Auditing - πŸ“– Postmortem Culture - πŸ“– Managing Incidents - πŸ“– Chaos Engineering

3. Automation and Infrastructure as Code (20%)

  • Infrastructure as Code (IaC) principles
  • Terraform on IBM Cloud
  • IBM Cloud Schematics
  • Configuration management
  • Automation scripts and tools
  • CI/CD pipeline design
  • GitOps practices
  • Policy as Code
  • Automated testing
  • Self-healing systems
  • Runbook automation

Automation Resources: - πŸ“– Terraform on IBM Cloud - πŸ“– Terraform Provider - πŸ“– IBM Cloud Schematics - πŸ“– Schematics Actions - πŸ“– Continuous Delivery - πŸ“– Toolchains - πŸ“– Tekton Pipelines - πŸ“– Ansible on IBM Cloud - πŸ“– IBM Cloud CLI - πŸ“– GitOps with ArgoCD - πŸ“– Runbook Automation

4. High Availability and Reliability (15%)

  • High availability design patterns
  • Fault tolerance and resilience
  • Disaster recovery planning
  • Backup and restore strategies
  • Multi-zone and multi-region architectures
  • Load balancing and failover
  • Circuit breakers and bulkheads
  • Retry logic and exponential backoff
  • Graceful degradation
  • Capacity planning
  • Performance optimization

HA and Reliability Resources: - πŸ“– High Availability Architecture - πŸ“– Disaster Recovery - πŸ“– Multi-Zone Regions - πŸ“– Load Balancers - πŸ“– Auto-Scaling - πŸ“– Kubernetes HA - πŸ“– Database HA - πŸ“– Backup Services - πŸ“– Reliability Patterns

5. Performance and Capacity Management (10%)

  • Performance monitoring and tuning
  • Capacity planning methodologies
  • Resource optimization
  • Load testing strategies
  • Performance profiling
  • Database performance tuning
  • Network performance optimization
  • Cost vs performance trade-offs
  • Scalability patterns
  • Bottleneck identification

Performance Resources: - πŸ“– Performance Optimization - πŸ“– Monitoring Metrics - πŸ“– Kubernetes Performance - πŸ“– Database Performance - πŸ“– Load Testing - πŸ“– CDN Performance - πŸ“– Capacity Planning

6. Container and Kubernetes Operations (10%)

  • Kubernetes cluster management
  • Container orchestration
  • Pod scheduling and placement
  • Resource quotas and limits
  • Horizontal and vertical pod autoscaling
  • Cluster autoscaling
  • Rolling updates and deployments
  • StatefulSets and DaemonSets
  • Kubernetes monitoring and logging
  • Helm chart management
  • Service mesh operations (Istio)
  • Container security operations

Kubernetes Resources: - πŸ“– IBM Cloud Kubernetes Service - πŸ“– Kubernetes Best Practices - πŸ“– Cluster Autoscaler - πŸ“– Horizontal Pod Autoscaler - πŸ“– Vertical Pod Autoscaler - πŸ“– OpenShift Operations - πŸ“– Helm on Kubernetes - πŸ“– Istio Service Mesh - πŸ“– Kubernetes Monitoring

Core SRE Concepts

The Four Golden Signals

  1. Latency: Time to service a request
  2. Traffic: Demand on your system
  3. Errors: Rate of failed requests
  4. Saturation: How "full" your service is

Resources: - πŸ“– Four Golden Signals

SLIs, SLOs, and SLAs

Service Level Indicators (SLIs): - Quantitative measures of service level - Examples: latency, error rate, throughput, availability

Service Level Objectives (SLOs): - Target values for SLIs - Example: 99.9% availability, 95th percentile latency < 200ms

Service Level Agreements (SLAs): - Contractual agreements with consequences - Usually less stringent than SLOs

Error Budgets: - 100% - SLO = Error Budget - Used to balance innovation and reliability

Resources: - πŸ“– SLO Implementation - πŸ“– SLO Monitoring

Toil and Toil Reduction

Toil: Manual, repetitive, automatable work with no enduring value

Characteristics: - Manual - Repetitive - Automatable - Tactical - No enduring value - Scales linearly with service growth

Toil Reduction Strategies: - Automation - Self-service tools - Process improvement - Platform engineering

Resources: - πŸ“– Eliminating Toil

On-Call Best Practices

  • Reasonable on-call load (50% for tickets, <2% for pages)
  • Clear escalation paths
  • Comprehensive runbooks
  • Post-incident reviews
  • On-call compensation
  • Mental health support

Resources: - πŸ“– Being On-Call

IBM Cloud SRE Tools

Monitoring and Observability

Automation and IaC

CI/CD and DevOps

Kubernetes and Containers

Common SRE Scenarios

Scenario 1: Implementing SLOs and Error Budgets

Challenge: Define and implement SLOs for a microservices application.

Implementation: 1. Identify SLIs: - Availability: % of successful requests - Latency: 95th percentile response time - Throughput: requests per second - Error rate: % of 5xx errors

  1. Set SLOs:
  2. Availability: 99.9% (3 nines)
  3. Latency: 95th percentile < 200ms
  4. Error rate: < 0.1%

  5. Calculate Error Budget:

  6. 100% - 99.9% = 0.1% error budget
  7. For 30 days: 43.2 minutes of downtime

  8. Monitor SLOs:

  9. Set up monitoring dashboards
  10. Configure alerts at 50%, 75%, 90% error budget consumption
  11. Track error budget burn rate

  12. Use Error Budget:

  13. Deploy new features when error budget available
  14. Freeze deployments when error budget exhausted
  15. Focus on reliability improvements

Resources: - πŸ“– SLO Implementation Guide - πŸ“– Monitoring SLOs

Scenario 2: Building Automated Incident Response

Challenge: Automate detection and initial response to common incidents.

Implementation: 1. Detection: - Configure monitoring alerts in Sysdig - Set up log-based alerts in LogDNA - Implement synthetic monitoring - Create custom metrics

  1. Notification:
  2. Integrate with PagerDuty for on-call
  3. Send alerts to Slack channels
  4. Use Event Notifications service
  5. Implement escalation policies

  6. Automated Response:

  7. Create Schematics actions for remediation
  8. Use Cloud Functions for automated fixes
  9. Implement runbook automation
  10. Auto-scaling based on metrics

  11. Documentation:

  12. Maintain up-to-date runbooks
  13. Document incident response procedures
  14. Create troubleshooting guides
  15. Track incident history

Resources: - πŸ“– Incident Response Automation - πŸ“– Schematics Actions

Scenario 3: Implementing GitOps for Kubernetes

Challenge: Deploy and manage Kubernetes applications using GitOps.

Implementation: 1. GitOps Setup: - Store all configurations in Git - Use separate repos for app and config - Implement branch protection - Use pull request workflow

  1. Deployment Automation:
  2. Deploy ArgoCD or Flux
  3. Configure automatic sync
  4. Implement health checks
  5. Set up rollback policies

  6. Environment Management:

  7. Use Kustomize for overlays
  8. Separate dev/staging/prod configs
  9. Implement promotion workflow
  10. Use Helm for package management

  11. Monitoring:

  12. Monitor sync status
  13. Track deployment metrics
  14. Alert on sync failures
  15. Audit all changes via Git

Resources: - πŸ“– GitOps with Kubernetes - πŸ“– ArgoCD Documentation

Scenario 4: Capacity Planning and Auto-Scaling

Challenge: Implement comprehensive auto-scaling strategy.

Implementation: 1. Resource Monitoring: - Track CPU, memory, network, disk - Monitor application-specific metrics - Collect historical data - Identify usage patterns

  1. Horizontal Scaling:
  2. Configure Horizontal Pod Autoscaler (HPA)
  3. Use custom metrics for scaling
  4. Set min/max replicas
  5. Configure scale-up/down policies

  6. Vertical Scaling:

  7. Implement Vertical Pod Autoscaler (VPA)
  8. Set resource requests and limits
  9. Monitor right-sizing recommendations
  10. Plan maintenance windows for VPA changes

  11. Cluster Scaling:

  12. Enable cluster autoscaler
  13. Configure node pools
  14. Set cluster size limits
  15. Monitor node utilization

  16. Load Testing:

  17. Perform capacity testing
  18. Identify breaking points
  19. Validate scaling policies
  20. Test disaster recovery scenarios

Resources: - πŸ“– Kubernetes Autoscaling - πŸ“– Cluster Autoscaler

Scenario 5: Disaster Recovery Testing

Challenge: Implement and test disaster recovery procedures.

Implementation: 1. DR Strategy: - Define RTO and RPO requirements - Choose DR approach (backup/restore, pilot light, warm standby, hot standby) - Document DR procedures - Assign roles and responsibilities

  1. Backup Implementation:
  2. Enable automated backups for databases
  3. Backup Kubernetes resources
  4. Store backups in separate region
  5. Test backup encryption
  6. Implement backup retention policies

  7. Failover Procedures:

  8. Document failover steps
  9. Automate failover where possible
  10. Configure DNS for failover
  11. Implement health checks
  12. Test failover mechanisms

  13. DR Testing:

  14. Schedule regular DR tests
  15. Test backup restoration
  16. Validate failover procedures
  17. Measure actual RTO and RPO
  18. Update procedures based on findings

  19. Documentation:

  20. Maintain DR runbooks
  21. Document test results
  22. Track improvements
  23. Train team members

Resources: - πŸ“– Disaster Recovery Planning - πŸ“– Backup Strategies

SRE Best Practices

Monitoring Best Practices

  1. Monitor the Four Golden Signals: Latency, traffic, errors, saturation
  2. Use Multiple Data Sources: Metrics, logs, traces
  3. Create Actionable Alerts: Every alert should require action
  4. Avoid Alert Fatigue: Tune thresholds, use alert suppression
  5. Dashboard Design: Show what matters, avoid vanity metrics
  6. Synthetic Monitoring: Proactive monitoring of user journeys
  7. Distributed Tracing: Understand request flows across services

Resources: - πŸ“– Monitoring Best Practices

Incident Management Best Practices

  1. Clear Escalation: Well-defined escalation procedures
  2. Blameless Culture: Focus on systems, not people
  3. Post-Incident Reviews: Learn from every incident
  4. Action Items: Track and complete follow-up tasks
  5. Communication: Keep stakeholders informed
  6. Documentation: Maintain incident history
  7. Continuous Improvement: Use incidents to improve systems

Resources: - πŸ“– Postmortem Culture

Automation Best Practices

  1. Automate Toil: Eliminate repetitive manual work
  2. Infrastructure as Code: All infrastructure in version control
  3. Idempotency: Operations safe to repeat
  4. Testing: Test automation thoroughly
  5. Documentation: Document automation behavior
  6. Monitoring: Monitor automation execution
  7. Gradual Rollout: Test automation on non-production first

Resources: - πŸ“– Automation Philosophy

Capacity Planning Best Practices

  1. Data-Driven: Base decisions on metrics
  2. Forecasting: Project future capacity needs
  3. Organic Growth: Plan for natural growth
  4. Inorganic Growth: Plan for launches and campaigns
  5. Load Testing: Validate capacity assumptions
  6. Regular Review: Update capacity plans quarterly
  7. Buffer Capacity: Maintain headroom for unexpected growth

Resources: - πŸ“– Capacity Planning

On-Call Best Practices

  1. Reasonable Load: < 50% of time on-call work
  2. Quality Runbooks: Clear, tested procedures
  3. Post-Incident Reviews: Learn and improve
  4. Escalation Paths: Clear when to escalate
  5. Compensation: Fair on-call compensation
  6. Support: Mental health and burnout prevention
  7. Rotation: Fair rotation schedule

Resources: - πŸ“– On-Call Best Practices

Exam Tips and Strategies

General Preparation

  1. Read SRE Books: Google SRE Book and Workbook
  2. Hands-On Practice: Implement monitoring and automation
  3. Real Incidents: Learn from real-world scenarios
  4. Tool Proficiency: Master IBM Cloud monitoring tools
  5. Kubernetes Operations: Deep Kubernetes operational knowledge

Key Study Areas

  • SLIs, SLOs, SLAs, and error budgets
  • Monitoring with Sysdig and LogDNA
  • Incident management procedures
  • Terraform and Infrastructure as Code
  • Kubernetes operations and autoscaling
  • CI/CD with Tekton and Toolchains
  • High availability patterns
  • Disaster recovery planning
  • Performance optimization
  • Capacity planning methodologies
  • Toil reduction strategies
  • Blameless postmortem culture

Common Pitfalls to Avoid

  • Confusing SLIs with SLOs
  • Over-alerting leading to alert fatigue
  • Not understanding error budget concept
  • Missing the importance of blameless culture
  • Not automating repetitive tasks
  • Ignoring capacity planning
  • Poor incident communication
  • Not testing DR procedures
  • Inadequate monitoring coverage
  • Not tracking toil

Scenario-Based Questions

  • Read scenarios carefully - identify key requirements
  • Look for SRE principles (reliability, scalability, observability)
  • Consider automation opportunities
  • Think about monitoring and alerting needs
  • Evaluate high availability requirements
  • Consider incident response procedures
  • Think about capacity and performance implications

Exam Day Tips

  1. Focus on SRE principles, not just tools
  2. Think about reliability and scalability
  3. Consider monitoring and observability
  4. Look for automation opportunities
  5. Evaluate incident response approaches
  6. Consider capacity planning needs
  7. Think about error budgets
  8. Focus on blameless culture
  9. Review flagged questions
  10. Manage time effectively

Additional Learning Resources

Books and Publications

Training and Courses

Community and Support

Tools and Documentation

Important Exam Topics by Priority

High Priority (Study First)

  1. SLIs, SLOs, SLAs, and error budgets
  2. Monitoring with Sysdig (metrics, dashboards, alerts)
  3. Log Analysis with LogDNA
  4. Incident management procedures
  5. Blameless postmortem culture
  6. Kubernetes operations and troubleshooting
  7. Autoscaling (HPA, VPA, cluster autoscaler)
  8. Terraform and Infrastructure as Code
  9. CI/CD with Tekton pipelines
  10. High availability patterns

Medium Priority (Study Second)

  1. Activity Tracker for auditing
  2. Disaster recovery planning
  3. Capacity planning methodologies
  4. Performance optimization techniques
  5. Toil identification and reduction
  6. GitOps practices
  7. Runbook automation
  8. Chaos engineering principles
  9. Service mesh operations (Istio)
  10. Load testing strategies

Lower Priority (If Time Permits)

  1. Advanced Kubernetes features
  2. Advanced Terraform patterns
  3. Specific CI/CD tool integrations
  4. Advanced monitoring queries
  5. Custom metric collection

SRE Checklists

Production Readiness Checklist

  • SLIs and SLOs defined
  • Monitoring and alerting configured
  • Logging aggregation set up
  • Distributed tracing enabled
  • Runbooks documented
  • On-call rotation established
  • Incident response procedures defined
  • Backup and restore tested
  • Disaster recovery plan documented
  • Load testing completed
  • Capacity planning done
  • Auto-scaling configured
  • CI/CD pipeline operational
  • Security scanning integrated
  • Documentation complete

Incident Response Checklist

  • Incident detected and acknowledged
  • Severity level assigned
  • Stakeholders notified
  • Incident commander assigned
  • Communication channel established
  • Troubleshooting underway
  • Mitigation steps taken
  • Service restored
  • Root cause identified
  • Post-incident review scheduled
  • Action items documented
  • Runbooks updated
  • Monitoring improved
  • Prevention measures implemented

Postmortem Checklist

  • Timeline documented
  • Root cause analysis completed
  • Contributing factors identified
  • Impact assessment documented
  • What went well documented
  • What could be improved identified
  • Action items created and assigned
  • Lessons learned shared
  • Monitoring gaps addressed
  • Automation opportunities identified
  • Documentation updated
  • Follow-up meeting scheduled

Final Preparation Checklist

  • Read Google SRE Book key chapters
  • Master SLI/SLO/error budget concepts
  • Practice with IBM Cloud Monitoring
  • Configure Log Analysis for applications
  • Implement Terraform for infrastructure
  • Create CI/CD pipelines with Tekton
  • Practice Kubernetes troubleshooting
  • Configure autoscaling (HPA, VPA, cluster)
  • Set up incident response procedures
  • Write and test runbooks
  • Practice capacity planning
  • Review high availability patterns
  • Study disaster recovery procedures
  • Understand blameless culture
  • Take practice exams
  • Schedule exam appointment

Good luck with your IBM Cloud Site Reliability Engineer certification exam!

Remember: SRE is about balancing reliability with innovation. Focus on data-driven decision making, automation, and continuous improvement. The goal is not 100% uptime, but sustainable, reliable systems that enable rapid innovation.