SRE Engineer defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems. It guides the user through assessing reliability, defining quantitative SLOs, implementing golden signal monitoring, automating repetitive toil, and testing system resilience via chaos engineering. Reach for this skill when establishing production reliability frameworks, managing error budgets, designing alerts, or reducing operational toil.
Key Features
Quantitative SLO and SLI definitions
Multiwindow burn rate alerting rules
Golden signal Prometheus queries
Toil automation script patterns
Incident response and chaos testing guidance
Privacy & Security
Data Collection
This tool follows industry-standard security practices and only collects data necessary for functionality.