Description
You will lead Site Reliability Engineering efforts focused on observability and incident management optimization.
Responsibilities
- Configure and enable SLI, SLO, SLA, and Error Budgets.
- Set up alerts based on SLO requirements and optimize alert noise.
- Implement event correlation to improve incident response.
- Monitor application performance using industry-standard tooling.
Required Skills
- 5+ years of experience in SRE or related reliability roles.
- Hands-on experience with BigPanda.
- Proficiency with Splunk.
- Experience with AppDynamics.
- Direct experience with Application Performance Monitoring (APM) tools.
- Strong understanding of alert optimization and event correlation.
- Ability to manage SLI/SLO/SLA frameworks.
- Bachelor's degree or equivalent experience.