Specialize in ensuring the reliability, resilience, and recoverability of enterprise services and systems across cloud, on-prem, and store environments.
Serve as the Site Reliability Engineering (SRE) voice within incident and problem management, leading major incident response.
Drive root cause analysis and partner with Software Engineers and platform teams to reduce recurrence and improve system health in production.
Utilize observability and monitoring tools such as Dynatrace and Azure Monitor to detect, triage, and diagnose production issues.
Collaborate in an Agile environment, tracking work and metrics via platforms such as Jira, and support new go-lives and pilots onsite in Blue Ash, OH.
What's Needed?
3+ years of experience in Site Reliability Engineering, incident management, or production support for enterprise systems.
Experience leading or participating in major incident management, including incident commander roles and executive communication during outages.
Solid understanding of observability and monitoring concepts and tooling (e.g., Dynatrace, Azure Monitor).
Working knowledge of Linux and scripting languages such as BASH and Python for troubleshooting and automation.
Familiarity with Kubernetes, Docker, and cloud platforms like Azure and GCP to troubleshoot and reason about system behavior in production.