3+ years of experience in Site Reliability Engineering, incident management, or production support for enterprise systems
Experience leading or participating in major incident management (incident commander, bridge/war-room facilitation, executive communication during outages)
Experience with problem management and root cause analysis (RCA), including driving corrective/preventive actions to closure
Solid understanding of observability and monitoring concepts and tooling (e.g., Dynatrace, Azure Monitor) to detect, triage, and diagnose production issues
Working knowledge of Linux and scripting (e.g., BASH, Python) for troubleshooting and diagnostic automation
Familiarity with Kubernetes, Docker, and cloud platforms (Azure, GCP) sufficient to troubleshoot and reason about system behavior in production
Experience working in an Agile environment, tracking work and metrics via a platform such as Jira
Strong communication skills — able to translate technical incident details for both engineering teams and business stakeholders under time pressure
In office, Blue Ash, OH 5 days a week
Willingness to travel and provide onsite support for new go-lives and pilots
Familiarity with Point of Sale Systems in Enterprise environments