Analyze system monitoring data, lead incident triage, and drive reliability improvements across enterprise infrastructure.
This role is remote.
Responsibilities
Lead enterprise-level incident triage, troubleshooting, and root-cause analysis.
Analyze monitoring and observability data to identify performance issues and reliability risks.
Evaluate system designs, workflows, and application architectures to implement reliability enhancements.
Collaborate with DevOps, infrastructure, and application teams to resolve complex technical issues.
Mentor teams in resolving operational challenges and support continuous improvement initiatives.
Required Skills
8+ years of experience deploying, maintaining, and troubleshooting complex enterprise-scale applications.
3+ years of deep expertise in at least two enterprise monitoring/observability tools: Dynatrace, Splunk, SolarWinds, ServiceNow, or Operator Workspace.
Extensive experience in one or more domains: Networking, Windows Systems, Desktop Infrastructure, Unix/Linux, AWS Cloud, Azure Cloud, Middleware, Java/JavaScript Development, or Database Administration.
Proficiency with IT operational metrics, system reliability indicators, and application performance monitoring.
Bachelor's degree in Computer Science, Engineering, IT, or a related technical discipline.
Strong experience working with cross-functional technical teams in large-scale environments.
Proficiency with Microsoft Office applications (Word, Excel, PowerPoint).
Preferred Skills
Experience with test-driven development (TDD), distributed systems, microservices, and cloud-native architectures.
Familiarity with additional monitoring and performance management tools beyond the core stack.
Experience supporting regulated or large-scale government, healthcare, or enterprise environments.