← Back to jobs
Sterling, VA, USA
No related jobs found
Roles and Responsibilities
- Design and implement a P0–P3 severity taxonomy with ISSUE/CHANGE/WARNING categories, replacing the current alerting model
- Build a multi-signal correlation engine that combines 4 data sources (APM, infrastructure monitoring, network health, cluster status) with cross-layer confidence scoring
- Develop AWS Lambda polling functions for network device APIs and Kubernetes cluster APIs at configurable intervals
- Create 10+ observability monitors (pod crash loops, OOM kills, GPU anomalies, NVMe health, 5xx rates) in a Datadog-like monitoring platform
- Implement PagerDuty escalation with L1/L2/L3 tiers and automatic incident creation for critical alerts
- Build persona-based routing (Ops, Engineering, Partners, End Customers) with confidence-gated customer notifications
- Implement incident lifecycle management with append-only timelines, MTTA/MTTR/MTTD tracking, and post-mortem framework
- Design and execute 10 end-to-end test scenarios validating the full alert lifecycle
- Deliver architecture diagrams, operational runbooks, and monitor specifications
Required Technical Skills
- Backend: Python, REST APIs, AWS Lambda, DynamoDB, schema design
- Observability: Datadog, Grafana, Prometheus, or similar (alert rules, dashboards, webhook integrations, SLI/SLO concepts)
- Incident Management: PagerDuty or Opsgenie (escalation policies, Events API, on-call rotation)
- Cloud: AWS (Lambda, DynamoDB, EventBridge, SES, API Gateway), IAM
- Kubernetes: OpenShift or EKS (cluster health, operator status, pod lifecycle)
- Networking: REST API polling, webhook handlers, network device monitoring (Peplink, Meraki, or similar)
- Testing: pytest, end-to-end test design, API testing
- Communication: Slack API/webhooks, email (SES), ticketing systems (Freshworks, Jira, ServiceNow)
Experience
- 7+ years engineering, 3+ years in observability/SRE/platform
- Built or operated a multi-source alerting or correlation system
- Experience with severity models, escalation workflows, and incident management
Any Graduate
No related jobs found
← Back to jobs