Description
Manage Azure-based production environments, implement SRE practices, and automate infrastructure and CI/CD pipelines to ensure platform reliability and performance.
This role is on-site.
Responsibilities
- Manage and support Azure production environments, ensuring stability, performance, and availability.
- Implement SRE practices including SLIs, SLOs, error budgets, alert tuning, and toil reduction.
- Develop automation scripts and internal tooling using Python for reporting and operational workflows.
- Deploy, manage, and troubleshoot Azure services (App Services, Functions, AKS, Storage, Key Vault, etc.).
- Troubleshoot incidents across application, infrastructure, and network layers, performing root cause analysis.
Required Skills
- Strong hands-on experience with Microsoft Azure cloud services and networking (VNets, NSGs, Private Endpoints, Load Balancers, App Gateway, Azure Front Door).
- Proven experience in SRE, DevOps, or platform engineering roles with 5+ years of experience.
- Strong Python scripting experience for automation, integrations, and tooling.
- Experience with Azure Monitor, Log Analytics, Application Insights, KQL, dashboards, and alerts.
- Strong knowledge of Terraform, Bicep, ARM templates, or similar Infrastructure as Code tools.
- Hands-on experience building and maintaining CI/CD pipelines using Azure DevOps, GitHub Actions, or Jenkins.
- Good understanding of Linux and Windows environments.
- Strong expertise in incident management, troubleshooting, RCA, and operational excellence.
Preferred Skills
- Experience with Kafka, Service Bus, Event Hub, or high-volume messaging systems.
- Knowledge of advanced observability tools (Prometheus, Grafana, Datadog).
- Experience collaborating with globally distributed teams or in regulated environments.