Design, build, and maintain Azure landing zones and platform services such as VNet, Private Endpoints, Key Vault, Azure Firewall/NSGs, and Application Gateway/WAF.
Implement Infrastructure as Code (IaC) with Terraform and/or Bicep, enforce GitOps workflows, and create reusable modules and pipelines to support automation-first approaches.
Define and measure SLIs/SLOs, error budgets, and reliability roadmaps for critical cloud services, ensuring high availability and performance.
Develop and optimize observability solutions using Azure Monitor, Log Analytics, Application Insights, and Prometheus/Grafana to ensure system health and performance.
Support incident response, perform root cause analysis, and lead post-incident reviews to continuously improve platform reliability and security.
What's Needed?
At least 5 years of hands-on experience with Azure infrastructure and services in a production environment.
Proficiency in Azure services such as AKS, App Services, Functions, Azure SQL/MI, Cosmos DB, Storage, Event Hub/Service Bus, Redis, and networking components.
Strong expertise in Infrastructure as Code (Terraform preferred) and Git-based workflows using GitHub or Azure DevOps.
Experience with SRE principles, including SLIs/SLOs, incident management, and capacity planning.
Proficiency in scripting languages such as PowerShell and Python, along with Linux fundamentals and security best practices