Drive the execution of our internal cloud migration, proactively communicating across organizations to clear roadblocks, align on technical dependencies, and move workloads efficiently.
Leverage existing Infrastructure-as-Code (Terraform) and automation frameworks, actively identifying failure points and improving them as we migrate.
Assist heavily in debugging existing deployment pipelines and infrastructure tools, writing code (C#) and scripting (PowerShell) to resolve issues and increase reliability.
Adopt an unwavering "own it, make it better" mentality; take responsibility for the stability and efficiency of the tools and processes you interact with.
Take a data-driven approach to your decisions using telemetry and metrics, working effectively in a dynamic environment where you need to think on your feet.
Establish strict, objective success metrics for environment parity and migration speed, ensuring we have the data to prove our strategic wins and honor our timeline commitments.
What You Bring
Senior-level engineering experience with a proven track record of managing and rapidly deploying services in cloud environments (specifically Azure).
Deep expertise in Cloud Infrastructure-as-a-Service (IaaS) and migrating/provisioning environments with a focus on efficiency and speed.
Exceptional communication skills, with a proven ability to articulate complex technical challenges across different organizational boundaries to resolve blockers.
Highly motivated self-starter who thrives as part of a broader team but requires minimal direction to dive into existing systems, figure things out, and drive results.
Strong proficiency in utilizing and refactoring existing Terraform code for Infrastructure-as-Code (IaC) implementation.
Highly adept at PowerShell scripting and debugging to improve system automation and configuration.
Strong working knowledge of C# to effectively debug and enhance existing object-oriented infrastructure tooling.
Proven track record of using telemetry and metrics to ground decisions in objective data.
Core Responsibilities
Lifecycle Orchestration: Automate the end-to-end lifecycle of StatefulSets: provisioning, seamless volume expansion, graceful termination, and automated re-attachment during node failures.
High Availability & Uptime: Implement advanced scheduling logic (Pod Topology Spread Constraints, Anti-affinity) to ensure stateful workloads survive zonal outages and maintenance windows.
Storage Performance & Tuning: Optimize Azure Disk (Premium/Ultra) and Azure NetApp Files integration via CSI drivers to minimize IOPS bottlenecks and latency.
Disaster Recovery Automation: Develop and test automated "Snapshot-to-Restore" pipelines. Ensure that the Actual State of data volumes can be recovered to the Goal State in minutes, not hours.
Infrastructure as Code: Utilize Terraform to provision the hardened Azure foundation (Disk Encryption Sets, Proximity Placement Groups, and Networking) required for high-performance stateful clusters.
Observability & Health: Build deep-visibility dashboards and alerting for Persistent Volume (PV) utilization, disk pressure, and stateful replication lags.
Technical Qualifications
Kubernetes Internal Mastery: Expert-level understanding of StatefulSet controllers, Persistent Volume Claims (PVCs), and the Container Storage Interface (CSI).
Azure AKS Specialist: Deep experience with Azure Kubernetes Service, specifically around persistent storage integration and Azure-specific networking constraints.
Automation & Scripting: Proficient in Go or Python/Bash for writing custom controllers or maintenance hooks (PreStop/PostStart) that ensure data consistency during updates.
Reliability Engineering: Proven track record of managing production databases or distributed systems (e.g., Postgres, ClickHouse, Elasticsearch) on Kubernetes.
GitOps & State Management: Experience using GitOps to manage complex stateful deployments where order of operations matters