← Back to jobs
Long Lake, WA, USA
No related jobs found
• Experience supporting large-scale distributed systems (200+ microservices).
• Strong hands-on experience with FluxCD for GitOps-based continuous delivery.
• Experience implementing and managing:
o GitOps deployment patterns
o Automated Kubernetes application reconciliation
o Multi-environment deployment strategies
o Progressive delivery and automated rollbacks
o Configuration and secret management
o Helm-based application deployments through FluxCD
• Expertise integrating FluxCD with Enterprise GitHub, Jenkins, Azure Container Registry (ACR), and Kubernetes clusters.
• Experience managing platform releases for 200+ microservices using GitOps principles.
• Ability to troubleshoot deployment drift, reconciliation failures, and cluster synchronization issues.
• Strong understanding of:
o SLI/SLO/SLA management
o Incident Management
o Problem Management
o Capacity Planning
o Reliability Engineering
o Root Cause Analysis (RCA)
o Chaos Engineering principles
________________________________________
Candidate Profile
• We are seeking a highly skilled Senior Platform Site Reliability Engineer (SRE) to support and operate a large-scale distributed platform consisting of 200+ Java-based microservices, event-driven architectures, modern UI applications, and enterprise-grade databases.
• The candidate will drive platform reliability, scalability, observability, automation, and operational excellence across cloud-native environments.
• The ideal engineer will have strong expertise in Kubernetes, FluxCD, Terraform, Java ecosystem, Kafka/ActiveMQ messaging, CI/CD automation, observability, and database operations, with a mindset focused on reliability engineering, automation, and production support
Bachelor's degree
No related jobs found
← Back to jobs