Platform Operations: Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.
Cluster Management: Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
Linux Systems Administration: Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
Infrastructure as Code: Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
Continuous Improvement: Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.
Top requirements:
Bachelors is preferred, but not required.
Minimum of 5 years professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
Candidates must have 5 years strong system admininstration with Linux! Rancher for Kubernetes experience is a must