Description
Job Description
Stefanini Group is seeking a Senior Compute Engineer specialized in Red Hat OpenShift to strengthen our Compute Operations team and provide Level 3 (L3) expert support for enterprise customers running critical workloads on container and virtualization platforms.
This role is a key technical position focused on day-to-day operations, stability, and continuous improvement of OpenShift-based platforms. The engineer will act as the highest escalation point for complex incidents and problems, support platform lifecycle activities (upgrades, patching, performance tuning), and contribute to platform modernization initiatives - including VMware-to-OpenShift virtualization transformation programs.
The ideal candidate combines strong troubleshooting skills, deep infrastructure understanding, and hands-on OpenShift expertise, with the ability to work in a structured operational environment (ITIL/ managed services), while also supporting automation and standardization.
Job Responsibilities:
Level 3 Operations & Technical Escalation (Core Responsibility):
- Act as the L3 escalation point for complex technical issues related to:
- Red Hat OpenShift clusters (control plane, worker nodes, networking, storage, authentication)
- OpenShift Virtualization (KubeVirt) and VM-based workloads hosted on OpenShift
- Linux OS level issues impacting cluster stability or workloads.
- Own and drive resolution of:
- Major Incidents (P1/P2) with deep technical investigation and rapid recovery focus
- Recurring incidents through Problem Management (root cause analysis and permanent fixes).
- Lead deep troubleshooting activities:
- cluster degradation, node failures, API instability, etcd performance issues
- networking issues (ingress, routes, DNS, CNI, service connectivity)
- storage issues (persistent volumes, performance bottlenecks, CSI failures)
- workload failures (pods, operators, deployments, stateful applications).
- Provide clear technical updates during incidents, including impact assessment, recovery plan/ workaround, risks and next steps.
Platform Lifecycle Management (Upgrades, Patching, Stability):
- Plan and execute OpenShift lifecycle activities such as version upgrades (cluster upgrades and operator upgrades), patching and security hardening and certificate management and renewal processes.
- Validate platform readiness before changes: capacity, compatibility, performance, known issues.
- Maintain high availability and resilience: backup/restore strategy support (including etcd backup practices), disaster recovery readiness and operational runbooks.
- Ensure operational compliance with defined maintenance windows and change governance.
VMware-to-OpenShift Virtualization Transformation Support (While the role is operational, this is a strategic area of expertise required in the team):
- Support enterprise modernization initiatives involving migration from traditional virtualization platforms (VMware) to OpenShift Virtualization
- Contribute to:
- migration approach definition and technical design support
- workload onboarding, validation, and stabilization on OpenShift
- performance tuning and operational model definition for VM-based workloads on OpenShift.
- Ensure production-grade operational readiness: monitoring, alerting, backup, patching and support model aligned with managed services standards.
Standardization, Automation & Operational Improvement:
- Develop and maintain operational documentation, including troubleshooting guides, standard operating procedures (SOPs), build standards and reference architectures, operational runbooks for recurring tasks.
- Support automation initiatives using tools such as: Ansible / Automation Platform (preferred), GitOps practices (ArgoCD) where applicable and scripting (Bash / Python) to reduce manual operations.
- Proactively identify improvements to increase platform stability, recovery speed (MTT), repeatability and reduction of human error.
Monitoring, Observability & Performance Management:
- Support and improve observability across the platform, including:
- OpenShift monitoring stack (Prometheus / Alertmanager / Grafana)
- log management (e.g., EFK / Loki or enterprise logging platforms).
- Troubleshoot performance issues related to compute resource constraints, scheduling and resource requests/limits and cluster scaling and capacity planning.
- Work with customer stakeholders and internal teams to define alert thresholds, reduce noise and false positives and improve operational dashboards and health reporting.
Security & Compliance Support:
- Ensure the platform is operated in a secure manner aligned with enterprise expectations:
- RBAC best practices
- integration with enterprise identity providers (LDAP / AD / SSO)
- secure cluster configuration and segregation
- Support vulnerability remediation and platform hardening initiatives
- Collaborate with Security teams for audits, compliance requests, and evidence collection.
Job Requirements
Mandatory Technical Skills:
- Strong hands-on experience with Red Hat OpenShift administration and operations.
- Strong Linux background (RHEL preferred), including troubleshooting OS performance, services, networking, and storage.
- Solid understanding of Kubernetes fundamentals: pods, deployments, services, ingress, namespaces, RBAC, operators.
- Experience troubleshooting infrastructure-related issues across compute, network, storage, and platform services.
- Experience working in production environments with uptime and SLA commitments.
Mandatory Professional Skills:
- Proven ability to operate as Level 3 support, including deep troubleshooting, structured root cause analysis and ownership until resolution.
- Ability to communicate clearly with customers (technical and non-technical stakeholders) and internal teams (L1/L2/architects/project teams).
- Strong documentation discipline and operational mindset.
Preferred/ Nice-to-Have Skills:
- Experience with OpenShift Virtualization (KubeVirt) and VM-based workloads.
- Experience supporting VMware environments and understanding virtualization concepts: vSphere architecture, clusters, HA/DRS, storage/datastores, VM lifecycle.
- Experience with automation tools:
- Ansible / Red Hat Ansible Automation Platform
- GitOps tools (ArgoCD)
- Infrastructure as Code practices.
- Experience with enterprise storage and CSI integrations.
- Experience with enterprise networking topics (DNS, routing, firewall constraints, load balancing).
- Experience with public cloud OpenShift deployments (optional): ROSA / ARO / OCP on AWS/Azure/GCP.
Certifications (Preferred):
- Red Hat Certified Specialist in OpenShift Administration (preferred)
- Red Hat Certified Engineer (RHCE) (strong advantage)
- Kubernetes certifications (CKA/CKAD) (nice to have)