Diagnose and resolve hard Day 2 operations problems at scale alongside partner engineering teams
Help partners prepare operating models for new platforms, capacity, and services before customer adoption
Improve reliability, performance, and cost efficiency using metrics such as incident frequency, recovery time, utilization, and cost per token
Assess and improve partner Day 2 maturity across people, process, tooling, telemetry, security, and incident response
Convert validated solutions into operating procedures, reference architectures, automation scripts, and agentic workflows reusable across the partner ecosystem
Required Skills
8+ years in production infrastructure, cloud engineering, solutions architecture, SRE, or HPC; or 5+ years of specialist-level work in large-scale GPU or AI infrastructure
Hands-on experience building, operating, or improving distributed infrastructure under real production load
Deep expertise in at least one Day 2 stack area: GPU monitoring (DCGM), BMC/Redfish, firmware/driver lifecycle; InfiniBand, NCCL, UFM; or high-performance storage (Lustre, IBM Storage Scale, WEKA, VAST Data)
Proficiency across Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry
Automation experience with Terraform, Ansible, Argo CD, or similar tooling
Strong Linux knowledge; Python, Bash, or equivalent scripting for measurement, diagnosis, and remediation
BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or equivalent experience