← Back to jobs

InfiCare Technologies Logo
GPU Solution Architect

InfiCare Technologies

 

Santa Clara, CA, USA

Posted On: 30+ days ago
Experience: 8+ years
Availability: Remote
Openings: 1
Category: Solution Architect
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

Key Responsibilities


 

  • Diagnose and resolve hard Day 2 operations problems at scale alongside partner engineering teams
  • Help partners prepare operating models for new platforms, capacity, and services before customer adoption
  • Improve reliability, performance, and cost efficiency using metrics such as incident frequency, recovery time, utilization, and cost per token
  • Assess and improve partner Day 2 maturity across people, process, tooling, telemetry, security, and incident response
  • Convert validated solutions into operating procedures, reference architectures, automation scripts, and agentic workflows reusable across the partner ecosystem

Required Skills


 

  • 8+ years in production infrastructure, cloud engineering, solutions architecture, SRE, or HPC; or 5+ years of specialist-level work in large-scale GPU or AI infrastructure
  • Hands-on experience building, operating, or improving distributed infrastructure under real production load
  • Deep expertise in at least one Day 2 stack area: GPU monitoring (DCGM), BMC/Redfish, firmware/driver lifecycle; InfiniBand, NCCL, UFM; or high-performance storage (Lustre, IBM Storage Scale, WEKA, VAST Data)
  • Proficiency across Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry
  • Automation experience with Terraform, Ansible, Argo CD, or similar tooling
  • Strong Linux knowledge; Python, Bash, or equivalent scripting for measurement, diagnosis, and remediation
  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or equivalent experience

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs