← Back to jobs

NVIDIA Logo
Senior AI Infrastructure Engineer

NVIDIA

 

Santa Clara, CA, USA

Posted On: Just posted
Experience: 6+ years
Availability: Hybrid
Openings: 1
Category: AI Infrastructure Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will design, build, and operate internal tooling for large-scale AI training and inference platforms on cloud infrastructure.

This role is hybrid.

Responsibilities

  • Conduct performance characterization and analysis on large multi-GPU and multi-node clusters.
  • Improve service lifecycles from inception through deployment, operation, and refinement.
  • Support pre-launch phases via system design consulting, tool development, and capacity management.
  • Monitor availability, latency, and overall system health for live services.
  • Scale systems sustainably using automation and drive reliability improvements.

Required Skills

  • 6+ years of professional experience in infrastructure automation and distributed systems design.
  • Developed tools for large-scale private or public cloud systems in production.
  • Proficiency in Python, Go, C/C++, or Java.
  • In-depth knowledge of Linux, Networking, Storage, and Containers.
  • Experience with Public Cloud, Infrastructure as Code (IaC), and Terraform.
  • Ability to balance independent project initiation with effective collaboration.

Preferred Skills

  • Experience with large-scale AI/ML infrastructure platforms.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs