← Back to jobs

NVIDIA Logo
Compute Cluster SRE Engineer

NVIDIA

 

Bengaluru, Karnataka, India

Posted On: 6 days ago
Experience: 3+ years
Availability: Hybrid
Openings: 1
Category: Technical Lead Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

Maintain large-scale GPU compute clusters to ensure high efficiency and availability.

This role is on-site.

Responsibilities

  • Design and implement large-scale infrastructure, including monitoring, logging, and alerting systems.
  • Manage the full service lifecycle from design and deployment to operation and refinement.
  • Measure and monitor availability, latency, and system health to maintain live infrastructure.
  • Scale systems through automation and evolve architecture to improve reliability and velocity.
  • Respond to incidents and conduct blameless postmortems to drive iterative improvements.

Required Skills

  • 3+ years of hands-on industry experience in systems engineering.
  • Linux system administration expertise with Ubuntu and CentOS/RedHat.
  • HPC cluster scheduler experience, specifically with Slurm or LSF.
  • Proficiency in scripting with Python, Perl, or Bash.
  • Experience with open-source IT automation tools such as Ansible.
  • Bachelor’s Degree in Computer Science or a related technical field.

Preferred Skills

  • Experience with Bright Cluster Manager (BCM).
  • Knowledge of Infiniband or Ethernet concepts.
  • Experience with high-speed storage solutions like Lustre or GPFS.

Education

Bachelor’s Degree

Related Jobs

No related jobs found

← Back to jobs