← Back to jobs

NVIDIA Logo
Senior Site Reliability Engineer

NVIDIA

 

Santa Clara, CA, USA

Posted On: 1 day ago
Experience: 8+ years
Availability: Hybrid
Openings: 1
Category: Senior Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You design and implement on-premises HPC infrastructure supplemented with cloud computing to meet growing IT needs.

This role is on-site.

Responsibilities

  • Design and implement advanced storage solutions, including high-performance NFS, S3-compatible object storage, and distributed storage systems.
  • Develop tooling to automate deployment, manage large-scale infrastructure environments, and enable self-service resource consumption.
  • Automate operational monitoring and alerting processes.
  • Collaborate with teams to understand developer workflows and gather infrastructure requirements.
  • Guide methodologies for building, testing, and deploying applications for optimal performance.

Required Skills

  • 8+ years of relevant experience with a BS in Computer Science or equivalent.
  • Deep experience with storage protocols such as NFS, NVMe/TCP, S3, and Lustre (LNet).
  • Experience with containerization technologies like Kubernetes and their integration with storage.
  • Proficiency in Python or Go.
  • Experience with configuration management tools like Chef, Ansible, Puppet, or Saltstack.
  • Background with cloud infrastructure (AWS, Azure, or Google Cloud).
  • Experience with monitoring stacks such as Prometheus+Grafana or Elasticsearch+Kibana.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs