← Back to jobs

NVIDIA Logo
Senior Site Reliability Engineer - DGX Cloud

NVIDIA

 

Santa Clara, CA, USA

Posted On: 1 day ago
Experience: 10+ years
Availability: Remote
Openings: 1
Category: Senior Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the operational and reliability aspects of large-scale Kubernetes clusters, focusing on performance at scale, monitoring, logging, and alerting.

This role is remote.

Responsibilities

  • Design, implement, and support large-scale Kubernetes cluster operations.
  • Improve the full service lifecycle from inception through refinement.
  • Maintain service availability by measuring and monitoring latency and system health.
  • Scale systems sustainably using automation mechanisms.
  • Participate in on-call rotations to support production systems and conduct blame-free postmortems.

Required Skills

  • 10+ years of professional experience.
  • Experience with infrastructure automation and distributed systems design.
  • Proficiency in at least one of: Python, Go, Perl, or Ruby.
  • In-depth knowledge of Linux, Networking, and Containers.
  • Experience developing tools for large-scale private or public cloud systems in production.
  • Experience with Kubernetes, Docker, and OpenStack.
  • Ability to debug and optimize code and automate routine tasks.

Preferred Skills

  • Any Graduate.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs