← Back to jobs

Snowrelic Inc Logo
AI Site Reliability Engineer

Snowrelic Inc

 

Reston, VA, USA

Posted On: 15+ days ago
Experience: 5+ years
Availability: Remote
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will maintain service level objectives for internal NVIDIA DGX and Cisco-UCS AI platforms.

This role is remote.

Responsibilities

  • Build and run automated CI/CD pipelines to deliver operational capabilities.
  • Ensure availability, latency, scalability, and efficiency of NVIDIA and Cisco UCS infrastructure using fault-tolerant approaches.
  • Drive capacity planning, performance analysis, and instrumentation for non-functional system requirements.
  • Implement metrics-driven processes to meet service quality targets.

Required Skills

  • 5+ years administering and supporting Linux-based operating systems.
  • Experience deploying and administering NVIDIA (DGX) or equivalent High-Performance Compute (HPC) clusters.
  • Proficiency with Python, Golang, or C/C++.
  • Deep knowledge of Kubernetes and Docker.
  • Experience with infrastructure-as-code tools such as Terraform and Ansible.
  • Proficiency with GIT and CI/CD systems like Jenkins or GitLab.
  • Bachelor’s degree in Computer Science, IT, or equivalent experience.

Preferred Skills

  • Familiarity with high-performance compute environments.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs