← Back to jobs

Snowrelic Inc Logo
AI Site Reliability Engineer

Snowrelic Inc

 

Reston, VA, USA

Posted On: 30+ days ago
Experience: 5+ years
Availability: Remote
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will operate as an AI Site Reliability Engineer, focusing on maintaining Service Level Objectives for internal NVIDIA DGX and Cisco-UCS based AI platforms.

Responsibilities

  • Lead, build, and run fully automated pipelines through the CI/CD system to deliver operational capabilities.
  • Handle availability, latency, scalability, and efficiency of NVIDIA and Cisco UCS infrastructure using fault-tolerant approaches.
  • Drive capacity planning, performance analysis, and instrumentation for non-functional system requirements.
  • Automate operational capabilities using Python, Ansible, Terraform, or Go.
  • Implement metrics-driven processes to ensure service quality targets are met.

Required Skills

  • 5+ years administering and supporting Linux-based operating systems.
  • Experience deploying and administering NVIDIA (DGX) or equivalent High-Performance Compute (HPC) clusters.
  • Experience writing code in Python, Golang, or C/C++.
  • Proficiency with GIT and CI/CD systems like Jenkins or GitLab.
  • Deep knowledge of Kubernetes and Docker.
  • Experience with infrastructure-as-code tools such as Terraform and Ansible.
  • Familiarity with high-performance compute environments.
  • Bachelor’s degree in Computer Science, IT, or equivalent experience.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs