← Back to jobs

CTS Logo
Site Reliability Engineer

CTS

 

San Jose, CA, USA

Posted On: 15+ days ago
Experience: 15+ years
Availability: Remote
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own the reliability, automation, and scalability of high-performance compute infrastructure.

This role is remote.

Responsibilities

  • Ensure availability, latency, scalability, and efficiency of NVIDIA and Cisco UCS infrastructure using fault-tolerant approaches.
  • Drive capacity planning, performance analysis, instrumentation, and non-functional system requirements.
  • Automate operational capabilities using Python, Ansible, Terraform, or Go.
  • Deliver automation through CI/CD pipelines and chatbot integration.
  • Implement metrics-driven processes to meet service quality targets.

Required Skills

  • 15+ years of experience in infrastructure engineering.
  • Expertise with DevOps Automation and CI/CD systems (GitLab, GitHub Actions, Jenkins).
  • Proficiency in infrastructure-as-code using Terraform and Ansible.
  • Strong scripting and automation skills in Python.
  • Experience managing Docker and building CI/CD Pipelines.
  • Knowledge of Enterprise Grade Kubernetes clusters, preferably RedHat OpenShift or Google Anthos.
  • Familiarity with NVIDIA (DGX) hardware (A100/H100/H200) and Cisco UCS infrastructure.

Preferred Skills

  • Experience with Go for automation tools.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs