← Back to jobs

United IT Solutions Logo
Staff Site Reliability Engineer

United IT Solutions

 

San Francisco, CA, USA

Posted On: 15+ days ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Lead reliability and automation for the entire compute ecosystem, driving 10x improvements in performance and efficiency.

This role is on-site.

Responsibilities

  • Design and lead large-scale, cross-functional projects to improve core service reliability and infrastructure efficiency.
  • Reduce toil by developing automation tools and championing the "everything as code" philosophy.
  • Serve as the primary technical escalation point for critical incidents, conducting deep-dive root cause analyses and implementing corrective measures.
  • Define and implement SLOs, SLIs, and Error Budgets for critical services.
  • Enhance monitoring, logging, and tracing systems to provide comprehensive visibility into system health.

Required Skills

  • 8+ years of progressive experience in Site Reliability Engineering or Production Engineering.
  • Expert-level proficiency with AWS, including networking, compute, and storage.
  • Deep expertise in Kubernetes and the cloud-native ecosystem.
  • Fluency in at least one major scripting/programming language: Python, Go, or Java.
  • Solid experience with monitoring and logging solutions, specifically Datadog.
  • Proven ability to design and implement robust, highly available distributed systems.
  • Demonstrated experience with Infrastructure as Code tools, specifically Terraform.

Preferred Skills

  • Experience implementing Service Mesh technologies (e.g., Istio, Linkerd).
  • Strong understanding of security principles and practices in a cloud environment.
  • CKA (Certified Kubernetes Administrator) or CKAD (Certified Kubernetes Application Developer) certification.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs