← Back to jobs

Traceable Inc Logo
SRE Leader
Posted On: 1 day ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: SRE Leader
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Lead a team of Site Reliability Engineers to ensure the availability, performance, and scalability of cloud-based distributed systems handling tens of billions of events per day.

Responsibilities

  • Manage and mentor SREs, establishing goals aligned with strategic objectives and supporting career growth.
  • Own availability, performance, monitoring, capacity planning, and emergency response for cloud services and infrastructure.
  • Build and maintain modern CI/CD and DevOps infrastructure to automate deployment pipelines.
  • Debug and resolve production issues and escalations in collaboration with product engineering teams.

Required Skills

  • 10+ years of experience in SRE and DevOps working with large-scale distributed systems.
  • Expertise with cloud-native technologies including AWS, GCP, microservices, containers, and Kubernetes.
  • Hands-on experience with streaming systems such as Kafka Streams or Flink.
  • Experience operationalizing and scaling data systems like MongoDB, Apache Pinot, Apache Trino, Spark, and Apache Iceberg.
  • Proficiency in infrastructure as code using Terraform, Helm, or Ansible.
  • Strong knowledge of Linux systems and troubleshooting production environments.
  • Proficiency in Java and scripting languages.

Preferred Skills

  • Bachelor's or Master's degree in Computer Science.

Education

Bachelor's degree in Computer Science

Related Jobs

No related jobs found

← Back to jobs