← Back to jobs

Zibal Logo
Site Reliability Engineer

Zibal

 

Leesburg, VA, USA

Posted On: 30+ days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Manage system reliability and performance through automation, capacity planning, and collaboration with development teams.

Responsibilities

  • Analyze operating system and application metrics to drive performance tuning and fault finding.
  • Partner with developers to improve service quality through rigorous testing and release procedures.
  • Consult on system design, platform management, and capacity planning.
  • Build sustainable systems using automation and technical uplifts.
  • Maintain well-defined service-level objectives to balance feature delivery speed with reliability.

Required Skills

  • 5+ years of experience in relevant engineering roles.
  • Proficiency in programming using Python, Java, C/C++, Ruby, or JavaScript.
  • Experience with distributed storage technologies including NFS, HDFS, Ceph, or Amazon S3.
  • Hands-on experience with resource management frameworks such as Kubernetes, Apache Mesos, or Yarn.
  • Strong understanding of structured and OOP programming principles.
  • Ability to identify performance bottlenecks and system improvement areas.
  • Bachelor’s degree or equivalent in computer science or a related discipline.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs