← Back to jobs

Cloud Space LLC Logo
Site Reliability Engineer (SRE)

Cloud Space LLC

 

New Jersey, USA

Posted On: 2 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the reliability, observability, and automation of application infrastructure.

This role is on-site.

Responsibilities

  • Forecast demand and conduct capacity planning to optimize resource utilization.
  • Develop and maintain automation tools to reduce manual intervention in infrastructure management.
  • Build observability frameworks using Prometheus and Grafana to define metrics, trends, and alerts.
  • Manage incident response, including post-mortems and the implementation of corrective actions.
  • Collaborate with engineering teams to optimize system performance and ensure adherence to SLOs and SLIs.

Required Skills

  • 5+ years of Site Reliability Engineering experience.
  • Expertise in infrastructure management and building high-availability systems.
  • Hands-on experience with Prometheus and Grafana.
  • Proficiency in automation scripting.
  • Experience with capacity planning and performance optimization.
  • Proven track record in incident management and response.
  • Bachelor's degree or equivalent experience.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs