← Back to jobs

RulesIQ Logo
Senior Site Reliability Engineer

RulesIQ

 

Jersey City, NJ, USA

Posted On: 1 day ago
Experience: 12+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own the reliability and operational integrity of large-scale distributed systems.

This role is on-site.

Responsibilities

  • Administer and optimize Kubernetes clusters, specifically Amazon EKS and Red Hat OpenShift.
  • Define SLIs/SLOs and manage error budgets, using data to balance reliability against feature velocity.
  • Design and implement observability stack components using Prometheus, Grafana, ELK, and Splunk.
  • Lead on-call rotations, incident response, and drive continuous improvement through blameless root cause analysis.

Required Skills

  • 12+ years of overall industry experience, with 6+ years in SRE, DevOps, or Production Engineering.
  • Expertise administering Kubernetes internals, networking, Helm charts, and Operators.
  • Hands-on experience managing and tuning Apache Kafka and Redis Enterprise Clusters.
  • Strong proficiency with IaC tools including Terraform, Ansible, and Helm/ArgoCD.
  • Proficiency in Python, Shell Scripting, or Groovy for automation.
  • Proven track record running large-scale production systems with minimal downtime.

Preferred Skills

  • Enforce container security and policy governance using tools like OPA/Gatekeeper or Kyverno.
  • Demonstrable experience instrumenting Java, Node.js, and Python applications with metrics and tracing.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs