← Back to jobs

Staffactory Logo
Site Reliability Engineer

Staffactory

 

Jersey City, NJ, USA

Posted On: 30+ days ago
Experience: 5+ years
Availability: Onsite
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the reliability and scalability of distributed systems.

This role is on-site.

Responsibilities

  • Design, implement, and maintain scalable and resilient Apache Flink deployments on Kubernetes.
  • Develop automation tools and scripts for deployment, monitoring, and maintenance of Flink jobs and infrastructure.
  • Manage and troubleshoot production Flink job operations.
  • Ensure system observability using monitoring stacks.

Required Skills

  • 5+ years of professional experience in SRE or similar infrastructure role.
  • Expert-level experience with Apache Flink in production environments.
  • Deep hands-on knowledge of Kubernetes, including Helm and Operators.
  • Proficiency in scripting languages: Python, Bash, or Go.
  • Experience implementing monitoring with Prometheus and Grafana.
  • Familiarity with logging solutions like ELK Stack.
  • Working knowledge of cloud platforms (AWS or GCP).
  • Solid understanding of container orchestration and networking principles.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs