← Back to jobs

Compunnel Logo
Site Reliability and Operations Engineer

Compunnel

 

Hoboken, NJ, USA

Posted On: 11 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will manage and optimize Kubernetes-based distributed caching and compute grid systems to ensure high availability and scalability.

This role is on-site.

Responsibilities

  • Design, build, and enhance distributed caching and compute grid solutions on Kubernetes and OpenShift.
  • Orchestrate microservices and container workloads using Docker and Helm.
  • Implement observability and monitoring frameworks using Prometheus, Grafana, ELK, or OpenTelemetry.
  • Automate infrastructure provisioning and deployments using Ansible and Helm Charts.
  • Troubleshoot complex system and infrastructure issues within Kubernetes environments.

Required Skills

  • 5+ years of experience in infrastructure or site reliability engineering.
  • Deep expertise with Kubernetes and OpenShift in on-prem and cloud environments.
  • Proficiency in Java, Go, or Python.
  • Hands-on experience with Docker and Helm.
  • Proven experience with CI/CD tools (Jenkins, ArgoCD, GitHub Actions) and pipeline integration.
  • Expertise with observability tools: Prometheus, Grafana, Loki, and Jaeger.
  • Experience with service meshes such as Istio or Linkerd.
  • Knowledge of multi-cluster and hybrid cloud Kubernetes deployments.

Preferred Skills

  • Experience with high-performance computing platforms or grid computing frameworks.
  • Familiarity with distributed caching strategies and data sharding.
  • Relevant certifications such as CKAD, CKA, or Red Hat Certified Specialist in OpenShift.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs