← Back to jobs

VBeyond Corporation Logo
SRE Engineer

VBeyond Corporation

 

Edison, NJ, USA

Posted On: 1 day ago
Experience: 5+ years
Availability: Remote
Openings: 2
Category: SRE Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will support large-scale GCP/AWS/EKS platforms, focusing on observability, Kubernetes, and cloud infrastructure reliability.

This role is remote.

Responsibilities

  • Own the observability stack, including Prometheus, Grafana, OpenTelemetry, Loki/ELK/Splunk, Jaeger, Alertmanager, and SLO frameworks.
  • Build intelligent monitoring pipelines to ensure reliable metric, log, and trace ingestion, alongside analytics.
  • Develop Terraform modules for observability infrastructure, Kubernetes components, cluster add-ons, and monitoring services.
  • Improve cluster reliability through automation, performance tuning, capacity modeling, and event-driven remediation.
  • Lead operational readiness, SLO reporting, incident management, and root cause analysis for platform outages.

Required Skills

  • 5+ years in SRE, Infrastructure, or Kubernetes operations.
  • Strong knowledge of EKS, ECS, GKE, Kubernetes internals, and cluster operations.
  • Expertise in observability stacks including Prometheus, OTel, Grafana, ELK, Datadog, and Splunk.
  • Advanced Terraform IaC and automation skills; Python or Go preferred.
  • Experience with CI/CD, cloud networking, and service mesh (Istio).
  • Experience with capacity planning for large-scale systems.
  • Familiarity with GCP and AWS environments.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs