← Back to jobs

Netorbit Inc Logo
Site Reliability Engineer
Posted On: 1 day ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will own the operational reliability of a highly scalable Infrastructure as a Service (IaaS) platform.

Responsibilities

  • Define and instrument Service Level Objectives (SLOs) and alerts for service owners.
  • Build tooling and dashboards, and facilitate postmortems to continuously enhance system reliability.
  • Troubleshoot distributed and cloud production environments, driving automated problem resolution.
  • Handle customer escalations, production outages, and on-call responsibilities to maintain 4x9 uptime.
  • Participate in product roadmap planning and drive team initiatives to improve service availability.

Required Skills

  • 5+ years of industry experience, including 3+ years managing large-scale virtualized data center environments.
  • Experience with VMware technologies, specifically vSphere, ESXi, vSAN, and/or NSX.
  • Comfortable using Python or other scripting languages for automation.
  • Solid background troubleshooting distributed and cloud production environments.
  • Experience resolving issues in large-scale distributed environments using automation.
  • Proven ability to conduct post-mortems to prevent recurrence of failures.
  • Experience handling production incidents and on-call support.
  • Familiarity with testing methodologies.

Preferred Skills

  • Experience with large-scale IaaS platform operations.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs