← Back to jobs

ArrowCore Group Logo
Data Center Site Reliability Engineer

ArrowCore Group

 

Portland, OR, USA

Posted On: Just posted
Experience: 5+ years
Availability: Onsite
Openings: 8
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will own the reliability and uptime of on-premises and cloud-based data center environments.

This role is on-site.

Responsibilities

  • Design and manage monitoring, logging, and alerting systems using Prometheus, Grafana, and PagerDuty.
  • Develop infrastructure-as-code with Pulumi and Terraform, and maintain CI/CD pipelines via Buildkite and ArgoCD.
  • Participate in on-call rotations, respond to incidents, and perform root cause analysis.
  • Analyze system performance, forecast capacity needs, and optimize resource utilization.
  • Collaborate with hardware, networking, and software engineering teams to implement resilient solutions.

Required Skills

  • 5+ years in site reliability engineering or large-scale infrastructure management.
  • Expert knowledge of Kubernetes (on-prem and cloud) and infrastructure-as-code tools (Pulumi, Terraform).
  • Proficiency in systems programming (Rust, C++, or Go) with strong automation skills.
  • Deep understanding of monitoring and observability technologies.
  • Experience with CI/CD systems, specifically Buildkite.
  • Strong troubleshooting skills across hardware, networking, and distributed systems.
  • Experience in incident management and root cause analysis.
  • Bachelor's degree in Computer Science, Engineering, or a related field.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs