← Back to jobs

Arista Networks Inc Logo
Site Reliability Engineer
Posted On: 15+ days ago
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will design, build, and deploy production systems focusing on scalability, reliability, observability, and performance while meeting security standards.

This role is remote.

Responsibilities

  • Develop and maintain automation to eliminate toil and streamline operational efficiency.
  • Monitor production systems, establish alerting strategies, and implement automated incident response.
  • Create incident runbooks and conduct postmortem analyses to prevent recurrence.
  • Collaborate with engineering teams to resolve infrastructural bottlenecks.
  • Manage and optimize monitoring infrastructure for system visibility.

Required Skills

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience (5+ years in infrastructure/systems).
  • Proficiency in Go, Python, or bash shell scripting for automation.
  • Strong knowledge of Linux or UNIX administration and debugging.
  • Hands-on experience operating large-scale production systems and applications.
  • Demonstrated expertise in infrastructure-as-code principles.
  • Experience with server provisioning, including storage and networking.
  • Experience with incident response and postmortem analysis.
  • Proficiency with Terraform, Kubernetes, Prometheus, Grafana, Docker, PostgreSQL, Microsoft Azure, and CI/CD.
  • Proven ability to work collaboratively in cross-functional teams.

Preferred Skills

  • Experience with container orchestration platforms like Kubernetes.
  • Proficiency in managing monitoring stacks including Prometheus and Grafana.
  • Familiarity with cloud platforms such as Google Cloud Platform or Amazon Web Services.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs