← Back to jobs

ConnectWise LLC Logo
Site Reliability Engineer

ConnectWise LLC

 

United States

Posted On: 1 day ago
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the reliability, performance, and operational health of large-scale distributed systems.

This role is remote.

Responsibilities

  • Build systems and infrastructure to monitor complex, large-scale distributed environments.
  • Triage critical production issues by identifying stability and performance bottlenecks.
  • Represent the SRE organization in design reviews and operational readiness exercises.
  • Devise methods to actively monitor system throughput, capacity, and reliability.
  • Debug complex systems and evolve running environments without causing downtime.

Required Skills

  • 5+ years of experience as a System Administrator with programming skills.
  • Proficiency in Python scripting.
  • Experience with infrastructure-as-code tools like Terraform.
  • Strong understanding of Linux system administration and Unix networking concepts (TCP/IP, HTTP).
  • Experience with monitoring and logging solutions including Prometheus, Grafana, and the ELK Stack.
  • Experience analyzing logs and troubleshooting large-scale distributed systems.
  • Fundamental knowledge of virtualization, storage, networking, server, and security concepts.
  • Experience supporting multi-tier web application architectures.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs