← Back to jobs

Triplecom Inc Logo
Senior Site Reliability Engineer (Cloud)

Triplecom Inc

 

Arlington, TX 76001, USA

Posted On: 12 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Cloud Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will manage and automate cloud infrastructure and Kubernetes clusters to ensure system resilience and reliability.

Responsibilities

  • Develop automated solutions for infrastructure management and runbook automation.
  • Create monitoring frameworks, dashboards, and alerting templates for cloud services.
  • Collaborate with application teams to define and implement SLOs and SLIs.
  • Establish and document High Availability (HA) and Disaster Recovery (DR) best practices.
  • Troubleshoot complex system issues and drive rapid incident response.

Required Skills

  • 5+ years of experience in site reliability or cloud engineering.
  • Hands-on experience with Kubernetes (K8s) clusters.
  • Proficiency with Azure cloud solutions.
  • Experience using Splunk, Azure Monitor, or Grafana for observability.
  • Strong scripting and automation skills for infrastructure management.
  • Experience designing and managing large-scale, distributed systems.
  • Ability to design complex system solutions and implement monitoring frameworks.
  • Background in system troubleshooting and incident response.

Preferred Skills

  • Experience implementing High Availability and Disaster Recovery architectures.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs