← Back to jobs

TechDigital Logo
SRE Specialist

TechDigital

 

Dallas, TX, USA

Posted On: 15+ days ago
Experience: 10+ years
Availability: Onsite
Openings: 1
Category: SRE Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the operational stability and automation of large-scale distributed systems.

This role is on-site.

Responsibilities

  • Guide engineering teams on Site Reliability Engineering (SRE) principles.
  • Implement and maintain disaster recovery, failover mechanisms, and backup strategies.
  • Troubleshoot complex production issues using scripting and observability data.
  • Manage ITSM processes including incident, problem, and change management.
  • Maintain and advance automation tooling within CI/CD pipelines.

Required Skills

  • 10+ years of experience in operations or SRE roles.
  • Expertise in Azure Cloud operations.
  • Hands-on experience with Dynatrace, Newrelic/DataDog, Prometheus, and Grafana.
  • Proficiency in scripting languages: Python and Shell.
  • Experience with infrastructure-as-code tools: Terraform and Ansible.
  • Familiarity with GitOps workflows and CI/CD pipeline implementation.
  • Solid understanding of ITSM processes.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs