← Back to jobs

Data Science Technologies Logo
Site Reliability Engineering (SRE) Architect

Data Science Technologies

 

Atlanta, GA, USA

Posted On: Just posted
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Site reliability engineering
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will architect and drive reliability strategies across the platform.

Responsibilities

  • Architect highly available, scalable, and secure infrastructure patterns on AWS.
  • Define and evangelize SRE best practices, including SLIs, SLOs, and Error Budgets.
  • Design solutions to reduce operational toil through automation and system redesign.
  • Lead blameless postmortems, ensuring systemic architectural improvements are implemented.
  • Provide senior technical guidance on reliability, scalability, and performance during design phases.

Required Skills

  • 5+ years of experience in an architectural role focused on reliability and performance.
  • Deep practical knowledge of SRE principles (SLIs/SLOs, toil reduction, incident management).
  • Expertise in cloud computing, specifically AWS infrastructure and security services.
  • Strong experience with containerization (Kubernetes, Docker) and serverless computing.
  • Solid experience implementing observability solutions using Prometheus, Grafana, Dynatrace, or OpenTelemetry.
  • Proficiency in scripting and automation using Python, Go, or Bash.
  • Proven ability to influence technical direction and lead architectural reviews.

Preferred Skills

  • Experience designing and implementing chaos engineering practices and platforms.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs