← Back to jobs

ACL Digital Logo
Site Reliability Engineer

ACL Digital

 

Atlanta, GA, USA

Posted On: 12 days ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Manage the reliability, scalability, and performance of cloud-based systems and applications.

This role is on-site.

Responsibilities

  • Implement and improve monitoring, alerting, and logging solutions to detect and respond to incidents.
  • Automate deployment, configuration management, and troubleshooting processes.
  • Participate in on-call rotations, triage production incidents, lead RCAs, and implement preventive actions.
  • Collaborate with development teams to deploy services that meet reliability and performance standards.
  • Conduct capacity planning and performance analysis to manage growing traffic and data volumes.

Required Skills

  • 10+ years of experience as a Site Reliability Engineer or in a similar role.
  • Deep expertise in AWS cloud infrastructure, specifically Lambda, S3, SQS, IAM, and Route 53.
  • Proficiency in Infrastructure as Code using Terraform or CloudFormation.
  • Hands-on experience with Python and Bash for scripting and automation.
  • Experience with containerization using Docker and Kubernetes.
  • Knowledge of serverless architecture and AWS Lambda.
  • Experience managing CI/CD pipelines and version control via Git.
  • Proficiency with monitoring tools such as CloudWatch, SumoLogic, Dynatrace, or Grafana.
  • Bachelor's degree in Computer Science, Engineering, or equivalent work experience.

Preferred Skills

  • AWS Certification.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs