← Back to jobs
Austin, TX, USA
No related jobs found
Job Responsibilities include:
Design, deploy, and maintain scalable AWS cloud infrastructure using services such as EC2, EKS, S3, RDS, Lambda, IAM, and VPC.
Manage and optimize Kubernetes environments, including container orchestration, autoscaling, configuration management, and service reliability.
Build and maintain CI/CD pipelines and Infrastructure-as-Code (IaC) solutions using Terraform and CloudFormation to support automated deployments.
Implement and manage monitoring, logging, alerting, and observability solutions using Prometheus, Grafana, CloudWatch, ELK, and related tools.
Lead incident response, root cause analysis, and post-incident reviews while driving improvements in system resiliency and operational excellence.
Collaborate with development teams to automate processes, improve application performance, optimize cloud costs, and embed reliability and security best practices throughout the SDLC.
Required Qualifications:
Bachelor’s degree in computer science, Information Technology, Software Engineering, Engineering, or a related technical field.
6–8 years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Infrastructure, or Platform Engineering.
Strong expertise in AWS cloud services, Kubernetes, Infrastructure-as-Code (Terraform/CloudFormation), CI/CD pipelines, and cloud automation.
Proven experience with monitoring and observability tools (Prometheus, Grafana, CloudWatch, ELK), incident management, root cause analysis, and maintaining highly available production systems
Any Graduate
No related jobs found
← Back to jobs