← Back to jobs
Sunnyvale, CA, USA
No related jobs found
Job Responsibilities include:
Design, deploy, and maintain highly available and scalable cloud infrastructure on AWS using services such as EC2, EKS, S3, RDS, Lambda, IAM, and VPC.
Manage and optimize Kubernetes-based environments, including container orchestration, autoscaling, configuration management, and application deployments.
Build, maintain, and enhance CI/CD pipelines and Infrastructure-as-Code (IaC) solutions using Terraform and cloud automation tools.
Implement and support monitoring, logging, alerting, and observability platforms to ensure system reliability, performance, and operational visibility.
Lead incident response, root cause analysis, post-mortem reviews, and reliability improvement initiatives to minimize downtime and operational risks.
Collaborate with development and engineering teams to embed reliability, scalability, security, and performance best practices throughout the software development lifecycle.
Required Qualifications:
Bachelor’s degree in computer science, Information Technology, Software Engineering, Engineering, or a related technical field.
6–8 years of experience in Site Reliability Engineering (SRE), Cloud Operations, DevOps, or Infrastructure Engineering.
Strong expertise in AWS cloud services, Kubernetes, Infrastructure as Code (Terraform/CloudFormation), CI/CD pipelines, and cloud automation.
Proven experience with monitoring and observability tools, incident management, root cause analysis, and maintaining highly available, scalable production systems
Any Graduate
No related jobs found
← Back to jobs