Description
You will manage and optimize cloud-based infrastructure and automation workflows.
Responsibilities
- Design and maintain scalable cloud solutions using AWS.
- Develop Python scripts and tools to automate system tasks and improve reliability.
- Monitor system performance and observability using New Relic, Grafana, or Prometheus.
- Troubleshoot and resolve performance issues, incidents, and infrastructure bottlenecks.
- Optimize infrastructure for availability, scalability, and cost-efficiency while contributing to CI/CD pipeline enhancements.
Required Skills
- 5+ years of experience in SRE or DevOps roles.
- Proficiency in Python for scripting and automation.
- Strong experience with AWS services including EC2, S3, Lambda, CloudWatch, and IAM.
- Hands-on experience with monitoring and observability tools such as New Relic, Grafana, or Prometheus.
- Experience with containerization using Docker and orchestration with Kubernetes.
- Knowledge of Infrastructure as Code (IaC) using Terraform or CloudFormation.
- Familiarity with CI/CD pipelines and automation frameworks.
- Understanding of DevOps principles and Site Reliability Engineering practices.
Preferred Skills
- Experience with ELK Stack (Elasticsearch, Logstash, Kibana) for logging and analytics.
- Knowledge of networking, security best practices, and compliance frameworks.
- Exposure to incident management, on-call support, and GitOps practices.