Description
You will own the infrastructure and reliability practices for our services.
Responsibilities
- Design and maintain CI/CD pipelines using Jenkins and IaC practices.
- Implement and manage monitoring stacks including OpenTelemetry, ELK, Splunk, and CloudWatch.
- Automate infrastructure provisioning using Terraform.
- Apply Site Reliability Engineering (SRE) principles to production systems.
Required Skills
- 8+ years of professional experience in DevOps or SRE.
- Proficiency in Python and Bash scripting.
- Experience with cloud environments, specifically AWS.
- Hands-on experience with Kubernetes (EKS) and ECS.
- Familiarity with infrastructure as code (IaC) tools like Terraform.
- Knowledge of Java and Spring Boot environments.
- Experience with monitoring tools such as OpenTelemetry, ELK, or Splunk.