Description
You will design and implement site reliability engineering architectures focused on 5-9s availability and resiliency.
Responsibilities
- Design high availability and scalability patterns to meet strict reliability requirements.
- Build and maintain CI/CD pipelines using Jenkins, Git, or GitHub.
- Develop automation solutions for system tasks and infrastructure management.
- Deploy, monitor, and observe production systems using Splunk and Datadog.
- Optimize application performance through continuous monitoring and predictive analysis.
Required Skills
- 5+ years of experience in SRE or DevOps roles.
- Expertise in AWS and Red Hat OpenShift (OCP).
- Proficiency with Linux command line and shell scripting.
- Strong programming skills in Java, Python, Bash, or Perl.
- Hands-on experience with Ansible and Terraform.
- Experience with AWS EMR and cloud-native environments.
- Deep understanding of networking, firewalls, and production environments.
- Knowledge of observability tools including Splunk and Datadog.
- Experience working within Agile Engineering teams.
Preferred Skills
- Background in studying architectural patterns at scale, including API design and repeatable delivery pipelines.