Description
You will apply SRE and DevOps principles to maintain the stability and performance of core data platforms.
Responsibilities
- Drive operational excellence by automating CI/CD pipelines and infrastructure (IaC).
- Lead incident response for key systems including Databricks, Informatica, and Power BI.
- Maintain high service reliability by monitoring SLIs/SLOs for the data and analytics ecosystem.
- Utilize observability tools, such as Dynatrace, to track system health.
Required Skills
- 4+ years of hands-on experience in DevOps, SRE, or Cloud Infrastructure.
- Practical experience with AWS and Azure.
- Proficiency in Python or scripting languages (e.g., Bash, Go) for automation.
- Experience with Configuration Management tools, specifically Ansible.
- Strong knowledge of containerization technologies like Docker and Kubernetes.
- Solid understanding of Linux systems and networking fundamentals (TCP/IP, DNS, Load Balancing).
- Working knowledge of relational, cloud-native (e.g., AWS RDS), and NoSQL databases.
- Experience with CI/CD practices.
- Familiarity with Power BI.