Description
You will manage and maintain the Kafka streaming platform as a Site Reliability Engineer.
Responsibilities
- Perform SRE duties for the Kafka streaming platform, including cluster maintenance and implementing changes via documented installation and validation plans.
- Monitor platforms and follow runbooks and SOPs to manage platform and application issues.
- Troubleshoot and debug Kafka platform services and application problems to identify root causes.
- Conduct root cause analysis for major production incidents and implement proactive measures to improve reliability.
- Automate routine tasks using Ansible, shell scripting, or Python to reduce manual intervention and human error.
Required Skills
- 6 to 10 years of professional experience.
- Deep knowledge of Kafka architecture, including producers, consumers, topics, and partitions.
- Experience working as a Site Reliability Engineer specifically for Kafka platforms.
- Proficiency in writing Ansible playbooks for task automation.
- Experience with shell scripting and Python.
- Strong understanding of Unix/Linux system internals.
- Knowledge of networking and distributed systems.
- Ability to perform thorough debugging and root cause analysis.