You will manage and maintain the Kafka streaming platform as a Site Reliability Engineer.
Responsibilities
Perform SRE duties for the Kafka Streaming Platform, including cluster maintenance and executing changes based on documented installation and validation plans.
Monitor platforms and follow runbooks and SOPs to resolve platform and application issues.
Troubleshoot and debug Kafka platform services and applications to identify and rectify root causes.
Conduct root cause analysis for production incidents and implement proactive measures to improve system reliability.
Automate routine tasks using scripts and automation tools to reduce manual intervention and human error.
Required Skills
6+ years of professional experience.
Extensive experience with Kafka and KafkaAdmin.
Deep knowledge of Kafka architecture, including producers, consumers, topics, and partitions.
Proficiency in writing Ansible playbooks for task automation.
Experience with Shell scripting and Python.
Strong understanding of Unix/Linux system internals, networking, and distributed systems.
Proven ability in troubleshooting both platform services and application-level problems.