Description
Lead the design and implementation of efficient ETL solutions and data pipelines on the Cloudera Data Platform.
Responsibilities
- Build and schedule data pipelines using Apache Spark, Python, Kafka, and Hive to handle large volumes of data within SLAs.
- Convert functional requirements into high-level and low-level technical designs, including source-to-target documentation.
- Optimize existing Spark-based ingestion and aggregation pipelines to improve performance and meet delivery targets.
- Lead code reviews, test case reviews, and implement data ingestion and governance frameworks.
- Manage production implementation, execute change requests, and oversee large-scale data migrations or history rebuilds.
Required Skills
- 5+ years of experience in data engineering or related technical leadership roles.
- Expertise in Apache Spark for large-scale data aggregation and performance tuning.
- Proficiency in Python and building ETL solutions.
- Hands-on experience with Hadoop and Kafka.
- Experience working with Hive and the Cloudera Data Platform.
- Strong ability to translate business use cases into technical tasks and effort estimates.
- Proven troubleshooting skills for resolving complex technical issues and bugs in production.
- Ability to work independently with global teams to mitigate project delivery risks.
- Any Graduate degree.