You will design and manage large-scale data processing workflows within a Hadoop ecosystem.
Responsibilities
- Develop and maintain data pipelines using Spark and Scala.
- Write and optimize scripts in Python, Perl, and Shell to automate data tasks.
- Manage real-time data streaming using Kafka.
- Build and scale distributed data processing jobs on Hadoop clusters.
Required Skills
- 12+ years of experience in data engineering.
- Expertise in Hadoop ecosystems.
- Proficiency in Spark and Scala.
- Strong Python programming skills.
- Experience with Kafka for stream processing.
- Ability to write Shell scripts for automation.
- Experience with Perl scripting.
Education
- Bachelor's degree or equivalent graduate level education.