Description
You will design, develop, and maintain scalable data pipelines and infrastructure to support analytics and processing needs across the Hadoop ecosystem.
This role is hybrid.
Responsibilities
- Design and implement ETL processes for both batch and streaming analytics.
- Build and manage data infrastructure within Hadoop ecosystems, ensuring scalability and performance.
- Optimize and troubleshoot distributed systems for data ingestion, storage, and processing.
- Collaborate with engineering and analytics teams to translate business requirements into technical solutions.
- Maintain comprehensive documentation and support data warehouse operational excellence.
Required Skills
- 5+ years of direct experience in data engineering.
- Proficiency in Hadoop components: Spark, HDFS, Hive, and Iceberg.
- Expertise in Spark SQL for complex data transformations.
- Strong command of Python and Unix shell scripting.
- Experience with SQL Server or similar relational databases.
- Familiarity with CI/CD tools, specifically Jenkins.
- Proficiency in testing frameworks such as JUnit or pytest.
- Solid understanding of data structures and algorithms.
Preferred Skills
- Bachelor's degree in Computer Science or related field.