Description
You will build and manage data infrastructure to support analytics and processing needs.
Responsibilities
- Design, develop, and maintain data pipelines within Hadoop ecosystems, ensuring scalability and performance.
- Implement ETL processes for both batch and streaming analytics.
- Optimize and troubleshoot distributed systems for data ingestion, storage, and processing.
- Collaborate with engineering and analytics teams to meet business requirements.
- Maintain documentation and support data warehouse operational excellence.
Required Skills
- 5+ years of direct experience in data engineering.
- Proficiency in Hadoop components including Spark, HDFS, Hive, and Iceberg.
- Expertise in Spark SQL.
- Strong command of Python and Unix.
- Experience with SQL Server or similar relational databases.
- Familiarity with CI/CD tools like Jenkins.
- Proficiency in testing frameworks such as JUnit or pytest.
- Solid understanding of data structures and algorithms.