Description
You will build and maintain scalable data pipelines and workflows.
Responsibilities
- Design, build, and maintain scalable data pipelines and workflows.
- Manage and optimize cloud-native data platforms on Azure using Databricks and Apache Spark.
- Implement CI/CD workflows and monitor data pipelines for performance and accuracy.
- Develop and maintain ETL/ELT pipelines using open-source frameworks like Apache Spark and Apache Airflow.
- Integrate and process data streams from message queues and streaming platforms like Kafka and RabbitMQ.
Required Skills
- 6+ years of experience in data engineering or related field.
- Strong programming skills in Python, including Pandas and NumPy.
- Proficiency in SQL and experience with relational databases (Sybase, DB2, Snowflake, PostgreSQL, SQL Server).
- Hands-on experience with pipeline monitoring and CI/CD workflows.
- Experience with Azure Databricks and Apache Spark.
- Familiarity with Git for version control.
- Experience with Kafka and RabbitMQ.
- Ability to work independently and collaborate across distributed teams.