Description
You will architect and develop scalable data platforms and pipelines within the Azure ecosystem.
Responsibilities
- Architect and implement scalable data warehouses on Azure Databricks using dimensional modeling.
- Design and optimize ETL/ELT pipelines using Python and PySpark for large-scale data processing.
- Establish ingestion and transformation workflows following medallion architecture (Bronze, Silver, Gold).
- Manage Azure Data Lake Storage (ADLS) and Delta Lake with Unity Catalog for governance and security.
- Optimize Databricks clusters, jobs, and workflows to ensure performance and reliability.
- Implement CI/CD pipelines for deployment and testing using Azure DevOps and GitHub.
Required Skills
- 8+ years of experience in Data Engineering.
- Expertise in Azure Databricks and PySpark.
- Deep understanding of medallion architecture (Bronze, Silver, Gold).
- Hands-on experience with ADLS, Delta Lake, and Unity Catalog.
- Strong proficiency in Python and Spark optimization.
- Experience with distributed computing and cluster performance tuning.
- Proficiency in GitHub and Azure DevOps for CI/CD pipelines.
- Proven experience in ETL/ELT pipeline development.
- Strong knowledge of data warehouse design and dimensional modeling.