Description
Build and maintain scalable data pipelines within the Azure ecosystem.
This role is on-site.
Responsibilities
- Design and develop ETL/ELT pipelines using Azure Databricks and PySpark.
- Write optimized SQL for complex transformations, aggregations, and analysis.
- Integrate structured and unstructured data from ADLS Gen2 with Azure SQL DB, Synapse, and Data Factory.
- Optimize pipeline performance and troubleshoot bottlenecks or data quality issues.
- Translate business requirements into technical specifications and maintain documentation.
Required Skills
- 10+ years of professional experience in data engineering.
- 6+ years of hands-on experience with Azure Databricks and PySpark.
- Strong proficiency in SQL, including window functions, CTEs, indexing, and partitioning.
- Direct experience with Azure Data Factory, Synapse Analytics, and Azure SQL DB.
- Practical knowledge of Azure Data Lake Storage Gen2.
- Experience with performance tuning and data quality frameworks in cloud pipelines.
- Solid understanding of data modeling and big data architectures.
- Familiarity with Git, CI/CD pipelines, and Agile methodologies.
Preferred Skills
- Understanding of DevOps practices within the Azure ecosystem.