Description
You will build and maintain scalable data pipelines within the Azure ecosystem.
Responsibilities
- Design and develop ETL/ELT pipelines using Azure Databricks and PySpark to ingest and transform data from diverse sources.
- Write optimized SQL queries for complex data transformations, aggregations, and analysis.
- Integrate structured and unstructured data from Azure Data Lake (ADLS Gen2) with Azure SQL DB, Synapse Analytics, and Data Factory.
- Optimize pipeline performance and troubleshoot bottlenecks, failures, or data quality issues.
- Translate business requirements into technical specifications while maintaining documentation for workflows and schemas.
Required Skills
- 10+ years of professional experience in data engineering.
- 6+ years of hands-on experience with Azure Databricks and PySpark.
- Strong proficiency in SQL and Advanced SQL (window functions, CTEs, indexing, and partitioning).
- Direct experience with Azure Data Factory, Synapse Analytics, and Azure SQL DB.
- Practical knowledge of Azure Data Lake Storage Gen2.
- Experience with performance tuning and data quality frameworks in cloud pipelines.
- Solid understanding of data modeling and big data architectures.
- Familiarity with Git, CI/CD pipelines, and Agile methodologies.
- Understanding of DevOps practices within the Azure ecosystem.