Description
Key Responsibilities
- Lead the design and development of scalable ETL/ELT pipelines, Data Lake, Lakehouse, and Data Warehouse solutions.
- Develop and optimize data pipelines using Azure Databricks, PySpark, Python, SQL, Microsoft Fabric, Azure Data Factory (ADF), and Azure Synapse Analytics.
- Design, develop, and optimize Snowflake Data Warehouse solutions and analytical data models.
- Implement Delta Lake, Medallion Architecture, CDC, incremental data processing, schema evolution, and data quality frameworks.
- Build curated datasets, semantic models, and analytical data layers supporting Power BI reporting and enterprise analytics.
- Optimize Spark jobs, SQL queries, data pipelines, cluster utilization, storage, and cloud costs.
- Implement data integration solutions across Azure services, Snowflake, Databricks, Microsoft Fabric, and enterprise source systems.
- Develop batch and near real-time data processing solutions using modern cloud data engineering technologies.
- Implement CI/CD pipelines, Git-based development, automated deployments, monitoring, logging, and production support.
- Establish data governance, security, lineage, metadata, and access-control best practices.
- Conduct code reviews, architecture discussions, technical design sessions, and mentor Data Engineers.
- Collaborate with Data Architects, BI Developers, Data Scientists, Business Analysts, and business stakeholders to deliver enterprise data solutions.
Required Skills
- 8–12+ years of Data Engineering experience.
- Strong hands-on expertise in Databricks, PySpark, Python, and SQL.
- Strong experience with Microsoft Fabric, OneLake, Fabric Data Factory, Fabric Lakehouse, and Fabric Warehouse.
- Strong hands-on Snowflake experience including data warehousing and analytical workloads.
- Experience with Azure Synapse Analytics and Azure Data Factory (ADF).
- Strong understanding of Power BI, Semantic Models, Data Modeling, and DAX.
- Hands-on experience with Delta Lake, Lakehouse Architecture, Medallion Architecture, CDC, and incremental processing.
- Experience with Azure cloud services, CI/CD, Git, DevOps, data governance, and production support.
- Strong understanding of ETL/ELT, Data Warehousing, Data Lakes, Data Integration, and Distributed Data Processing.
- Strong communication, problem-solving, technical leadership, and stakeholder management skills.
Preferred Skills
- Unity Catalog
- Fabric Lakehouse / Fabric Warehouse
- Snowflake Streams & Tasks
- Microsoft Purview
- Terraform
- Azure DevOps
- GitHub Actions
- Apache Kafka
- Azure Event Hubs
- Spark Structured Streaming
- Experience with Data Quality, Data Lineage, Metadata Management, and Data Governance
Technical Environment
Azure Databricks, Apache Spark, PySpark, Python, SQL, Delta Lake, Microsoft Fabric, OneLake, Fabric Data Factory, Fabric Lakehouse, Fabric Warehouse, Snowflake, Azure Synapse Analytics, Azure Data Factory, Power BI, DAX, Azure Data Lake Storage (ADLS), Unity Catalog, Microsoft Purview, Azure DevOps, GitHub Actions, Git, Terraform, Kafka, Azure Event Hubs, CI/CD, ETL/ELT, Data Warehousing, Lakehouse Architecture, Data Modeling, CDC, Incremental Processing, Data Governance, Data Quality, Data Lineage, Monitoring & Production Support