Description
Key Skills: Azure Databricks, PySpark, Apache Spark, SQL, Azure Synapse, Apache Kafka, Big Data, Python, Java, Azure Data Factory
Good to Have Skills: Cassandra, Microservices, API development, Apache Flink, ZeroMQ, Protocol Buffers, CI/CD with Jenkins and GitHub, Dynatrace for observability, Delta Lake, Unity Catalog, Azure Data Lake Storage Gen2, dimensional modeling, data warehouse design, data governance and security, Lakehouse and Lambda architectures, capital markets and securities processing experience.
Roles & Responsibilities:
- Architect and deliver enterprise-scale Azure Lakehouse platforms supporting petabyte-scale data processing.
- Design real-time streaming solutions using Kafka and Spark Structured Streaming for high-volume data processing.
- Lead migration of on-premise data warehouses to Azure Synapse and Databricks to improve performance and reduce costs.
- Build reusable PySpark frameworks for data ingestion, transformation, and validation across multiple projects.
- Optimize Spark jobs through partitioning, caching, and memory tuning to achieve significant runtime improvements.
- Implement end-to-end Azure Data Factory pipelines with CI/CD and parameter-driven orchestration.
- Provide technical mentorship to senior and junior engineers and architectural guidance across multiple teams.
- Collaborate with business stakeholders to translate analytics requirements into scalable technical designs.
- Manage and lead teams of 40+ members, govern delivery processes and drive innovation initiatives.
- Act as Module lead for India operations and liaise with North America teams to drive business objectives