Design and implement scalable and reliable data pipelines to ingest, process, and store diverse data at scale, utilizing technologies such as Apache Spark, Hadoop, and Kafka.
Work within cloud environments like AWS or Azure to leverage services including EC2, RDS, S3, Lambda, and Azure Data Lake for efficient data handling and processing.
Architect and operationalize data pipelines following Medallion Architecture best practices within a Lakehouse framework—ensuring data quality, lineage, and usability across Bronze, Silver, and Gold layers.
Develop and optimize data models and storage solutions (Databricks, Data Lakehouses) to support operational and analytical applications, ensuring data quality and accessibility.
Lead the Data Management Community of Practice, serving as the primary facilitator, coordinator, and spokesperson. Drive knowledge sharing, establish best practices, and represent the data engineering discipline across the organization.
What's Needed?
Bachelor's Degree in Computer Science, MIS, or other business discipline and 10+ years of experience in data engineering, with a proven track record in designing and operating large-scale data pipelines and architectures.
Demonstrated experience designing and implementing Medallion Architecture in a Databricks Lakehouse environment, including layer transitions, data quality enforcement, and optimization strategies.
Expertise in developing ETL/ELT workflows.
Comprehensive knowledge of platforms and services like Databricks, Dataiku, and AWS native data offerings.
Solid experience with big data technologies (Apache Spark, Hadoop, Kafka) and cloud services (AWS, Azure) related to data processing and storage