You own end-to-end data pipelines, architecture, and production reliability for high-volume batch and streaming workloads.
This role is on-site.
Responsibilities
Design, build, and operate critical batch and streaming pipelines using Glue, Redshift, Spark, Kinesis, and Lambda, handling schema evolution and late data.
Define data architecture and Lakehouse patterns for specific business domains, selecting storage and query engines (S3, Redshift, Athena) based on performance needs.
Tune Spark jobs, Glue configurations, and Redshift performance; investigate production issues and reduce latency and cloud costs.
Implement domain-level data quality checks, monitoring, and alerting; own production incidents and drive permanent fixes.
Lead by example through coding, PR reviews, and design guidance; mentor senior engineers and set engineering standards.
Required Skills
10+ years of experience in data engineering.
Deep expertise with AWS services: Glue, Redshift, Spark, Kinesis, and Lambda.
Strong experience with S3, Athena, and data modeling strategies.
Proven ability to optimize complex ETL/ELT pipelines for high-volume workloads.
Experience with schema evolution, CDC, and handling edge cases in distributed systems.
Ability to translate business requirements into scalable, reliable data solutions.
Bachelor’s degree in a relevant field or equivalent experience.