Description
You will design, build, and maintain scalable data pipelines and data lake infrastructure on AWS.
This role is on-site.
Responsibilities
- Design and maintain ETL/ELT pipelines using AWS Glue, Lambda, Step Functions, and EMR.
- Develop scalable ingestion solutions for structured and unstructured data sources.
- Build data transformation workflows using PySpark and Spark-based frameworks.
- Manage large-scale data lakes on S3, implementing partitioning and cataloging via Glue Data Catalog.
- Optimize Amazon Redshift data models and build high-performance ELT workloads using SQL, Spectrum, and COPY commands.
Required Skills
- 7+ years of hands-on experience in data engineering.
- Deep expertise with AWS services: S3, Glue, Redshift, Lambda, IAM, CloudWatch, and EMR.
- Strong SQL skills with experience in data modeling (star/snowflake schemas).
- Hands-on experience with PySpark/Spark or similar distributed processing frameworks.
- Experience with streaming technologies such as Kinesis or Kafka.
- Deep knowledge of Amazon Redshift optimization and performance tuning.
- Strong understanding of ETL/ELT architecture and data integration patterns.
Preferred Skills
- Experience with large-scale data lake management and optimization.
- Background in building high-performance ELT workloads.