Description
You will design and build scalable data solutions across our AWS environment.
Responsibilities
- Design and develop ETL/ELT pipelines for data ingestion, processing, and transformation from diverse sources into AWS data warehouses and data lakes.
- Architect and optimize scalable data architectures using data lakes (S3, Delta Lake, Iceberg) and data warehouses (Redshift, Snowflake) on AWS.
- Collaborate with data scientists and analysts to define data models and schemas supporting data-driven decision-making.
- Implement data quality and governance practices to ensure data accuracy and compliance.
- Troubleshoot complex data pipeline failures and recommend system improvements.
Required Skills
- 5+ years experience as a Data Engineer focused on AWS cloud services for data solutions.
- Expertise in designing and optimizing complex data pipelines using AWS Glue, Lambda, EMR, S3, Redshift, Athena, Step Functions, DynamoDB, and Lake Formation.
- Proficiency in Python (including PySpark) and SQL (advanced, query tuning).
- Experience with big data technologies such as Apache Spark, Hadoop, and Kafka.
- In-depth knowledge of data modeling techniques (relational, dimensional) and schema design.
- Experience with containerization using Docker and Kubernetes.
- Familiarity with DevOps practices and CI/CD pipelines (AWS CodePipeline, Jenkins).
- Demonstrated experience supporting ML/AI projects, including feature engineering pipelines.
- Strong skills in problem-solving, analytical thinking, and cross-functional communication.