Description
You will design, build, and optimize scalable data pipelines and warehouses on AWS.
This role is remote.
Responsibilities
- Design and implement ETL processes using AWS Glue, Python, and SQL for data ingestion, transformation, and loading.
- Develop and maintain data lakes and warehouses on Amazon S3 and Amazon Redshift, ensuring scalability and security.
- Optimize data storage, performance, and cost efficiency using Redshift Spectrum and query tuning.
- Define data models, governance standards, and solution designs in collaboration with cross-functional teams.
- Troubleshoot performance bottlenecks and ensure system reliability adhering to AWS Well-Architected practices.
Required Skills
- 6+ years of hands-on experience with the AWS analytics stack, including AWS Glue, Amazon Redshift, and Amazon S3.
- Proficiency in Python for data processing, automation, and building reusable frameworks.
- Strong SQL expertise for complex queries, transformations, and data validation.
- Experience with data modeling, schema design, and dimensional modeling (Star/Snowflake schemas).
- Good understanding of data architecture, integration patterns, and solution design principles.
- Exposure to data governance, cataloging, and security best practices within the AWS environment.
- Experience implementing data quality checks and monitoring across ETL pipelines.
Preferred Skills
- Familiarity with Tableau and DataIQ for data analytics and visualization.
Bachelor's degree required.