Description
You will architect and implement data pipelines to process and transform large datasets from databases, APIs, and web services.
Responsibilities
- Architect data pipelines to move data from various sources into data lakes or warehouses.
- Build scalable data processing solutions using Snowflake, AWS, and GCP.
- Develop data quality checks and monitoring systems to ensure accuracy throughout the pipeline.
- Implement Machine Learning algorithms for anomaly detection, data quality, and continuous monitoring.
- Contribute to data governance, including lineage tracking, access controls, and retention policies.
Required Skills
- 7+ years of experience with data warehouse architectures, ETL/ELT, and scripting.
- 5+ years of data modeling experience with advanced SQL and query performance tuning.
- Proficiency in Python and SQL for data manipulation and processing.
- Experience with Snowflake, including multi-cluster architecture and shareable data features.
- Hands-on experience with AWS services including S3, Lambda, and Data-pipeline.
- Working knowledge of BigQuery, Databricks, and GCP.
- Experience with Apache Airflow or AWS MWAA.
- Familiarity with CI/CD pipelines and automation tools.
- Bachelor's degree in Computer Science, Information Technology, or a related field.
Preferred Skills
- Expertise in distributed processing frameworks like Apache Spark.
- Experience with ETL tools such as Databricks or IBM DataStage.