Description
You will design, implement, and maintain end-to-end data pipelines and warehouses to support web applications and databases.
This role is on-site.
Responsibilities
- Design and implement data warehouses and data marts to serve internal and external data consumers.
- Execute the full lifecycle of data pipelines, moving them from design through operationalization and ongoing maintenance.
- Model data logically at both macro and micro levels to ensure efficient consumption.
- Tune database performance and manage the complete data lifecycle.
- Support and enhance existing data pipelines and database infrastructure.
Required Skills
- 6-10 years of experience in data integration teams.
- 3+ years developing data pipelines using Apache Spark, with Databricks preferred.
- 2+ years of active work specifically with Databricks.
- 2+ years of experience with data warehouse modeling techniques.
- Strong proficiency in PySpark, Python, and SQL, including distributed computing principles.
- Experience designing and implementing ETL/ELT processes using SSIS or similar tools.
- Fluency in complex SQL queries and database performance tuning.
- Knowledge of cloud platforms (AWS or Azure) and big data technologies like Hadoop.
- Bachelor's degree in a relevant field.