Description
You will design and maintain scalable data pipelines and analytics architectures using Databricks, dbt Cloud, and Apache Airflow.
Responsibilities
- Build and optimize Delta Lake-based data models using star schema, fact/dimension design, and SCD Type 1 & 2.
- Develop modular, reusable, and testable SQL transformations using dbt models, macros, and packages.
- Orchestrate end-to-end workflows via Airflow DAGs and Databricks Workflows, managing dependencies, scheduling, and fault tolerance.
- Implement robust data quality checks and testing frameworks (not null, unique, referential integrity) within dbt.
- Integrate pipelines with CI/CD workflows using Git-based version control and ensure adherence to data governance and security standards.
Required Skills
- 5+ years of experience in data engineering, with a focus on pharmaceutical or clinical trial analytics.
- Expert-level proficiency with AWS services (S3, RDS) and cloud-based data governance.
- Advanced experience with Databricks, Spark, and Delta Lake for high-performance data processing.
- Strong command of SQL and dbt Cloud for data transformation and modeling.
- Proficiency in Python for data manipulation, feature engineering, and automation.
- Experience with PostgreSQL and cloud architecture design.
- Knowledge of data governance tools such as Unity Catalog and enterprise policies.
Preferred Skills
- Experience with regression analysis, time-series forecasting, and classification models.
- Familiarity with MLOps, including model deployment, monitoring, and automated retraining pipelines.
- Background in clinical operations data systems, site selection, and non-enrollment prediction.