Description
Key Skills: AWS, Databricks, Python, PySpark, Data Engineering, SQL, CI/CD, Data Modeling, Terraform, Docker
Good to Have Skills: AWS certification (Developer, Solutions Architect, Data Analytics) and/or Databricks certification. Experience with data transformation frameworks (e.g., dbt) and workflow orchestration tools (e.g., Airflow). Experience building and operating data products with product thinking, domain-aligned datasets, documentation, SLAs, and adoption/usage metrics. Experience with data governance and access controls tools such as Collibra (catalog/lineage/metadata) and Immuta (policy-based data access). Streaming data experience (e.g., Kafka, Kinesis) and building near real-time pipelines. Exposure to modern data architecture patterns such as data mesh and data fabric, and experience defining reusable platform capabilities.
Roles & Responsibilities:
- Design, develop, and operate end-to-end data pipelines (ETL/ELT) to ingest data from diverse sources into an AWS-based data lakehouse and data warehouse.
- Build curated datasets using strong data modeling practices including dimensional modeling (star/snowflake), SCD patterns, and conformed dimensions to support BI and self-service analytics.
- Partner with product managers, analysts, and data scientists to understand requirements, define source-to-target mappings, and deliver datasets that are accurate, discoverable, and reusable.
- Define and implement data quality controls including validation rules, reconciliations, anomaly checks, data contracts, and SLAs while partnering with governance to maintain catalog, lineage, and business/technical metadata.
- Implement orchestration, logging, monitoring, and alerting to ensure reliable operations including data observability, pipeline health, backfills, and incident triage.
- Apply engineering best practices including automated unit/integration tests, code reviews, and CI/CD to promote changes safely across environments.
- Develop transformations on Databricks using Python/PySpark and Spark SQL while optimizing jobs for performance and cost through partitioning, file sizing, caching, and tuning.
- Write and optimize complex SQL for analysis, transformations, and warehouse/lakehouse consumption patterns.
- Build on AWS using services such as S3, IAM, Glue, Lambda, Step Functions, EMR/ECS/Fargate, and CloudWatch while implementing secure access patterns and least-privilege principles.
- Use Infrastructure as Code (Terraform) to provision and manage cloud resources while promoting reusable modules and automated deployments.
- Coach and mentor other data engineers through onboarding, pairing, code/design reviews, and knowledge-sharing sessions to raise the team's engineering bar.
- Provide technical leadership by proposing patterns/standards for lakehouse, warehousing, data quality, CI/CD, influencing the roadmap, and communicating trade-offs clearly to stakeholders.
Experience Required: 8+ years of experience in Data Engineering, building production-grade data platforms and pipelines. Strong experience working in Agile delivery teams, collaborating effectively across roles, and communicating clearly with stakeholders regarding requirements, trade-offs, and timelines.
Education: Bachelor's degree (or equivalent experience) in Computer Science, Engineering, or a related field