Description
You will design, build, and operate ingestion pipelines for batch and near-real-time data.
Responsibilities
- Implement CDC-based ingestion patterns for databases, SaaS platforms, and external partners.
- Standardize ingestion frameworks for files, APIs, event streams, and cross-account data sharing.
- Define and maintain raw and staging data models preserving source fidelity and lineage.
- Partner with source system owners to define ingestion SLAs, contracts, schemas, and change management.
- Automate infrastructure using AWS CDK and integrate CI/CD pipelines.
Required Skills
- 5+ years of experience in data engineering focused on data ingestion and integration.
- Strong understanding of Change Data Capture (CDC) concepts and ingestion patterns.
- Hands-on experience with AWS services including Lambda, Step Functions, MWAA, Glue, and Redshift.
- Experience building ingestion pipelines for APIs, files, databases, and event-based systems.
- Proficiency in Python and familiarity with JSON, Parquet, and Avro.
- Experience implementing infrastructure as code using AWS CDK.
- Working knowledge of CI/CD, version control, and automated deployments.
- Familiarity with APIs and general data engineering principles.