Description
You will design, develop, and maintain scalable ETL/ELT pipelines and Lakehouse architectures.
This role is on-site.
Responsibilities
- Design and maintain scalable ETL/ELT pipelines using Scala and Apache Spark or Flink for batch and streaming workloads.
- Optimize Spark jobs and SQL queries to improve performance, efficiency, and cost.
- Implement Lakehouse architectures using Apache Iceberg, Hudi, or Delta Lake, applying Medallion Architecture (Bronze/Silver/Gold).
- Enable data observability to ensure data freshness, lineage, and reliability.
Required Skills
- Strong proficiency in Scala and Apache Spark (Batch & Streaming).
- Solid understanding of SQL and distributed computing concepts.
- Experience with GCP (Dataproc, GCS, BigQuery) or equivalent cloud platforms (AWS/Azure).
- Hands-on experience with Docker and Kubernetes.
- Experience with Lakehouse table formats (Iceberg, Hudi, Delta).
- Familiarity with CI/CD practices.
- Bachelor's degree required.
- 1–5 years of experience as a Data Engineer or Big Data Engineer.
Preferred Skills
- Experience building data pipelines for ML or feature engineering.
- Exposure to workflow orchestration tools such as Airflow or Azkaban.