Description
You will design and maintain large-scale data pipelines and schemas to support evolving data sources.
Responsibilities
- Build batch data pipelines using big data frameworks with a focus on optimization, fault tolerance, and SLA adherence.
- Design data models and schemas that facilitate seamless joins across diverse datasets.
- Develop idempotent workflows using orchestrators such as Airflow, Luigi, or Automic.
- Write SQL to analyze, optimize, and profile data within BigQuery or Spark SQL.
- Collaborate with stakeholders to translate business requirements into technical data solutions.
Required Skills
- 4+ years of experience developing big data technologies and data pipelines.
- Proven experience managing and manipulating datasets in the terabyte (TB) scale.
- Proficiency with Apache Spark (Scala preferred), Hadoop, and Apache Hive.
- Hands-on experience with cloud platforms, specifically GCP.
- Strong SQL skills for data profiling and optimization.
- Expertise in data modeling and schema design.
- Ability to resolve complex issues during data integration and schema evolution.
Preferred Skills
- Experience building near real-time streaming pipelines using Apache Kafka, Spark streaming, and Kafka Connect.
- Working knowledge of REST APIs, Apache Druid, Redis, Elastic Search, or GraphQL.
- Exposure to Looker, Tableau, or the eCommerce domain.