← Back to jobs

Synechron Logo
PySpark Data Engineer

Synechron

 

Dubai - United Arab Emirates

Posted On: Just posted
Experience: 3+ years
Availability: Hybrid
Openings: 1
Category: PySpark Data Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will design, develop, and maintain scalable ETL pipelines using PySpark on the Cloudera Data Platform.

This role is on-site.

Responsibilities

  • Develop and maintain optimized ETL pipelines using PySpark, ensuring data integrity on the CDP.
  • Implement data ingestion from relational databases, APIs, and file systems into the data lake/warehouse.
  • Process, cleanse, and transform large datasets using PySpark for analytical needs.
  • Tune PySpark code and Cloudera components to optimize resource utilization and runtime.
  • Automate data workflows using orchestration tools like Apache Oozie or Airflow.

Required Skills

  • 3+ years of experience as a Data Engineer focused on PySpark.
  • Advanced proficiency in PySpark (RDDs, DataFrames, optimization).
  • Strong experience with Cloudera Data Platform (CDP), including Manager, Hive, Impala, HDFS, and HBase.
  • Proficiency in SQL and data warehousing concepts/ETL best practices.
  • Familiarity with Big Data technologies such as Hadoop and Kafka.
  • Experience with orchestration frameworks like Apache Oozie or Airflow.
  • Strong scripting skills in Linux.
  • Bachelor's or Master's degree in Computer Science or a related field.

Education

Bachelor's or Master's degrees

Related Jobs

No related jobs found

← Back to jobs