← Back to jobs

ValueLabs Logo
PySpark Developer

ValueLabs

 

Bengaluru, Karnataka, India

Posted On: 3 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: PySpark Developer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will build and maintain scalable ETL pipelines on the Cloudera Data Platform.

This role is on-site.

Responsibilities

  • Build and optimize ETL pipelines using PySpark to ensure data integrity and accuracy.
  • Implement data ingestion from relational databases, APIs, and file systems into the data lake or warehouse.
  • Transform and cleanse large datasets using PySpark to meet analytical requirements.
  • Tune PySpark code and Cloudera components to optimize resource utilization and reduce runtime.
  • Automate data workflows using Apache Oozie, Airflow, or similar orchestration tools.

Required Skills

  • 5+ years of experience in data engineering roles.
  • Advanced proficiency in PySpark, including RDDs, DataFrames, and optimization techniques.
  • Hands-on experience with Cloudera Data Platform (CDP) components: Hive, Impala, HDFS, and HBase.
  • Strong SQL skills and experience with Hive or Impala for data warehousing tasks.
  • Experience with big data technologies including Hadoop and Kafka.
  • Proficiency with orchestration frameworks like Apache Oozie or Airflow.
  • Strong Linux scripting skills for automation.

Preferred Skills

  • Experience with data quality validation routines and pipeline performance monitoring.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs