← Back to jobs

ValueLabs Logo
PySpark Developer

ValueLabs

 

Hyderabad, Telangana, India

Posted On: 30+ days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: PySpark Developer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will build and maintain scalable ETL pipelines within the Cloudera Data Platform ecosystem.

This role is on-site.

Responsibilities

  • Design, develop, and maintain optimized ETL pipelines using PySpark to ensure data integrity.
  • Implement data ingestion processes from relational databases, APIs, and file systems into data lakes or warehouses.
  • Process, cleanse, and transform large datasets using PySpark to meet analytical requirements.
  • Perform tuning of PySpark code and Cloudera components to optimize resource utilization and reduce runtime.
  • Automate data workflows using orchestration tools such as Apache Oozie or Airflow.

Required Skills

  • 5+ years of experience in data engineering.
  • Proficiency in PySpark for large-scale data processing.
  • Hands-on experience with ETL process development.
  • Experience with Data Pipeline Development and Data Transformation.
  • Practical knowledge of the Cloudera Data Platform.
  • Ability to manage data ingestion from diverse sources including APIs and relational databases.
  • Experience with orchestration tools like Apache Oozie or Airflow.
  • Any Graduate degree.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs