← Back to jobs

HARP Technologies & Services Pvt Ltd Logo
Data Engineer - Pyspark
Posted On: 2 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: Data Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Job Description

We are seeking a skilled PySpark Developer to join our Data Engineering team. The ideal candidate should possess strong hands-on experience in PySpark, Python, and SQL, with expertise in developing scalable data processing solutions and ETL pipelines. The candidate will be responsible for designing, developing, optimizing, and maintaining high-performance data applications that support enterprise analytics and business intelligence initiatives.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark.
  • Develop data processing solutions for large-scale structured and unstructured datasets.
  • Build, optimize, and troubleshoot Spark applications for performance and scalability.
  • Perform data extraction, transformation, validation, and loading activities.
  • Develop reusable and efficient data engineering frameworks and components.
  • Collaborate with business analysts, data architects, and cross-functional teams to understand data requirements.
  • Ensure data quality, integrity, and governance across data platforms.
  • Monitor and support production data pipelines and resolve performance bottlenecks.
  • Participate in code reviews and follow best practices for development and deployment.

Required Skills

  • Strong hands-on experience with PySpark
  • Proficiency in Python
  • Strong knowledge of SQL
  • Experience with ETL/Data Engineering projects
  • Good understanding of Data Warehousing concepts
  • Experience with Spark SQL, DataFrames, and distributed data processing
  • Knowledge of performance tuning and optimization techniques in Spark
  • Experience working in Linux/Unix environments
  • Strong analytical and problem-solving skills

Preferred Skills

  • Databricks
  • Delta Lake
  • Apache Kafka
  • Apache Airflow
  • Hadoop Ecosystem (Hive, HDFS)
  • Snowflake
  • Azure Data Factory (ADF)
  • AWS Glue, EMR, S3
  • Azure Databricks
  • CI/CD tools and Git

Technical Competencies

  • Spark SQL
  • DataFrame API
  • Window Functions
  • Joins and Aggregations
  • Partitioning and Bucketing
  • Caching and Persistence
  • Broadcast Joins
  • Structured Streaming
  • Spark Performance Optimization
  • Data Lake Architecture

Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
  • Excellent communication and collaboration skills.
  • Ability to work effectively in a fast-paced and dynamic environment

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs