← Back to jobs

Infosys Limited Logo
PySpark Developer

Infosys Limited

 

Bengaluru East, Karnataka, India

Posted On: 2 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Pyspark Developer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will build and maintain scalable data pipelines and ETL workflows using PySpark in distributed environments.

This role is on-site.

Responsibilities

  • Develop and maintain data pipelines using PySpark for large-scale datasets.
  • Design and implement ETL/ELT workflows and optimize Spark jobs for performance.
  • Ensure data quality, integrity, and governance while troubleshooting processing issues.
  • Automate workflows using scheduling tools like Airflow.
  • Collaborate with data engineers and analysts to deliver clean, efficient code.

Required Skills

  • 5+ years of experience with Python and PySpark.
  • Strong proficiency in Apache Spark (RDDs, DataFrames, Spark SQL).
  • Experience in ETL pipeline development and SQL database concepts.
  • Familiarity with data formats (Parquet, ORC, JSON, CSV).
  • Knowledge of Hadoop ecosystem components (HDFS, Hive).
  • Experience with AWS, Azure, or GCP cloud platforms.
  • Understanding of distributed computing concepts and Git.

Preferred Skills

  • Experience with Databricks, EMR, or Lakehouse architecture (Delta Lake).
  • Familiarity with workflow orchestration (Airflow) and CI/CD pipelines.
  • Exposure to Kafka, real-time streaming, or NoSQL databases (MongoDB, Cassandra).

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs