← Back to jobs

EZTEK Associates Logo
PySpark Developer

EZTEK Associates

 

Philadelphia, PA, USA

Posted On: 3 days ago
Experience: 4+ years
Availability: Onsite
Openings: 1
Category: PySpark Developer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will design and build scalable ETL/ELT pipelines using PySpark to ingest data from databases, logs, APIs, and files.

This role is hybrid.

Responsibilities

  • Build reusable, parameterized Spark jobs for batch and micro-batch processing.
  • Transform raw transactional and log data into analysis-ready datasets within the Data Hub and analytical data marts.
  • Optimize PySpark job performance and scalability across large data volumes.
  • Ensure data quality, consistency, and lineage across all ingestion flows.
  • Collaborate with Data Architects and Data Scientists to implement ingestion logic.

Required Skills

  • 4+ years of experience in data engineering with a focus on PySpark and Spark.
  • Proficiency in Python and data processing libraries.
  • Expertise in building pipelines from relational, semi-structured (JSON, XML), and unstructured sources.
  • Strong SQL skills for querying and validating data in Amazon Redshift or PostgreSQL.
  • Experience with distributed computing frameworks like Spark on EMR or Databricks.
  • Knowledge of data lake and data warehouse architectures.
  • Experience with workflow orchestration tools such as AWS Step Functions.
  • Ability to work with AWS S3, Glue, EMR, and Redshift.

Preferred Skills

  • Experience with Spark Structured Streaming or Kafka.
  • Familiarity with Delta Lake for large-scale data storage.
  • Knowledge of DevOps/CI-CD pipelines using Git or Jenkins.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs