← Back to jobs

Diverse Lynx Logo
PySpark Developer

Diverse Lynx

 

Dallas, TX, USA

Posted On: 15+ days ago
Experience: 7+ years
Availability: Onsite
Openings: 1
Category: PySpark Developer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will build and test scalable PySpark data processing solutions and validation systems.

This role is on-site.

Responsibilities

  • Design scalable PySpark test architectures and modular frameworks for batch processing ETL/data pipelines.
  • Architect end-to-end data validation systems in Hadoop environments, managing lineage and schema evolution.
  • Lead system design for Hadoop/Hive test environments, including YARN resource management and dynamic partitioning.
  • Design CI/CD test pipelines for PySpark/Hadoop jobs, incorporating artifact management, parallel execution, and blue-green deployments.
  • Mentor junior engineers on PySpark testing basics and contribute to overall testing strategy discussions.

Required Skills

  • 7+ years of experience in data and ETL testing with hands-on PySpark.
  • Proficiency with Hadoop and Hive environments.
  • Experience designing data quality systems using PySpark integrated with Hive metadata services.
  • Familiarity with Zephyr, Jira, and ServiceNow integrated test management systems.
  • Experience with API-driven automation within testing frameworks.
  • Ability to design testing platforms and test data generators.

Preferred Skills

  • Experience with dynamic partitioning strategies in Hadoop/Hive.
  • Background in mentoring technical teams on testing best practices.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs