← Back to jobs

Varite Inc Logo
Big Data Developer

Varite Inc

 

O'Fallon, MO, USA

Posted On: Just posted
Experience: 8+ years
Availability: Hybrid
Openings: 1
Category: Big Data Developer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will design, build, and maintain large-scale Spark applications and streaming data pipelines in a production environment.

This role is on-site.

Responsibilities

  • Develop and optimize Spark applications using Scala and PySpark for batch and real-time processing.
  • Build streaming pipelines with Kafka and Spark Structured Streaming, implementing windowing, watermarking, and checkpointing.
  • Create ingestion and routing flows using Apache NiFi and manage Kafka offsets for event replay.
  • Integrate workloads with distributed object storage systems like Apache Ozone and Ceph.
  • Optimize job performance through partitioning, memory tuning, and shuffle optimization while ensuring data quality.

Required Skills

  • 8+ years of experience with Big Data technologies.
  • Advanced proficiency in Scala and Python (PySpark).
  • Strong experience with Apache Spark Core, SQL, and Structured Streaming in production.
  • Solid understanding of Kafka-based streaming architectures and event-driven systems.
  • Hands-on experience with Apache NiFi for data ingestion and flow management.
  • Proficiency in SQL for structured and semi-structured data.
  • Experience with Linux, shell scripting, and Git-based version control.
  • Familiarity with CI/CD pipelines and monitoring tools.

Preferred Skills

  • Experience with Apache Ozone and/or Ceph as storage backends.
  • Background in Spark performance tuning (CPU, memory, I/O, shuffle).
  • Experience supporting mission-critical systems with strict SLAs.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs