← Back to jobs

Collabera Logo
Hadoop Data Engineer

Collabera

 

Chicago, IL, USA

Posted On: 1 day ago
Experience: 5+ years
Availability: Onsite
Openings: 2
Category: Hadoop Data Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will design, develop, and maintain batch and near-real-time data pipelines for analytics.

This role is on-site.

Responsibilities

  • Ingest data from Kafka, file shares, REST APIs, and relational databases.
  • Transform, clean, and validate data using HDFS, Hive, Impala, or Spark SQL.
  • Manage data formats including JSON, CSV, and XML for downstream consumption.
  • Perform data profiling and validation to ensure integrity and identify anomalies.
  • Troubleshoot failures and optimize performance in SQL jobs and Spark applications.

Required Skills

  • 5+ years of experience in data engineering.
  • Strong SQL proficiency with MySQL, Hive, Impala, and Spark SQL.
  • Hands-on programming experience in Scala or Python.
  • Experience with Spark Structured Streaming and distributed systems.
  • Working knowledge of Hadoop, Sqoop, and MapReduce.
  • Proficiency with data ingestion patterns and formats (JSON, CSV, XML).

Preferred Skills

  • Experience with HDFS and Hive ecosystem tools.
  • Background in designing near-real-time data pipelines.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs