← Back to jobs

Prophecy Technologies Logo
Data Engineer

Prophecy Technologies

 

Irvine, CA, USA

Posted On: 4 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Data System Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will build and maintain data pipelines while migrating workloads from on-premise systems to GCP.

This role is on-site.

Responsibilities

  • Develop Hive queries using joins and partitions to process large datasets and load filtered data to edge node tables.
  • Build Spark jobs using Scala and Python (PySpark) APIs and use Spark SQL to create structured data.
  • Write shell scripts to schedule full and incremental loads and automate metadata synchronization between GCS and GCP Hive.
  • Migrate data from GCP Google Cloud Storage Hive to BigQuery and flatten on-prem data for ingestion into Druid.
  • Troubleshoot performance issues, tune Spark application parallelism, and resolve production connectivity or data quality issues.

Required Skills

  • 5+ years of experience in data engineering.
  • Proficiency in HiveQL for data analysis and large-scale data processing.
  • Hands-on experience developing Spark jobs using Scala and Python (PySpark).
  • Experience with Spark SQL and performance tuning for memory and parallelism.
  • Practical knowledge of BigQuery and Google Cloud Storage (GCS).
  • Ability to develop shell scripts for job scheduling and automation.
  • Experience managing data ingestion from on-premise sources to Druid.
  • Experience with Hive optimization techniques and data validation.

Preferred Skills

  • Master’s degree in Computer Science, Computer Engineering, Data Analytics, or an MBA.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs