← Back to jobs

Cybrain Software Logo
Python with Spark
Posted On: Just posted
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Python with Spark
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will build and optimize large-scale data processing pipelines using Python and Spark.

This role is remote.

Responsibilities

  • Develop data processing logic using Spark DataFrames, RDDs, and Spark APIs.
  • Tune Spark queries and implement performance optimization techniques.
  • Build streaming applications using Spark Streaming and MLlib.
  • Manage data across various formats including Parquet, JSON, and CSV.
  • Work within the Hadoop ecosystem and cloud-based data platforms.

Required Skills

  • 5 to 10 years of experience in data engineering.
  • Proficiency in Python and the Spark Framework.
  • Deep knowledge of Spark query tuning and performance optimization.
  • Experience with MLlib and Spark Streaming.
  • Working knowledge of the Hadoop Ecosystem (HDFS, Hive, Impala).
  • Strong SQL skills and experience with NoSQL databases.
  • Hands-on experience with Redshift, Snowflake, or MongoDB.
  • Familiarity with Synapse Analytics on Azure, DynamoDB, or Databricks.
  • Experience with Cloud platforms including AWS and Azure.

Preferred Skills

  • Experience managing data in Parquet, JSON, and CSV formats.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs