← Back to jobs

Foray Software Private Limited Logo
Data Engineer
Posted On: 5 days ago
Experience: Not specified years
Availability: Onsite
Openings: 1
Category: Data Engineer / ETL Developer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will design and maintain large-scale data processing systems and ETL pipelines.

This role is on-site.

Responsibilities

  • Build PySpark applications using Spark DataFrames within Jupyter Notebook and PyCharm.
  • Develop ETL processes to ingest, copy, and structurally transform data across CSV, TSV, XML, JSON, and fixed-width formats.
  • Optimize Spark jobs to handle large-scale data volumes efficiently.
  • Deploy and manage applications within automated release pipelines using Jenkins.
  • Debug complex data workflows and processing errors.

Required Skills

  • 9+ years of overall IT experience with a heavy focus on Big Data technologies.
  • Hands-on expertise in Python and PySpark.
  • Experience with AWS Analytics services including Amazon EMR, Athena, and AWS Glue.
  • Proficiency with AWS Compute and Storage services including Lambda, EC2, S3, and SNS.
  • Experience with version control using Git.
  • Knowledge of columnar storage formats such as Parquet, Avro, and ORC.
  • Experience with compression techniques like Snappy and Gzip.
  • Familiarity with DevOps concepts and automated deployment environments.

Preferred Skills

  • Experience with Bash/Shell scripting.
  • Knowledge of data warehousing concepts including dimensions, facts, snowflake, and star schemas.
  • Experience with AWS databases such as Aurora, RDS, Redshift, ElastiCache, or DynamoDB.

Key Skills
Education

Not specified

Related Jobs

No related jobs found

← Back to jobs