Description

You will design, develop, and implement data models for enterprise-level applications while managing end-to-end ETL pipelines.

Responsibilities

  • Design and deploy serverless data pipelines using AWS Lambda and Glue Catalog.
  • Execute cloud migrations from on-premises environments to AWS and Azure.
  • Build and manage reporting and analytics infrastructure for internal business clients.
  • Establish and maintain multi-node Hadoop clusters and optimize algorithms using Spark.
  • Connect Azure to on-premises data centers via Azure Express Route.

Required Skills

  • 8+ years of experience in Data Engineering, ETL Development, or Software Engineering.
  • Extensive AWS expertise including S3, EMR, Redshift, DynamoDB, Athena, Glue, Kinesis, Lambda, and EC2.
  • Proficiency in Azure Cloud services including Data Factory, Data Lake Storage, Synapse Analytics, and Azure SQL.
  • Advanced Python programming with Object-Oriented principles and libraries like NumPy, Pandas, SciPy, and Matplotlib.
  • Experience with Big Data technologies including Apache Spark (Spark-SQL, DataFrames, RDD) and Hadoop (HDFS, MapReduce).
  • Hands-on experience with Snowflake and SnowSQL.
  • Infrastructure as Code and CI/CD experience using Terraform, Chef, Jenkins, GitHub, and Docker.
  • Ability to develop JSON-based RESTful and XML-based SOAP web services.

Preferred Skills

  • Experience with Databricks.

Education

Any graduate