You will design, develop, and implement data models for enterprise-level applications while managing end-to-end ETL pipelines.
Responsibilities
- Design and deploy serverless data pipelines using AWS Lambda and Glue Catalog.
- Execute cloud migrations from on-premises environments to AWS and Azure.
- Build and manage reporting and analytics infrastructure for internal business clients.
- Establish and maintain multi-node Hadoop clusters and optimize algorithms using Spark.
- Connect Azure to on-premises data centers via Azure Express Route.
Required Skills
- 8+ years of experience in Data Engineering, ETL Development, or Software Engineering.
- Extensive AWS expertise including S3, EMR, Redshift, DynamoDB, Athena, Glue, Kinesis, Lambda, and EC2.
- Proficiency in Azure Cloud services including Data Factory, Data Lake Storage, Synapse Analytics, and Azure SQL.
- Advanced Python programming with Object-Oriented principles and libraries like NumPy, Pandas, SciPy, and Matplotlib.
- Experience with Big Data technologies including Apache Spark (Spark-SQL, DataFrames, RDD) and Hadoop (HDFS, MapReduce).
- Hands-on experience with Snowflake and SnowSQL.
- Infrastructure as Code and CI/CD experience using Terraform, Chef, Jenkins, GitHub, and Docker.
- Ability to develop JSON-based RESTful and XML-based SOAP web services.
Preferred Skills
- Experience with Databricks.