You will build and maintain data pipelines to transform raw data into actionable information for data science and product teams.
Responsibilities
Extract data from Hadoop databases using Sqoop and Linux systems.
Stage real-time data from gateways into AWS S3 or Azure Blob storage.
Implement Spark using Scala, utilizing DataFrames and Spark SQL API for high-speed processing.
Optimize Spark session performance through effective use of partitions.
Develop ETL pipelines in ADF using Linked Services, Datasets, and Pipelines to move data between Azure SQL, Blob storage, and Azure SQL Data Warehouse.
Build Spark streaming pipelines using Java.
Required Skills
5+ years of experience in big data engineering.
Proficiency in Scala and Java.
Hands-on experience with Spark and Spark SQL.
Strong SQL skills for data analysis and transformation.
Experience with Azure SQL and Azure SQL Data Warehouse.
Experience managing data in Blob storage.
Knowledge of Hadoop databases and Sqoop.
Master's degree in Computer Science, IT, IS, Engineering, or a related field.
Ability to travel or relocate to unanticipated client sites if required.