Description
Key Skills: PySpark, AWS EMR, Airflow, SQL, Data Warehousing, ETL, Spark, Python, AWS
Good to Have Skills: Java, Scala, Terraform, Data Lake & Lake Formation, Apache Iceberg, CI/CD, GitHub/GitLab, pair programming, testing, clean code, agile software development frameworks
Roles & Responsibilities:
- Become an essential part of the Data Refinement team, which provides data integration services for most downstream Media Measurement departments
- Design, develop, test and deploy big data applications and pipelines with attention to scalability, efficiency and maintainability
- Collaborate in a distributed, cross-functional agile team of Data Engineers, Data Scientists and Cloud Engineers to reach shared goals
- Take ownership of technical implementations and architecture together with your team and in alignment with communities of practices
- Communicate with data producers and consumers to align on data exchange formats and interfaces for better integration
- Provide suggestions to business in the Media Measurement context on how existing and new data sources can be utilized
- Recommend different ways to constantly improve data reliability and quality across all data processing pipelines
- Improve ways of working with integrating AI based workflows in the engineering domain for enhanced automation
- Take ownership of the components, pipelines and the overall system so that you contribute to improve things over time
Experience Required: 6+ years of relevant experience in Data Engineering using Spark with solid hands-on experience in Python, Java or Scala