Description
You will own data processing workflows and cluster management within AWS EMR environments.
This role is hybrid.
Responsibilities
- Develop and maintain data processing workflows using Spark and PySpark.
- Write and optimize Python scripts for core data engineering tasks.
- Automate operational processes and job orchestration via shell scripting.
- Manage and troubleshoot data workloads within AWS EMR clusters.
Required Skills
- 5+ years of professional experience in data engineering.
- Proficiency with Apache Spark and PySpark.
- Strong Python programming skills for data manipulation.
- Experience with Shell Scripting for automation.
- Hands-on experience managing AWS EMR clusters.
- Bachelor's degree or equivalent graduate qualification.