Description
You will own the design, development, and maintenance of large-scale data processing pipelines using Apache Spark.
Responsibilities
- Build and maintain efficient Spark code to process, transform, and analyze large datasets.
- Optimize Spark jobs for performance, scalability, and resource utilization.
- Integrate Hadoop, Hive, Spring, Hibernate, Kafka, and ETL processes into Spark applications.
- Monitor and manage Spark clusters to ensure high availability and reliability.
- Implement data quality and validation processes to ensure accuracy and consistency.
Required Skills
- 10+ years of experience as a Spark Developer or in a similar big data role.
- Strong proficiency in Apache Spark, including Spark SQL, Spark Streaming, and Spark MLlib.
- Proficiency in Scala or Python for Spark development.
- Experience with Extract, Transform, Load (ETL) processes and data modeling.
- Knowledge of data warehousing and distributed computing principles.
- Experience with cluster management and large datasets.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
Preferred Skills
- Ability to provide technical guidance and mentorship to junior developers.