Description
You build and optimize scalable data processing applications using Python and PySpark.
Responsibilities
- Develop and maintain data processing applications.
- Design data solutions meeting performance and reliability standards.
- Monitor and troubleshoot pipeline performance issues.
- Implement CI/CD pipelines for automated testing and deployment.
Required Skills
- 8+ years of Python development experience.
- Expertise in PySpark, Apache Spark SQL, DataFrames, and Spark Streaming.
- Proficiency in Pandas and NumPy.
- Strong SQL skills and relational database experience.
- Hands-on experience with CI/CD pipelines.
- Version control with Git.
Preferred Skills
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.