Description
You will design, build, and maintain scalable batch ETL pipelines for data ingestion, transformation, and loading.
This role is on-site.
Responsibilities
- Develop modular, testable, and reusable PySpark pipelines in Python for large-scale data processing.
- Optimize Spark job performance through partitioning, caching, join strategies, and shuffle management.
- Implement data quality checks, handle schema evolution, and ensure dataset accuracy and completeness.
- Troubleshoot pipeline failures, analyze logs, and build robust error handling and retry mechanisms.
- Collaborate with cross-functional teams to define data mappings and document pipeline logic and dependencies.
Required Skills
- 2–5 years of hands-on experience building data pipelines using Python and PySpark.
- Strong understanding of ETL concepts, data transformations, and handling large-scale datasets.
- Proficiency in writing clean, maintainable, and debuggable production code.
- Working knowledge of data structures, algorithms, and software development best practices.
- Experience with performance tuning for distributed computing frameworks.
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or a related field (or equivalent practical experience).
Preferred Skills
- Experience with operational stability tools, including alerting systems and runbooks.
- Familiarity with continuous improvement of engineering standards and tooling.