You will engineer and optimize data pipelines using Azure services.
Responsibilities
- Develop solutions in PySpark and Python within the existing object-oriented and functional coding framework, specifically for Subrogation.
- Refactor legacy Cursor implementation code into PySpark solutions leveraging distributed computing.
- Conduct performance testing and optimization of data processes alongside architects.
- Design and implement data persistence models on ADLS, running tests to optimize Read and Write performance.
- Work hands-on with DataBricks, PySpark, Python, and ADF to deliver data solutions.
Required Skills
- 5+ years of professional experience in data engineering.
- Proficiency with Databricks.
- Strong command of PySpark and Python.
- Hands-on experience with Azure Data Factory (ADF).
- Experience building solutions using object-oriented and functional programming models.
- Ability to optimize data persistence layers on ADLS for I/O efficiency.
- Experience with distributed computing paradigms.