Description
You will build and optimize data pipelines and applications within the Azure ecosystem.
Responsibilities
- Develop and optimize PySpark-based ETL pipelines for large-scale data processing.
- Build, test, and maintain Python applications.
- Design and implement data solutions using Azure Data Lake, Azure Databricks, and Azure Data Factory.
- Implement data classification, policy enforcement, and metadata extraction within Purview.
- Troubleshoot Spark job performance bottlenecks to improve pipeline efficiency.
Required Skills
- 3+ years of experience in Python and PySpark for big data processing.
- Proven experience as a Python Developer or in a similar engineering role.
- Strong experience with Azure Data Services including Azure Data Lake and Azure Data Factory.
- Experience working with Azure Databricks.
- Ability to work within an agile environment.
- Degree in any graduate field.