You will own the testing strategy and execution for data pipelines.
Responsibilities
Own end-to-end test strategy for data pipelines including ingestion, transformation, and consumption across Dev/Test/Prod environments.
Design and implement automation frameworks in Python for ETL/ELT validation covering schema, data quality, business rules, transformations, aggregations, deduplication, and SCD logic.
Build reusable test libraries for Databricks notebooks/jobs, tables, and comprehensive validation processes.
Implement CI/CD-integrated tests including integration and regression testing for data workloads and pipelines.
Convert manual use cases to automated test cases and provide comprehensive end-to-end test coverage.
Required Skills
6-8 years in QA/Testing, with at least 3+ years focused on data/ETL testing in cloud environments.
Strong hands-on experience with Python for test automation and data manipulation.
Proficiency in SQL and ETL Testing principles.
Experience with Data Pipelines, Databricks, and PySpark.
Familiarity with Data Modeling and Cloud Computing concepts.