Description
You will build and maintain data monitoring pipelines to identify and resolve quality issues before they impact downstream products.
Responsibilities
- Design and implement data monitoring pipelines to proactively resolve data quality issues.
- Collaborate with stakeholders to define requirements, develop quality metrics, and negotiate data quality SLAs.
- Lead technical efforts for data observability and federated query systems to enable data democratization.
- Develop new methodologies to improve access to trustworthy data and enhance pipeline performance.
Required Skills
- 3+ years of experience in Data Engineering or similar roles working with big data pipelines and analytics.
- 2+ years of hands-on experience with Apache Spark.
- 2+ years of professional coding experience in Python, Java, or an equivalent language.
- 2+ years of experience using SQL in scalable data warehouses such as BigQuery or Snowflake.
- Proficiency in cloud technologies, specifically GCP or AWS.
- Experience with Apache Airflow.
- Deep understanding of Distributed Systems and Effective Data Management.
- Proven ability to implement data engineering best practices and optimize pipeline scalability.
Preferred Skills
- Bachelor’s Degree in Computer Science, Engineering, or a related STEM field.