5+ years of data engineering (or software engineering with a strong data focus), with strong SQL.
Expertise in Python (Java or Scala a plus) and technologies such as Airflow, Spark, Trino/Dremio, Iceberg, Kafka, Docker.
Hands-on experience designing and maintaining custom ETL / data pipelines and warehouse solutions.
Proven ability to independently troubleshoot and root-cause production data issues — across a shared pipeline estate, not only pipelines you personally built — driving problems to their true (often upstream) cause and a durable fix, not only executing prescribed steps.
Strong ownership and operational discipline: rigorous separation of development and production environments, careful low-rework changes, and consistent follow-through on issues you find or create.
Demonstrated ownership of data quality — designing validation checks and performing root-cause analysis on data discrepancies.
Experience operating pipelines in production: incident response, backfills/reprocessing, deployment/release activities.
Ability to work beyond narrowly-scoped tasks — take an ambiguous or new problem and carry it to completion with limited oversight.
Familiarity with SDLC best practices, version control (Git), and CI/CD.
Excellent oral and written communication; able to produce clear, structured operational communication (change plans, RCAs, runbooks) and work across cross-functional teams