Description
You will build and maintain large-scale data infrastructure.
Responsibilities
- Design and implement scalable ETL pipelines using PySpark, SQL, and Delta Lake on Databricks.
- Develop and optimize calculation engines that integrate with multiple source systems.
- Tune Spark jobs to handle datasets in the billions of records, utilizing partitioning and caching.
- Ensure data quality, lineage, and reconciliation across all data layers.
- Implement parameterized financial logic, including APR and compounding rules, based on business requirements.
Required Skills
- 15+ years of experience in Data Engineering, preferably within financial services.
- Strong hands-on experience with PySpark and SQL.
- Expertise with Databricks, including Delta Live Tables and Unity Catalog.
- Solid understanding of data modeling, partitioning strategies, and query optimization.
- Experience implementing CI/CD using tools like GitHub or Jenkins.
- Familiarity with workflow orchestrators such as Airflow or Dagster.
- Experience handling credit card or banking domain data.