Description
- We are seeking an experienced Data Engineer to design, build, and maintain scalable, high-performance data pipelines and platforms that support enterprise analytics, reporting, and AI/ML initiatives. The ideal candidate should have strong expertise in Python, SQL, Spark/PySpark, cloud platforms, ETL/ELT, data warehousing, and modern data engineering practices. Experience working in Agile teams and the financial services domain is preferred.
Must-Have Technical Skills
- Python
- SQL
- PySpark / Apache Spark
- ETL/ELT Development
- Azure Data Factory (ADF) or AWS Glue
- Azure Databricks or Databricks
- Azure Data Lake Storage (ADLS) / Amazon S3
- Azure Synapse Analytics, Snowflake, or Redshift
- Data Modeling
- Git
- CI/CD
- Docker (preferred)
- Roles & Responsibilities
- Design, develop, and maintain scalable data pipelines for batch and real-time processing.
- Build robust ETL/ELT workflows using Python, SQL, and Spark.
- Develop and optimize data ingestion processes from multiple structured and unstructured data sources.
- Implement scalable data models to support business intelligence and analytics.
- Work with large datasets using Databricks, Spark, and cloud-native data platforms.
- Optimize data processing performance, reliability, and cost efficiency.
- Ensure data quality, governance, security, and compliance standards are met.
- Collaborate with data scientists, analysts, application developers, and business stakeholders.
- Develop automated testing, monitoring, and deployment processes for data pipelines.
- Troubleshoot production issues and continuously improve data platform performance.
- Participate in Agile ceremonies, code reviews, and technical design discussions.
Required Qualifications
- 6 10 years of experience in Data Engineering.
- Strong programming experience in Python and SQL.
- Hands-on experience with Apache Spark/PySpark.
- Experience with Databricks and cloud-based data engineering solutions.
- Strong understanding of ETL/ELT design patterns.
- Experience with Azure or AWS cloud services.
- Knowledge of relational databases and data warehousing concepts.
- Experience with version control (Git) and CI/CD pipelines.
- Excellent analytical and problem-solving skills