← Back to jobs
Dallas, TX, USA
No related jobs found
Responsibilities
A Lead Databricks Developer designs| builds| and optimizes big data ETLELT pipelines using PySpark| SQL| and Delta Lake within cloud environments (Azure| AWS| or GCP). Key Responsibilities include creating scalable data solutions| optimizing Spark jobs| and managing data workflows to support analytics. Required skills include proficiency in Python SQL| Spark| and experience with data warehousing concepts. Pipeline Development Design| develop| and maintain robust| scalable ETLELT data pipelines using Databricks and PySpark.Data Processing Work with structured and unstructured data| utilizing Delta Lake for efficient storage and management. Optimization Perform performance tuning and optimization on Databricks clusters and Spark jobs for cost-efficiency. Orchestration Use tools like Azure Data Factory (ADF) or Airflow for workflow orchestration and scheduling. Integration Integrate Databricks with cloud data sources (Azure Data Lake Gen2| S3) and BI tools. Data Quality Security Implement data quality checks| validation| and security protocols. Required Skills Qualifications Technical Skills Expert-level knowledge of PySpark| Apache Spark| and SQL. Cloud Platforms Hands-on experience with Azure Databricks (most common)| AWS Databricks| or GCP. Data Engineering Strong understanding of Data Modeling| ETLELT patterns| and Data Warehousing concepts. Tools Languages Python| Azure Data Factory (ADF)| Delta Lake| GitAzure DevOps. Experience 6-8 years of experience in data engineering or software development
Bachelor's degree
No related jobs found
← Back to jobs