← Back to jobs

Lorven Technologies Logo
Data Engineer with AI

Lorven Technologies

 

Boston, MA, USA

Posted On: 15+ days ago
Experience: 6+ years
Availability: Remote
Openings: 1
Category: AI Data Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will build and scale lakehouse and AI data pipelines on Databricks.

This role is remote.

Responsibilities

  • Design, build, and maintain batch/streaming pipelines in Python and PySpark on Databricks using Delta Lake, Autoloader, and Structured Streaming.
  • Implement Bronze/Silver/Gold data models, optimizing performance with partitioning, Z-ORDER, and indexing.
  • Enable ML/AI workflows including feature engineering, MLflow experiment tracking, and supporting RAG pipelines with embeddings and vector stores.
  • Establish data quality checks using Great Expectations, lineage, and governance via Unity Catalog and RBAC.
  • Champion CI/CD and IaC practices while troubleshooting pipeline performance and cost issues.

Required Skills

  • 6+ years in data engineering with production pipeline experience.
  • Expert in Python and PySpark, including UDFs, Window functions, and Spark SQL.
  • Deep hands-on experience with Databricks: Delta Lake, Jobs/Workflows, and Structured Streaming.
  • Strong SQL and data modeling skills (dimensional, medallion, CDC).
  • Experience enabling ML/AI workflows using MLflow and feature stores.
  • Cloud proficiency on AWS, Azure, or GCP (object storage, IAM, networking).
  • Familiarity with CI/CD tools (GitHub/GitLab/Azure DevOps) and testing (pytest).
  • Experience with LLM workflows, including embeddings and vectorization.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs