You will design and build scalable data ingestion pipelines from enterprise systems.
This role is remote.
Responsibilities
Design and build scalable data ingestion pipelines from enterprise systems, including ERP, documentation tools, version control, and project management tools.
Develop connectors for structured, semi-structured, and unstructured data, implementing CDC and real-time sync.
Engineer knowledge graphs by designing ontologies and schemas for complex enterprise relationships.
Extract and model enterprise metadata, parsing and semantically indexing documents and code artifacts.
Prepare and contextualize data for LLM consumption, designing embedding strategies and managing vector databases.
Required Skills
5+ years of Data Engineering experience with production-grade pipelines.
Strong Python skills for writing clean, testable, maintainable code.
Expertise in MongoDB, including schema design, aggregation pipelines, indexing, and performance tuning.
Experience with Vector databases such as Qdrant, Pinecone, Weaviate, or pgvector.
Strong SQL skills, including complex queries, joins, window functions, and optimization.
Experience with ETL/ELT at scale, focusing on incremental loads, CDC, and idempotent pipelines.
Proficiency with pipeline orchestration tools like Airflow, Dagster, Prefect, or similar.
Document processing experience, including chunking, metadata extraction from PDFs/Word/HTML, and using LangChain or similar.
Preferred Skills
Proficiency with Graph Databases such as Neo4j or Jena.