Description
You will build and maintain data pipelines for production processing and ML training flows.
Responsibilities
- Ingest, structure, and analyze diverse unstructured data sources.
- Design and orchestrate data pipelines within an AWS environment.
- Evaluate, test, and improve the quality, privacy, and performance of data systems.
- Contribute to API and systems architecture alongside ML processing workflows.
Required Skills
- 3+ years of experience ingesting and structuring varied data sources.
- Production experience building and maintaining data pipelines.
- Strong proficiency in Python, SQL, and Pandas.
- Extensive experience with AWS, containers, and data orchestration.
- Machine Learning Engineering experience using PyTorch and scikit-learn.
- LLM expertise including LangChain, agents, and prompt engineering.
- Full stack development experience, specifically with JS/TS/Node.
- Significant experience handling healthcare data.
Preferred Skills