Description
You will lead the development and deployment of generative AI systems, focusing on model fine-tuning, RAG architecture, and scalable infrastructure.
This role is hybrid.
Responsibilities
- Design and implement Retrieval-Augmented Generation (RAG) architectures, integrating diverse data sources and optimizing query processing.
- Develop and fine-tune large language models and image generation models, applying quantization, pruning, and knowledge distillation for optimization.
- Build scalable AI systems, including APIs, microservices, and data pipelines, collaborating with MLOps for production deployment.
- Provide technical guidance on generative AI strategies and mentor junior engineers to ensure adherence to best practices.
- Conduct experiments to improve model performance and stay current with AI research advancements.
Required Skills
- 7+ years of overall experience in AI/ML, with specific focus on generative AI and RAG architectures.
- Strong programming proficiency in Python.
- Hands-on experience with deep learning frameworks such as TensorFlow and PyTorch.
- Practical experience with GenAI technologies including OpenAI, Anthropic, or Llama.
- Experience with cloud platforms (AWS, Azure, GCP) and their AI/ML services.
- Demonstrable skills in Natural Language Processing (NLP) and computer vision.
- Bachelor's degree in Computer Science, Computer Engineering, or a related field.
Preferred Skills
- Experience building and managing production-grade AI microservices.
- Proven track record of mentoring engineering teams and driving technical strategy.