Description
You will design, train, and deploy LLMs and Generative AI models, ensuring they meet enterprise standards for latency, throughput, and cost efficiency.
Responsibilities
- Build and maintain Retrieval-Augmented Generation (RAG) pipelines to improve contextual accuracy and response quality.
- Convert AI models into scalable APIs and microservices using FastAPI, Flask, or Spring Boot.
- Optimize model performance for production environments, focusing on reliability and cost efficiency.
- Develop reusable ML code, data pipelines, and feature engineering frameworks for training and inference.
- Collaborate with data engineers and product teams to integrate AI solutions into enterprise workflows.
Required Skills
- 3+ years of hands-on experience developing and deploying machine learning models.
- 1+ year of focused experience with LLMs, Generative AI, or agentic AI solutions.
- Expert-level Python (3.8+) for model development and scripting.
- Proficient Java for enterprise system integration.
- Strong experience with Google Cloud Platform, specifically BigQuery and Vertex AI.
- Proven track record of deploying RAG architectures in production.
- Deep learning frameworks: TensorFlow, PyTorch, or Keras.
- Experience with MLOps tools like MLFlow for experiment tracking and model registry.
- Familiarity with SQL and NoSQL databases.
Preferred Skills
- Containerization and orchestration using Docker and Kubernetes.
- Vector databases such as FAISS, Pinecone, or Chroma.
- GCP or AI/ML certifications (e.g., Google Cloud Professional Data Engineer).