You will design and build a Model-as-a-Service platform enabling non-experts to construct AI solutions via drag-and-drop components.
Responsibilities
Architect and optimize end-to-end Retrieval-Augmented Generation (RAG) pipelines, including advanced chunking strategies and vector database management.
Develop and maintain MCP (Model Control Protocol) libraries, clients, and servers to integrate diverse data sources with the AI engine.
Build and curate a repository for Agentic AI, allowing users to select or construct custom agents for specialized tasks.
Manage and optimize one of the largest on-premise GPU farms in the U.S. banking sector, comprising 500+ Nvidia nodes.
Integrate AI deployment pipelines with enterprise CI/CD tools like Jenkins and Ansible while implementing corporate guardrails within Model Risk Management (MRM) frameworks.
Required Skills
Expert-level Python with deep, hands-on production experience.
10-15+ years of experience building production-grade platforms for developers or business units.
Extensive data engineering background in massive data ingestion and processing.
Deep understanding of vector databases, inferencing, and advanced RAG chunking strategies.
Experience mimicking cloud capabilities (AWS/Azure) within strictly on-premise environments.
Proficiency with DevOps tools, specifically Jenkins and Ansible for automated deployment.
Proven ability to translate business use cases into scalable platform services.
Preferred Skills
Experience working with large-scale GPU farms and high-volume data environments.
Current knowledge of agentic frameworks and rag-less inferencing techniques.