← Back to jobs

Momento USA Logo
AI Architect

Momento USA

 

Bellmawr, NJ, USA

Posted On: 30+ days ago
Experience: 8+ years
Availability: Remote
Openings: 1
Category: AI Architect
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Role Overview

We are seeking a senior AI Architect to lead the design, evaluation, and deployment of enterprise-grade Generative AI systems and autonomous Agentic AI workflows. In this role, you will bridge cutting-edge research and production systems, building scalable, secure multi-agent systems, Retrieval-Augmented Generation (RAG) pipelines, and LLM-driven applications.

 

Key Responsibilities

  • Architecture & Strategy: Design and implement scalable architectures for Agentic workflows (ReAct, Planning, Tool-Use, Reflection) and multi-agent coordination frameworks (e.g., AutoGen, CrewAI, LangGraph).
  • GenAI Pipeline Engineering: Build robust enterprise RAG systems, incorporating advanced retrieval methods (hybrid search, re-ranking, graph RAG, and vector databases).
  • Model Selection & Fine-Tuning: Evaluate, benchmark, and fine-tune open-source and proprietary foundation models (e.g., Llama, GPT, Claude, Mistral) using techniques like LoRA/QLoRA and prompt optimization.
  • Governance & Reliability: Implement guardrails, evaluation pipelines (RAGAS, TruLens), AI safety standards, cost-optimization strategies, and observability (e.g., LangSmith, Phoenix).
  • Cross-Functional Leadership: Partner with product managers, data engineering, and DevOps teams to transition AI prototypes into secure, low-latency microservices.

 

Qualifications & Technical Skills

  • Experience: 8+ years in software design and machine learning, with 2+ years focused specifically on Generative AI and Agentic frameworks.
  • Agentic Frameworks: Hands-on experience with LangChain, LangGraph, LlamaIndex, AutoGen, CrewAI, or custom tool-calling agents.
  • Vector Databases & Retrieval: Proficiency with vector search engines (Pinecone, Qdrant, Milvus, Weaviate, Pgvector) and semantic search architectures.
  • Programming & Cloud: Expertise in Python, asynchronous APIs (FastAPI), cloud services (AWS Bedrock, Azure OpenAI, GCP Vertex AI), and containerization (Docker, Kubernetes).
  • LLMOps & Monitoring: Demonstrated experience in tracking model latency, token costs, drift, and agent decision-making pathways

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs