Description
You will design and build agentic AI systems, including multi-agent orchestration and reasoning workflows, while managing MCP server and model-serving infrastructure. You will develop LLM-powered applications such as RAG pipelines and autonomous agents, and build scalable ML pipelines for training, inference, and monitoring. Your work will involve deploying solutions on AWS, optimizing for performance and cost, and implementing evaluation frameworks for accuracy and safety.
This role is on-site.
Responsibilities
- Design and develop agentic AI systems, multi-agent orchestration, and tool integrations.
- Build and manage MCP servers and model-serving infrastructure for enterprise use.
- Develop LLM-powered applications including RAG pipelines, copilots, and autonomous agents.
- Build scalable ML pipelines in Python for training, inference, and monitoring.
- Deploy and manage solutions on AWS (SageMaker, Bedrock, EKS/ECS, EC2, S3, Lambda) and optimize for latency and cost.
Required Skills
- 10–15 years of experience in AI/ML engineering or data science.
- Strong Python expertise for core programming, ML/AI development, and data processing.
- Hands-on experience with LLMs (OpenAI or open-source), prompt engineering, and fine-tuning.
- Proven experience building agentic AI and multi-agent systems.
- Experience with MCP servers and model-serving frameworks.
- Strong experience with RAG pipelines and vector databases.
- Hands-on experience with AWS platform: SageMaker, Bedrock, EKS/ECS, EC2, S3, Lambda, IAM.
- Experience with PyTorch or TensorFlow.
- Solid understanding of distributed systems and system design.
Preferred Skills
- Experience with LangChain, AutoGen, CrewAI, or Semantic Kernel.
- Experience with MLOps tools such as MLflow, Kubeflow, or SageMaker Pipelines.
- Knowledge of knowledge graphs, semantic search, or real-time streaming systems.