← Back to jobs
Santa Clara, CA, USA
No related jobs found
Must-Have Skills
✅ Python
✅ LLMOps & MLOps
✅ AI/ML Platform Engineering
✅ LLM Inferencing & Model Hosting
✅ vLLM, SGLang, Triton, TGI, Ray Serve
✅ Kubernetes & Docker
✅ Azure ML & Databricks Model Serving
✅ MLflow
✅ Open-Source LLMs – Llama, Mistral, Gemma, Qwen
✅ RAG & Vector Databases
✅ GPU Optimization
✅ Quantization, KV Cache & PagedAttention
✅ Continuous & Dynamic Batching
🔹 Key Responsibilities
✔ Design, deploy, and optimize scalable LLM inference and model-serving platforms.
✔ Improve GPU utilization, latency, throughput, memory efficiency, and inference costs.
✔ Build and maintain production AI infrastructure using Kubernetes and Docker.
✔ Deploy, monitor, troubleshoot, and scale enterprise GenAI and LLM applications.
✔ Implement LLMOps/MLOps best practices, AI observability, governance, and responsible AI standards.
⭐ Preferred Skills
⭐ LLM Fine-Tuning – PEFT, LoRA, QLoRA
⭐ Azure AI Foundry & Azure OpenAI
⭐ Hugging Face & DeepSpeed
⭐ Distributed Training & Multi-GPU Environments
⭐ LangGraph, AutoGen & CrewAI
⭐ AI Observability & Governance
Bachelor's degree
No related jobs found
← Back to jobs