Description
You will design and operate ML training and deployment pipelines for LLM applications on Azure and SUSE Linux GPU clusters.
This role is on-site.
Responsibilities
- Manage SUSE Linux Enterprise (SLES) GPU clusters with NVIDIA H100 hardware, including driver installation and CUDA/NCCL tuning.
- Deploy GPU-based inference endpoints using Managed Online Endpoints, AKS GPU node pools, or Arc-enabled Kubernetes, handling traffic splits and rollbacks.
- Automate CI/CD pipelines in Azure DevOps from data preparation to model deployment using Infrastructure as Code (Terraform/Bicep).
- Integrate MLflow, Azure ML model registry, and Model Catalog for unified model versioning and promotion.
- Monitor model performance, data drift, and GPU metrics via Azure Monitor, Log Analytics, and NVIDIA DCGM Exporter.
Required Skills
- 7+ years in ML/AI engineering or MLOps with significant GPU workload experience.
- Hands-on experience with SUSE Linux (SLES) in production AI environments.
- In-depth knowledge of NVIDIA H100 architecture (HBM3, NVLink, MIG, multi-GPU orchestration).
- Proficiency in Azure ML, Azure AI Foundry, and Prompt Flow for LLM workflows.
- Expertise in deploying on Kubernetes with GPU node support (AKS, Arc-enabled K8s).
- Experience with Infrastructure as Code, specifically Terraform and Bicep.
- Familiarity with CI/CD workflows using Azure DevOps Pipelines.
- Knowledge of distributed training frameworks (DeepSpeed, Horovod, PyTorch DDP).
- Experience implementing governance and security for ML platforms.