← Back to jobs

The X4 Group Logo
Senior MLOps Engineer
Posted On: 1 day ago
Experience: 7+ years
Availability: Hybrid
Openings: 1
Category: MLOps Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will design and operate ML training and deployment pipelines for LLM applications on Azure and SUSE Linux GPU clusters.

This role is on-site.

Responsibilities

  • Manage SUSE Linux Enterprise (SLES) GPU clusters with NVIDIA H100 hardware, including driver installation and CUDA/NCCL tuning.
  • Deploy GPU-based inference endpoints using Managed Online Endpoints, AKS GPU node pools, or Arc-enabled Kubernetes, handling traffic splits and rollbacks.
  • Automate CI/CD pipelines in Azure DevOps from data preparation to model deployment using Infrastructure as Code (Terraform/Bicep).
  • Integrate MLflow, Azure ML model registry, and Model Catalog for unified model versioning and promotion.
  • Monitor model performance, data drift, and GPU metrics via Azure Monitor, Log Analytics, and NVIDIA DCGM Exporter.

Required Skills

  • 7+ years in ML/AI engineering or MLOps with significant GPU workload experience.
  • Hands-on experience with SUSE Linux (SLES) in production AI environments.
  • In-depth knowledge of NVIDIA H100 architecture (HBM3, NVLink, MIG, multi-GPU orchestration).
  • Proficiency in Azure ML, Azure AI Foundry, and Prompt Flow for LLM workflows.
  • Expertise in deploying on Kubernetes with GPU node support (AKS, Arc-enabled K8s).
  • Experience with Infrastructure as Code, specifically Terraform and Bicep.
  • Familiarity with CI/CD workflows using Azure DevOps Pipelines.
  • Knowledge of distributed training frameworks (DeepSpeed, Horovod, PyTorch DDP).
  • Experience implementing governance and security for ML platforms.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs