Description
You will manage and optimize distributed systems for large-scale deep learning workloads.
Responsibilities
- Design, deploy, and maintain distributed systems using Kubernetes and Slurm to optimize resource utilization.
- Configure and optimize Multi-GPU and Multi-Node Deep Learning job scheduling.
- Develop complex shell scripts to automate system tasks and reduce manual intervention.
- Monitor system performance and troubleshoot issues within distributed systems and job scheduling processes.
- Collaborate with cross-functional teams to translate project requirements into technical solutions.
Required Skills
- 3-6 years of experience in distributed systems and MLOps.
- Hands-on experience with MLOps tools including MLFlow, Kubeflow, and AutoML.
- Proficiency in Python and Linux systems, specifically Ubuntu.
- Proven experience managing Kubernetes and Slurm environments.
- Experience working with on-prem NVIDIA GPU servers.
- Strong understanding of AI model training, deployment, and Multi-GPU/Multi-Node scheduling.
- Experience with shell scripting for system automation and orchestration.
- Knowledge of logical networks and CI/CD processes.
Preferred Skills
- Understanding of PyTorch, TensorFlow, NLP, or Computer Vision.
- Experience with the NVIDIA ecosystem, including Triton Inference Server and CUDA.