← Back to jobs

NVIDIA Logo
Senior Machine Learning Engineer

NVIDIA

 

Yokneam, Israel

Posted On: 6 days ago
Experience: 12+ years
Availability: Onsite
Openings: 1
Category: Machine Learning Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will implement machine learning algorithms and models for large-scale deep learning training on NVIDIA supercomputers.

You will own the release pipeline.

Responsibilities

  • Implement deep learning models tailored for NVIDIA supercomputers with a focus on high-performance networking.
  • Profile, benchmark, and analyze models to identify optimizations for performance, efficiency, and accuracy.
  • Design and implement scalable training pipelines and frameworks leveraging high-performance networking.
  • Collaborate with software and hardware engineers to integrate networking solutions such as RoCE or InfiniBand.
  • Analyze large-scale training results to resolve networking bottlenecks and optimize model outcomes.

Required Skills

  • B.Sc. in Computer Science, Software Engineering, or 12+ years of equivalent experience.
  • Deep learning expertise with practical experience in TensorFlow and PyTorch.
  • Proficiency in CUDA programming for NVIDIA GPUs.
  • Experience with networking libraries like NCCL and protocols including RoCE and RDMA.
  • Strong understanding of parallel computing, distributed systems, and supercomputer architectures.
  • Ability to profile and optimize deep learning workflows to resolve networking-related bottlenecks.
  • Expertise in high-performance networking technologies like InfiniBand.

Preferred Skills

  • Proven experience profiling and optimizing large-scale deep learning training on NVIDIA supercomputers.
  • Expertise in optimizing networking parameters such as bandwidth, latency, or congestion control for deep learning workloads.

Education

Bachelor’s Degree

Related Jobs

No related jobs found

← Back to jobs