← Back to jobs

Technocraft Solutions Logo
Machine Learning Performance Engineer

Technocraft Solutions

 

New York, NY, USA

Posted On: Just posted
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Machine Learning Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will optimize end-to-end machine learning training performance, focusing on deep GPU architecture and low-level system tuning.

This role is on-site.

Responsibilities

  • Debug and optimize ML training performance across the full stack.
  • Profile and tune GPU kernels using PTX, SASS, warps, and Tensor Cores.
  • Implement optimizations for CUDA graph launches, warp-level synchronization, and asynchronous memory operations.
  • Tune high-performance networking for GPU clusters using InfiniBand, RoCE, GPUDirect, and NVLink.
  • Leverage NCCL and MPI for distributed GPU training and collective communication.

Required Skills

  • 5+ years of experience in systems performance engineering or ML infrastructure.
  • Deep knowledge of GPU architecture including memory hierarchy and cooperative groups.
  • Hands-on experience with CUDA GDB, NVIDIA Nsight Systems, and Nsight Compute.
  • Proficiency with GPU libraries: Triton, CUTLASS, CUB, Thrust, cuDNN, and cuBLAS.
  • Strong understanding of latency and throughput optimization in GPU environments.
  • Experience with distributed training frameworks and high-performance networking technologies.
  • Bachelor's degree in Computer Science or related field.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs