← Back to jobs

NVIDIA Logo
Senior Software Engineer, AI Inference Systems

NVIDIA

 

Toronto, ON, Canada

Posted On: Just posted
Experience: 7+ years
Availability: Onsite
Openings: 1
Category: Senior Software Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the optimization and orchestration of large-scale AI inference systems, focusing on vLLM and GPU hardware capabilities.

This role is on-site.

Responsibilities

  • Profile and optimize the vLLM inference framework using speculative decoding and various parallelism strategies.
  • Develop, optimize, and benchmark GPU kernels through fusion, autotuning, and memory layout optimization.
  • Build and extend high-level DSLs and compiler infrastructure to improve kernel productivity and hardware utilization.
  • Define inference benchmarking methodologies and contribute to MLPerf submissions.
  • Architect scheduling and orchestration for containerized large-scale inference deployments across cloud GPU clusters.

Required Skills

  • 7+ years of professional experience with strong programming skills in Python and C/C++.
  • Solid CS fundamentals including algorithms, data structures, operating systems, computer architecture, and distributed systems.
  • Knowledge of performance engineering in ML frameworks like PyTorch and inference engines like vLLM.
  • Familiarity with GPU programming, CUDA, memory hierarchy, streams, and NCCL.
  • Proficiency with profiling and debugging tools such as Nsight Systems and Nsight Compute.
  • Experience with containers and orchestration tools including Docker, Kubernetes, and Slurm.
  • Experience with cloud platforms AWS, GCP, or Azure.

Preferred Skills

  • Experience with Go or Rust.
  • Hands-on work with ML compilers and DSLs such as Triton, MLIR/LLVM, or XLA.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs