← Back to jobs

NVIDIA Logo
AI/ML Performance Engineer

NVIDIA

 

Redmond, WA, U.S.

Posted On: 3 days ago
Experience: 3+ years
Availability: Onsite
Openings: 1
Category: Sr. AI/ML Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will build performance models to prototype Generative AI algorithms and drive model-hardware co-design.

Responsibilities

  • Develop high-fidelity performance models to prototype emerging algorithmic techniques in Generative AI.
  • Design targeted optimizations for inference deployment to maximize accuracy, throughput, and interactivity.
  • Model end-to-end performance impact of GenAI workflows, including Agentic Pipelines and inference-time compute scaling.
  • Quantify performance benefits of optimizations to guide software and hardware roadmaps.
  • Collaborate with DL researchers, hardware architects, and software engineers.

Required Skills

  • Master's degree in Computer Science, Electrical Engineering, or a related field.
  • 3+ years of experience in system evaluation of AI/ML workloads or performance analysis and modeling.
  • Strong background in computer architecture, roofline modeling, queuing theory, and statistical performance analysis.
  • Solid understanding of LLM internals, including attention mechanisms, FFN structures, model parallelism, and inference serving.
  • Proficiency in Python for simulator design and data analysis.
  • Experience defining metrics, designing experiments, and visualizing large performance datasets.
  • Ability to distill complex analyses into clear recommendations for technical and non-technical stakeholders.

Preferred Skills

  • Proficiency in C++.
  • Experience with GPU computing (CUDA).
  • Experience with deep learning frameworks including PyTorch, TRT-LLM, VLLM, and SGLang.

Education

Master's degree

Related Jobs

No related jobs found

← Back to jobs