Description
You will benchmark AI models and optimize inference pipelines across diverse hardware architectures. Your work focuses on measuring latency, throughput, memory usage, and power efficiency to drive system performance improvements.
This role is on-site.
Responsibilities
- Benchmark AI models using PyTorch and ONNX Runtime to measure latency, throughput, memory, and power efficiency.
- Identify system bottlenecks and optimize AI inference pipelines on CPU, GPU, and NPU architectures.
- Compare and analyze performance metrics across NVIDIA, AMD, Qualcomm, and other AI platforms.
- Automate benchmarking processes, reporting, and performance analysis using Python and C/C++.
Required Skills
- 7+ years of experience in systems engineering or AI infrastructure.
- Strong proficiency in Python and C/C++ programming.
- Hands-on experience with PyTorch and ONNX Runtime.
- Deep understanding of Linux systems, CPU/GPU/NPU architecture, and performance tuning.
- Proficiency with profiling tools such as Perf, VTune, or Nsight.
- Experience with AI inference optimization techniques.
Preferred Skills
- Familiarity with MLPerf, YOLOv8, Whisper, or OpenVLA.
- Experience in Edge AI or robotics applications.