Description
You will lead technical implementation for AI and GPU workloads, deploying large language models and managing GPU infrastructure.
This role is on-site.
Responsibilities
- Deploy and configure LLMs using PyTorch and Hugging Face Transformers.
- Run GPU-based inference workloads utilizing NVIDIA CUDA.
- Set up and manage Linux-based environments, including shell scripting and package management.
- Manage GPU instances across cloud providers like AWS EC2 or GCP.
- Collaborate with the team to integrate models into production pipelines.
Required Skills
- 5+ years of professional experience in software engineering or ML operations.
- Expertise in Python, PyTorch, and Hugging Face Transformers.
- Mastery of Linux, specifically Ubuntu environments.
- Proficiency in shell scripting and package management.
- Hands-on experience deploying LLMs (Llama preferred).
- Experience running GPU-based inference using CUDA.
- Familiarity with cloud GPU platforms including AWS and Google Cloud Platform (GCP).
- Ability to create and maintain isolated Python environments using Conda and pip.