← Back to jobs

ApTask Logo
AI Engineer

ApTask

 

San Francisco, CA, USA

Posted On: 30+ days ago
Experience: 12 years
Availability: Hybrid
Openings: 1
Category: Generative AI Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will manage and optimize large-scale GPU clusters and AI infrastructure.

Responsibilities

  • Develop and maintain GPU-based clusters ranging from 10 to 1000 nodes to ensure availability and performance.
  • Administer AI platforms including distributed client services, LLMs, Vector-DB, and AI inferencing.
  • Manage deployments, resource allocation, monitoring, and security for client and AI platforms.
  • Monitor AI system performance to ensure adherence to industry best practices.
  • Document procedures and publish recommendations to improve AI infrastructure and internal delivery tools.

Required Skills

  • 12 years of professional experience.
  • Proficiency in RoCEv2 and 200G/400G cluster interconnect networking.
  • Hands-on experience with K8s, KVM, and Ubuntu.
  • Strong programming skills in Python, Go, Rust, and Shell.
  • Experience managing GPU drivers and optimizing GPU-based services.
  • Expertise in managing distributed client services and AI inferencing workloads.
  • Ability to support AI infrastructure requirements through cross-functional collaboration.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs