Collaborate with the customer's cloud partners to design, implement, and operationalize innovative GPU hardware and software solutions.
Partner with Sales Account Managers and other business leads to identify and secure business opportunities for GPU infrastructure products and solutions.
Act as the primary technical point of contact for the customer through the full lifecycle of developing, building, and bringing large-scale GPU cloud infrastructure into production.
Lead regular technical customer meetings covering project and product details, feature discussions, introductions to new technologies, and debugging sessions.
Work with the customer to build proofs of concept addressing critical business needs across networking and compute infrastructure.
Prepare and deliver technical content to the customer, including presentations and workshops.
Analyze and develop joint solutions for customer performance and scaling issues.
Required Qualifications:
BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or another Engineering field, or equivalent experience.
Motivation and skill to own and drive technical engagements with customers across the full customer lifecycle.
7+ years of Solution Engineering (or similar Sales Engineering / Cloud Engineering) experience working directly with partners and customers.
Experience crafting and deploying large-scale cluster environments.
Practical expertise in data center design, development, and execution for AI and HPC.
Strong time management skills and the ability to balance multiple tasks, with clear communication through documents and presentations.
Nice to Have:
Practical familiarity with GPU hardware and networking components (Ethernet/InfiniBand), storage, and other elements of large-scale AI and HPC cluster environments.
Practical knowledge of GPU systems management technologies such as NCCL, DCGM, UFM, Mission Control, and Base Command Manager.
Background with at-scale GPU systems, including performance testing and AI benchmarking.
Hands-on experience with cluster administration and orchestration (SLURM, Kubernetes)