← Back to jobs

Inherent Technology Logo
DevOps Engineer

Inherent Technology

 

Washington D.C., DC, USA

Posted On: 1 day ago
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Devops Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will build and maintain an inference platform for serving large language models across diverse GPU architectures.

This role is remote.

Responsibilities

  • Develop and maintain inference services optimized for NVIDIA and AMD GPUs.
  • Build tooling and observability to monitor system health and enable auto-tuning.
  • Construct benchmarking frameworks to assess model serving performance and guide infrastructure tuning.
  • Create native cross-platform inference support for various model architectures.
  • Contribute to open-source inference engines to improve performance on cloud infrastructure.

Required Skills

  • 5+ years of experience building distributed systems using Kubernetes and Docker.
  • Deep experience with CI/CD pipelines, APIs, authentication, and authorization.
  • Experience hosting and running inference for Large Language Models (LLMs).
  • Familiarity with inference engines such as vLLM, SGLang, and Modular Max.
  • Experience with distributed inference serving frameworks like llm-d, NVIDIA Dynamo, and Ray Serve.
  • Hands-on experience with AMD and NVIDIA GPUs, including CUDA, ROCm, AITER, NCCL, and RCCL.
  • Knowledge of distributed inference optimization techniques, including tensor/data parallelism and KV cache optimizations.

Preferred Skills

  • Proficiency with cloud environments and infrastructure as code practices.
  • Strong verbal and written communication skills for technical documentation.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs