← Back to jobs

Han Staffing Logo
Sr Site Reliability Engineer

Han Staffing

 

Bellevue, WA, USA

Posted On: 15+ days ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You own the design, deployment, and operation of enterprise AI Gateway infrastructure supporting OpenAI and internal LLM services.

This role is on-site.

Responsibilities

  • Implement and manage regional routing, failover strategies, and upstream host configurations for AI traffic.
  • Develop and maintain Helm charts, Kubernetes manifests, and Jinja templates for multi-environment deployments.
  • Enable per-API configuration for rate limiting, feature toggles, security credentials, and regional overrides.
  • Build and maintain monitoring and troubleshooting frameworks for AI workloads using Splunk and Grafana.

Required Skills

  • 10+ years of experience in site reliability engineering or related infrastructure roles.
  • Strong experience with Kubernetes, Helm, and cloud-native networking.
  • Hands-on experience with Istio/service mesh, routing rules, and traffic management.
  • Proficiency in Python, Bash, and Jinja templating for infrastructure automation.
  • Experience operating production-grade APIs under high reliability standards.
  • Deep understanding of SRE principles, monitoring, alerting, and incident management.
  • Experience building observability frameworks using Splunk and Grafana.
  • Experience working with AI/LLM APIs (OpenAI or similar) in an enterprise context.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs