← Back to jobs

Akkodis Logo
AI Reliability Engineer

Akkodis

 

Tampa, FL, USA

Posted On: 1 day ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: AI Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will design, scale, and maintain highly available infrastructure for LLM training, fine-tuning, and inference workloads.

This role is on-site.

Responsibilities

  • Architect and implement agentic AI systems for automated alert triage, root cause analysis, and self-healing.
  • Optimize GPU utilization, cluster health, and orchestration across Kubernetes-based environments.
  • Define and monitor AI-specific SLOs/SLIs such as latency, throughput, and cost efficiency.
  • Ensure reliability of vector databases and RAG pipelines for AI data processing.
  • Integrate AI-driven incident management with ChatOps and enforce security guardrails for GenAI systems.

Required Skills

  • 8–10 years of experience in Site Reliability Engineering, DevOps, or AI/ML infrastructure roles.
  • Strong expertise in Kubernetes and cloud platforms (AWS, GCP, or Azure).
  • Proficiency in infrastructure automation using Terraform and CI/CD pipelines.
  • Proven experience with GenAI/LLM systems and orchestration frameworks like LangChain or AutoGen.
  • Experience building reliable, scalable AI platforms and vector databases.
  • Bachelor’s or master’s degree in computer science, Engineering, Data Science, or a related field.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs