← Back to jobs

Square Hiring Logo
AI Reliability Engineer

Square Hiring

 

Tampa, FL, USA

Posted On: 7 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: AI Reliability Engineer
Tenure: Contract - W2
Related Jobs

No related jobs found

Description

You will build and maintain the application and infrastructure that powers large language models, bridging application development, cloud infrastructure, and ML operations.

This role is on-site.

Responsibilities

  • Design and implement autonomous, agentic AI workflows to eliminate operational toil and automate incident response.
  • Build and maintain the infrastructure and applications supporting large language model deployments.
  • Develop tool integrations via APIs and Model Context Protocol using Python or Go.
  • Manage distributed vector databases and apply cloud monitoring stacks to AI workloads.
  • Own CI/CD deployment pipelines and infrastructure as code strategies.

Required Skills

  • 5+ years of experience in Site Reliability Engineering or similar infrastructure roles.
  • Deep expertise in Kubernetes (EKS/GKE) and Infrastructure as Code with Terraform.
  • Strong proficiency in Python or Go for building system integrations.
  • Hands-on experience with LLM orchestration frameworks such as AutoGen, LangChain, or LlamaIndex.
  • Experience managing distributed vector databases like Pinecone, Milvus, Qdrant, or pgvector.
  • Advanced knowledge of cloud monitoring stacks (Datadog, Prometheus, OpenTelemetry).
  • Bachelor’s degree in Computer Science or related field.

Preferred Skills

  • Experience designing agentic AI workflows for incident automation.
  • Background in Gen AI systems and LLM infrastructure.

Education

Bachelor’s degree

Related Jobs

No related jobs found

← Back to jobs