← Back to jobs

GuideAI Logo
Site Reliability Engineer

GuideAI

 

Bengaluru, Karnataka, India

Posted On: 15+ days ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

Design, build, and operate highly reliable, scalable infrastructure for a multi-tenant SaaS platform.

This role is on-site.

Responsibilities

  • Automate deployment, provisioning, and operational workflows across cloud infrastructure and applications.
  • Build and maintain observability systems including metrics, logging, tracing, and dashboards.
  • Define and track Service Level Objectives (SLOs) and reliability metrics to drive system resilience.
  • Lead incident response, root cause analysis, and blameless postmortems to improve platform stability.
  • Develop internal tools and services to reduce manual effort and enhance engineering efficiency.

Required Skills

  • 5+ years of experience in infrastructure development and production operations.
  • Strong programming skills in Python or Go.
  • Deep experience with AWS and building production systems at scale.
  • Hands-on expertise with Kubernetes (EKS), Docker, Helm, CNI, and Ingress networking.
  • Proficiency with Infrastructure as Code using Terraform or Terragrunt.
  • Experience with CI/CD pipelines and version control (GitHub).
  • Strong understanding of Linux systems and networking fundamentals.
  • Familiarity with incident management practices in microservices environments.

Preferred Skills

  • Experience with observability platforms such as Datadog, Prometheus, or OpenTelemetry.
  • Knowledge of messaging/streaming systems (Kafka, SQS) and relational databases (Aurora, RDS).
  • Experience with secure access patterns including SSO, SAML, and OAuth.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs