← Back to jobs

HARP Technologies & Services Pvt Ltd Logo
DEVOPS SRE-AI

HARP Technologies & Services Pvt Ltd

 

Bangalore, Karnataka, India

Posted On: 11 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 1
Category: DevOPs SRE
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Job Description

We are looking for an experienced Site Reliability Engineer (SRE) / Platform Engineer with strong hands-on experience in SRE, Platform Engineering, DevOps, Cloud, Kubernetes, and Gen AI. The candidate will be responsible for building, scaling, and operating reliable, secure, and automated platform services. The role will focus on improving system reliability, reducing operational toil, enabling developer productivity, and contributing to COE initiatives, internal intellectual properties, and products to fulfil different *** requirements.

Key Responsibilities

  • Build and operate highly available, scalable, secure, and reliable platform services.
  • Implement SRE practices including SLIs, SLOs, error budgets, reliability engineering, and automation.
  • Manage, deploy, and scale Kubernetes-based workloads across cloud environments.
  • Develop and maintain CI/CD pipelines for application and infrastructure deployments.
  • Design, implement, and maintain Infrastructure as Code using Terraform and related tools.
  • Own production reliability, participate in on-call activities, incident response, troubleshooting, and Root Cause Analysis (RCA).
  • Enhance observability using metrics, logs, traces, monitoring, dashboards, and alerting solutions.
  • Identify and eliminate operational toil through automation and engineering best practices.
  • Work closely with development, DevOps, cloud, security, and infrastructure teams to improve platform reliability and availability.
  • Develop and maintain automation scripts and tools using Python, Go, Bash, or similar programming languages.
  • Support cloud-native application deployments and platform services across AWS, Azure, and GCP environments.
  • Contribute to platform engineering initiatives, internal COE projects, intellectual properties, and technology accelerators.
  • Support the development of internal platforms and products based on different *** requirements.
  • Monitor system health, capacity, availability, performance, and reliability of production environments.
  • Conduct incident analysis, troubleshooting, RCA, and implement preventive and corrective actions.
  • Contribute to continuous improvement of deployment, monitoring, reliability, and operational processes.
  • Collaborate with cross-functional teams to define and implement reliability and automation standards.

Mandatory Skills

  • 5–16 years of overall experience in SRE / Platform Engineering / DevOps / Cloud Engineering / AI Engineering.
  • Strong hands-on experience with Kubernetes and Docker.
  • Strong knowledge of Linux administration and troubleshooting.
  • Strong scripting/programming skills in Python, Go, Bash, or similar languages.
  • Hands-on experience with at least one major cloud platform such as AWS, Azure, or GCP.
  • Strong understanding of SRE principles including SLIs, SLOs, SLAs, error budgets, reliability, and availability.
  • Experience in CI/CD pipeline development and implementation.
  • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.
  • Experience with monitoring, observability, logging, metrics, tracing, and alerting tools.
  • Experience in production support, incident management, troubleshooting, and Root Cause Analysis (RCA).
  • Strong understanding of cloud-native architecture, automation, scalability, and high availability.
  • Strong analytical, troubleshooting, and problem-solving skills.

Good to Have

  • Experience with Internal Developer Platforms (IDP) and platform engineering frameworks.
  • Exposure to service mesh technologies such as Istio or Linkerd.
  • Knowledge of FinOps, cloud cost optimization, and resource management.
  • Experience with Jenkins/Git/GitLab/Azure DevOps or similar CI/CD tools.
  • Exposure to Prometheus, Grafana, ELK/EFK, OpenTelemetry, or similar observability tools.
  • Experience with AWS/Azure/GCP Kubernetes services such as EKS, AKS, or GKE.
  • Experience with Helm, ArgoCD, GitOps, or similar cloud-native deployment tools.
  • Exposure to Gen AI, LLMs, AI-powered DevOps/SRE tools, or AI-based operational automation.
  • Experience in developing AI/GenAI-based automation or intelligent SRE solutions is an advantage.
  • Cloud or Kubernetes certifications are an advantage.
  •  

Education
Bachelor’s/Master’s degree in Computer Science, Information Technology, Engineering, or a related field

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs