← Back to jobs

Persistent Systems Logo
Site Reliability Engineer

Persistent Systems

 

Pune, Maharashtra, India

Posted On: Just posted
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Manage and maintain production Kubernetes clusters, AWS infrastructure, and Linux systems to ensure high availability and scalability.

This role is on-site.

Responsibilities

  • Administer and troubleshoot Linux-based production systems and AWS services (EC2, EKS, S3, IAM, VPC, Load Balancers).
  • Manage Kubernetes cluster operations, ensuring high availability, scalability, and reliability of distributed systems.
  • Develop and maintain Ansible playbooks for infrastructure automation and configuration management.
  • Support and enhance CI/CD pipelines, implementing monitoring, alerting, and incident response processes.
  • Perform root cause analysis, resolve production incidents, and maintain operational runbooks during on-call rotations.

Required Skills

  • 5+ years of hands-on experience with Kubernetes cluster operations and troubleshooting.
  • Strong expertise in Linux system administration and administration.
  • Proficiency in AWS cloud infrastructure management.
  • Experience with Ansible for automation and configuration management.
  • Familiarity with Git-based version control workflows.
  • Understanding of container technologies such as Docker.
  • Knowledge of Infrastructure as Code principles.
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, Splunk, AppDynamics).
  • Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent experience).

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs