← Back to jobs

Hexaware Technologies Limited Logo
Site Reliability Engineer
Posted On: Just posted
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will combine software engineering with IT operations to ensure the reliability, availability, scalability, and performance of critical systems.

This role is on-site.

Responsibilities

  • Design and implement automated solutions for application and service high availability.
  • Respond to production incidents, conduct post-mortems, and implement preventive measures.
  • Set up monitoring systems to track performance metrics and proactively address issues.
  • Analyze system performance, identify bottlenecks, and optimize for speed and resource utilization.

Required Skills

  • 8+ years of total professional experience.
  • Expertise in end-to-end observability: Elastic Observability, Elastic APM, Distributed Tracing, OpenTelemetry.
  • Proficiency in Linux/Unix system administration and cloud infrastructure (AWS, Azure, Google Cloud).
  • Strong programming skills in Java, Python, Go, Bash, Spring Boot, or PySpark.
  • Experience with data management and warehousing (MongoDB, Snowflake).
  • Familiarity with CI/CD and configuration management (Jenkins, Ansible, Terraform).
  • Experience with containerization and orchestration (Docker, Kubernetes/EKS).

Preferred Skills

  • Solid understanding of networking, databases, and distributed systems.
  • Proven experience managing incident response and post-mortem processes.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs