← Back to jobs

CareerNet Technologies Pvt Ltd Logo
Site Reliability Engineering

CareerNet Technologies Pvt Ltd

 

Chennai, Tamil Nadu, India

Posted On: 10 days ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: Site reliability engineering
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Key Skills: Problem Management, Scripting, Observability, Monitoring, Site Reliability Engineer, Cloud, Incident, AWS, Bash Shell, Grafana, Azure, Prometheus, Python

Roles and Responsibilities:

  • Define and implement service-health metrics, monitoring, alerting, and dashboards for end-to-end observability.
  • Support incident response, problem management, and operational readiness, including root-cause analysis.
  • Identify and remediate reliability, performance, and resilience risks across critical services.
  • Engineer automation to reduce operational toil and improve MTTR.
  • Contribute to SLIs/SLOs, error-budget management, and continuous improvement of operational resilience.

Skills Required:

  • 8+ years in SRE/production-engineering/DevOps roles
  • Scripting expertise in Python or Bash
  • Strong monitoring and observability skills
  • Experience with cloud platforms like AWS or Azure
  • Incident management and problem-solving capabilities

Good to Have:

  • Experience with Grafana and Prometheus
  • Familiarity with CI/CD and infrastructure-as-code
  • Exposure to container orchestration with Kubernetes

Education: Bachelor's/Master's degree

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs