← Back to jobs

Intelliest Logo
Site Reliability Engineer

Intelliest

 

Hyderabad, Telangana, India

Posted On: 3 days ago
Experience: 4+ years
Availability: Hybrid
Openings: 2
Category: Site reliability engineering
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the observability stack and monitoring infrastructure to ensure system reliability.

This role is on-site.

Responsibilities

  • Design and maintain Grafana dashboards to visualize metrics, logs, and traces.
  • Configure integrations between Grafana and Prometheus, Loki, and Tempo.
  • Implement synthetic monitoring to simulate user interactions and monitor endpoint availability.
  • Set up Tempo for distributed tracing across microservices architectures.
  • Configure alerting rules and notifications to detect anomalies in real-time.

Required Skills

  • 4+ years of experience in Site Reliability Engineering (SRE).
  • Proficiency with Grafana for dashboarding and visualization.
  • Hands-on experience with Prometheus for metrics collection.
  • Experience with Loki for log management and analysis.
  • Expertise in using LogQL and PromQL query languages.
  • Practical knowledge of Tempo for distributed tracing.
  • Ability to integrate monitoring solutions with API Management (APIM) platforms.

Preferred Skills

  • Develop automation scripts for deploying and scaling monitoring infrastructure.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs