← Back to jobs

Hallmark Global Technologies Inc Logo
Site Reliability Engineering
Posted On: 2 days ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: site reliability engineering
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

Own enterprise-wide observability and reliability standards for complex hybrid systems.

This role is hybrid.

Responsibilities

  • Design and maintain Dynatrace service topology, dashboards, alerting frameworks, and health indicators.
  • Implement SRE best practices including SLIs, SLOs, and error budgets to drive platform resilience.
  • Lead incident detection, response, and root cause analysis (RCA) while facilitating post-incident reviews.
  • Identify and mitigate reliability, performance, and capacity risks proactively across cloud and on-prem environments.
  • Define observability standards and mentor engineering teams on SRE principles and tooling.

Required Skills

  • 8–10+ years of experience in Site Reliability Engineering, DevOps, or Observability.
  • Hands-on expertise with Dynatrace is mandatory.
  • Strong experience with Kubernetes-based environments and complex enterprise systems (ERP, WMS, eCommerce).
  • Deep understanding of observability practices: metrics, logs, traces, and APM.
  • Experience implementing SLIs, SLOs, and error budgeting frameworks.
  • Strong troubleshooting and root-cause analysis skills in hybrid (cloud + on-prem) settings.
  • Proficiency in scripting or automation using Python, Bash, or similar languages.

Preferred Skills

  • Experience with cloud platforms (AWS, Azure, GCP) and CI/CD pipelines.
  • Exposure to microservices architecture and distributed systems.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs