← Back to jobs

Stefanini IT Solutions Logo
Site Reliability Engineer

Stefanini IT Solutions

 

Dearborn, MI, USA

Posted On: 8 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

 Seeking an experienced SRE who is responsible for ensuring availability, reliability and performance of cloud and network systems and services by automating routine manual tasks

This position is focused on observability, monitoring, and technical consulting across our GCP-based data platforms. working hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to keep critical systems healthy, performant, and reliable at scale.

 

Responsibilities

  • Collaborate with Infrastructure teams in implementing critical solutions by automating routine tasks
  • Monitor and manage production environments, proactively identifying and resolving issues.
  • Participate in building advanced tooling for system access monitoring, log session recording, and administration of reliability across multiple geographically distributed data centers.
  • Engage with engineering teams to improve on-call efficiency, drive incident management and post-mortem analysis.
  • Perform capacity planning and optimization to support growing demand and traffic patterns.
  • Maintaining, monitoring and alerting systems for proactive system health checks.
  • Continuously improve system performance, stability, and security through data-driven analysis and optimization.
  • Facilitate knowledge sharing by creating and maintaining comprehensive documentation & diagrams.

 

Job Requirements

Details:

 

Experience Required

  • 4+ years of experience in development
  • Hands on experience with Big Query & Dynatrace
  • Hands-on experience with Google Cloud Platform (GCP).
  • Proficiency with monitoring/observability tools, ideally Dynatrace (or comparable, e.g., Datadog, New Relic).
  • Familiarity with ITSM tools such as ServiceNow (incident, problem, change management)

 

Experience Preferred

  • Familiarity with the use of AI tools - agents, skills, LLMs, copilot. Experience defining and tracking SLAs/SLOs/SLIs
  • Experience with GCP Cloud Run, Python, Troubleshooting (Problem Solving)

 

Education Required

  • Bachelor's Degree
     

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs