← Back to jobs

StaffXpert LLC Logo
Lead Site Reliability Engineer

StaffXpert LLC

 

San Francisco, CA, USA

Posted On: 15+ days ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: Lead Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own production stability and incident response for large-scale ecommerce systems.

This role is on-site.

Responsibilities

  • Coordinate production incident response across application, infrastructure, and support teams.
  • Use Dynatrace and Splunk to analyze logs, metrics, traces, and system dependencies.
  • Design operational dashboards and define monitoring thresholds to reduce alert noise.
  • Conduct root cause analysis and implement corrective actions for system failures.
  • Develop runbooks and automation scripts to improve operational efficiency.

Required Skills

  • 8+ years in Site Reliability Engineering, DevOps, or Platform Engineering.
  • Hands-on experience with Dynatrace and Splunk.
  • Strong understanding of microservices, APIs, Kubernetes, and cloud platforms.
  • Experience supporting large-scale ecommerce or enterprise production environments.
  • Proficiency in Python or shell scripting for automation.
  • Experience with ServiceNow, Jira, PagerDuty, and Microsoft Teams.
  • Solid troubleshooting skills for latency, throughput, and database performance.

Preferred Skills

  • Familiarity with Dynatrace DQL, Grail, Smartscape, Davis AI, and SLOs.
  • Experience with automated incident analysis or self-healing workflows.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs