← Back to jobs

TruEvan Technologies Logo
Site Reliability Engineer

TruEvan Technologies

 

New Jersey, United States

Posted On: 8 days ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own application reliability, observability, and incident response for distributed systems.

This role is on-site.

Responsibilities

  • Analyze application failures, latency issues, and degraded performance across distributed systems.
  • Perform root cause analysis (RCA) to isolate failing components and reduce Mean Time to Resolution.
  • Implement and enhance application-level observability using logs, metrics, and traces.
  • Build and maintain application topology maps to identify single points of failure.
  • Lead incident triage, improve escalation paths, and maintain troubleshooting runbooks.

Required Skills

  • 10+ years of experience in application engineering, production support, or SRE roles.
  • Strong experience troubleshooting Java, .NET, or Node.js applications.
  • Deep understanding of distributed systems, microservices architectures, and API behavior.
  • Hands-on experience with observability tools: Splunk, Dynatrace, AppDynamics, or Datadog.
  • Proficiency in CI/CD pipelines and SRE methodologies.
  • Experience with Azure and GCP cloud environments.

Preferred Skills

  • Experience implementing distributed tracing (OpenTelemetry, Jaeger, Zipkin).
  • Knowledge of resiliency patterns (circuit breakers, retries, fallbacks).

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs