← Back to jobs

E-Solutions Logo
Site Reliability/Observability Engineer

E-Solutions

 

Phoenix, AZ, USA

Posted On: 3 days ago
Experience: 5+ years
Availability: Hybrid
Openings: 2
Category: Site Reliability/Observability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will manage production environments by monitoring availability and system health for large-scale distributed software applications.

This role is on-site.

Responsibilities

  • Monitor system health and availability to ensure reliable production operations.
  • Build software and automated systems to manage platform infrastructure and applications.
  • Analyze metrics from operating systems and applications to perform tuning and fault finding.
  • Partner with development, Data Science, and MLOps teams to improve services through testing and release procedures.
  • Manage capacity planning, system design consulting, and troubleshooting of production issues.

Required Skills

  • 5+ years of professional experience in reliability or systems engineering.
  • Proficiency in programming using Python, Java, C/C++, Ruby, or JavaScript.
  • Experience managing cloud infrastructure using AWS CloudFormation or Terraform.
  • Hands-on experience with Amazon S3, SageMaker, and Amazon Bedrock.
  • Knowledge of cloud-native infrastructure including AWS Lambda and OpenShift.
  • Ability to manage APIs and troubleshoot API-related issues.
  • Experience balancing feature development speed with service-level objectives.
  • Degree in any graduate field.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs