← Back to jobs

Echo IT Solutions Logo
Site Reliability Engineer

Echo IT Solutions

 

Sunnyvale, CA, USA

Posted On: 8 days ago
Experience: 5+ years
Availability: Onsite
Openings: 2
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will maintain production systems on AWS and GCP, ensuring stability and performance for containerized applications.

This role is on-site.

Responsibilities

  • Maintain and debug Kubernetes clusters and containerized applications written in Golang, Java, or Python.
  • Manage infrastructure using Terraform, CloudFormation, and cloud deployment tools.
  • Implement and maintain monitoring solutions including Prometheus, Grafana, CloudWatch, and Datadog.
  • Support continuous integration practices using Jenkins, Travis CI, or CircleCI.
  • Automate system operations using Linux, Python, and Shell scripting.

Required Skills

  • 5+ years of experience with Linux, Python, and Shell scripting.
  • Hands-on experience maintaining production systems on AWS and/or GCP.
  • Proficiency in Kubernetes cluster management and container debugging.
  • Strong understanding of infrastructure as code (Terraform, CloudFormation).
  • Experience with monitoring and logging tools (Prometheus, Grafana, ELK, CloudWatch).
  • Familiarity with distributed systems components: Kafka, Spark, Zookeeper, Cassandra, Redis.
  • Knowledge of CI/CD pipelines and tools.

Preferred Skills

  • Experience with Storm, ElasticSearch, Nginx, or AWS S3/GCP Client libraries.
  • Background in Golang application development.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs