← Back to jobs

United Software Group Inc Logo
Site Reliability Engineer

United Software Group Inc

 

Halifax, NS, Canada

Posted On: 4 days ago
Experience: 5+ years
Availability: Remote
Openings: 1
Category: Site reliability engineering
Tenure: Full-time Only, Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

You will own the reliability and availability of mission-critical cloud infrastructure across GCP and Azure.

This role is remote.

Responsibilities

  • Support mission-critical applications in a 24/7 on-call rotation.
  • Design, deploy, and manage highly available cloud infrastructure across GCP and Azure.
  • Drive SRE practices including SLOs, SLIs, error budgets, incident response, and postmortems.
  • Manage Kubernetes environments (GKE/OpenShift) and service mesh technologies.
  • Implement Infrastructure as Code using Terraform and Config Connector.

Required Skills

  • 5+ years of experience in SRE, DevOps, or Cloud Engineering.
  • Strong expertise in Kubernetes, Docker, and container platforms.
  • Hands-on experience with GCP and cloud-native technologies.
  • Experience with Terraform and Infrastructure Automation.
  • Strong scripting skills in Python and/or Bash.
  • Experience with CI/CD pipelines and GitOps methodologies.
  • Knowledge of observability, monitoring, and incident management.
  • Experience supporting Akamai CDN and edge services.

Preferred Skills

  • Familiarity with Azure Pipelines, Argo CD, Argo Workflows, and OpenShift (OCP4).
  • Experience with Python, Shell Scripting, Ansible, AWX Tower, and Linux Administration.
  • Knowledge of Istio, Kiali, Gatekeeper, Grafana, Loki, Dynatrace, Prometheus, Vault, SOPS, External Secrets Operator, Kafka, Solr, Zookeeper, BigQuery, Bigtable, AlloyDB, Firestore, Pub/Sub, Apigee, Apache Web Server, PagerDuty, Camunda, WebMethods.

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs