← Back to jobs

eTeam Logo
Cloud DevOps Engineer

eTeam

 

Toronto, ON, Canada

Posted On: 13 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: Cloud DevOps Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

Key Responsibilities:
Azure Infrastructure and Platform Deployment:

  • Design, deploy, configure, and maintain infrastructure within Microsoft Azure.
  • Deploy and manage virtual machines, containers, Kubernetes clusters, networking, storage, and supporting platform services.
  • Support infrastructure across Development, QA, and Production environments.
  • Establish repeatable and reliable deployment processes using Infrastructure as Code and CI/CD automation.
  • Maintain secure, resilient, and appropriately sized platform environments.


Kubernetes and Container Management:

  • Deploy, configure, and operate containerized applications using Kubernetes.
  • Manage container lifecycle, configuration, secrets, networking, storage, and application dependencies.
  • Monitor container and cluster health, resource consumption, capacity, and performance.
  • Troubleshoot deployment, networking, configuration, and runtime issues.
  • Establish appropriate standards for container deployment and Kubernetes operations.


Performance, Load Management, and Scaling:

  • Monitor platform demand, workload patterns, resource utilization, and application performance.
  • Configure horizontal and vertical scaling policies for containers and supporting infrastructure.
  • Develop intelligent scaling approaches based on workload, queue depth, response time, resource utilization, and business demand.
  • Conduct capacity planning and identify potential performance bottlenecks before they affect production.
  • Help introduce predictive or AI-assisted scaling and platform management capabilities.


Dashboards and Platform Visibility:

  • Design and build advanced operational dashboards using tools such as **Grafana, Kibana, Azure Monitor, Application Insights**, and similar technologies.
  • Create clear executive, operational, application, and infrastructure views of platform health.
  • Build dashboards covering availability, performance, capacity, errors, latency, traffic, container health, AI workloads, and service dependencies.
  • Establish meaningful service-level indicators, service-level objectives, and reliability metrics.
  • Continuously improve dashboards so that issues, trends, and risks can be quickly identified.
  • Advanced dashboard design and dashboard-building experience is a core requirement for this role.


Monitoring and Alerting:

  • Implement monitoring and alerting across infrastructure, applications, containers, integrations, and AI platform services.
  • Configure actionable alerts that identify real production risks while minimizing unnecessary alert noise.
  • Establish thresholds, anomaly detection, health checks, synthetic monitoring, and automated remediation where appropriate.
  • Create operational runbooks and troubleshooting guidance.
  • Work with development and architecture teams to improve platform observability.


Production Reliability and Support:

  • Support the stability, availability, and operational readiness of the production AI platform.
  • Investigate and resolve platform, deployment, infrastructure, monitoring, and performance issues.
  • Participate in root-cause analysis and implement preventative improvements.
  • Ensure that production support processes, documentation, and escalation paths are established before platform usage increases.
  • Provide very light production support during 2026, with no regular after-hours support currently anticipated.
  • Help prepare the operating model for increased platform adoption and support requirements expected in 2027.


Required Qualifications:

  • Strong professional experience in DevOps, Site Reliability Engineering, cloud infrastructure, or platform engineering.
  • Advanced hands-on experience with Microsoft Azure.
  • Strong experience deploying and operating Kubernetes environments.
  • Strong knowledge of containerization technologies such as Docker.
  • Experience deploying and supporting containerized applications in Development, QA, and Production environments.
  • Advanced experience designing and building dashboards using Grafana, Kibana, Azure Monitor, Application Insights, or comparable tools.
  • Strong experience implementing monitoring, observability, logging, alerting, and operational health checks.
  • Experience managing application load, infrastructure capacity, performance, and automated scaling.
  • Experience with CI/CD pipelines and automated application deployment.
  • Experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates.
  • Strong troubleshooting skills across applications, containers, infrastructure, networking, and cloud services.
  • Ability to work independently while collaborating closely with developers, architects, AI engineers, and platform stakeholders.


Preferred Qualifications:

  • Experience supporting AI, machine learning, data, or high-compute platforms.
  • Experience monitoring AI models, inference services, token usage, GPU workloads, API consumption, queues, or model performance.
  • Experience implementing automated remediation, predictive monitoring, or AI-assisted platform operations.
  • Familiarity with AWS services and cloud operations.
  • Experience with Elasticsearch, Log Analytics, OpenTelemetry, Prometheus, or similar observability technologies.
  • Experience defining service-level indicators, service-level objectives, and reliability standards.
  • Experience with security, identity, secrets management, and cloud governance within Azure

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs