← Back to jobs
Toronto, ON, Canada
No related jobs found
Monitoring and Alerting:
Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.
Incident Response:
Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.
Automation:
Automate repetitive tasks and processes to improve efficiency and reduce manual effort.
Infrastructure Management:
Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources.
Capacity Planning:
Plan for future capacity needs to ensure systems can handle anticipated workloads.
Release Engineering:
Develop and maintain processes for deploying software updates and releases.
Collaboration:
Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.
Documentation:
Maintain clear and concise documentation of systems, processes, and procedures.
Continuous Improvement:
Identify areas for improvement and implement changes to enhance system reliability and performance.
Skills and Qualifications:
● Cloud Platform (OCP)
● 8+ Years experience in production support handling Prod incidents.
● Excellent knowledge of OCP and windows environment.
● Monitoring tools ( Dynatrace )
● Operating System (Windows, Linux)
● Scripting (Shell Scripting, Python, Power Shell)
● Database (SQL database management, MongoDb)
● Container Services (Kubernetes)
● Disaster Recovery Planning and execution
Any Graduate
No related jobs found
← Back to jobs