We are seeking an experienced Cloud Operations and Systems Engineer with strong expertise in observability, monitoring, and infrastructure operations. The ideal candidate will design and implement proactive monitoring and alerting capabilities across critical hybrid infrastructure platforms, ensuring timely detection and resolution of service-impacting events. This role requires deep experience with Dynatrace, infrastructure engineering, and Site Reliability Engineering (SRE) practices.
Key Responsibilities
Design, implement, and optimize enterprise monitoring and observability solutions using Dynatrace.
Develop meaningful service-level monitoring and alerting for critical infrastructure services including:
Active Directory
DNS
VMware vCenter
Backup Infrastructure
Windows Servers
Linux Servers
Configure and implement approximately 20+ advanced monitoring and alerting use cases across the enterprise infrastructure estate.
Establish proactive alerting mechanisms to reduce Mean Time to Detect (MTTD) and minimize operational risks.
Create dashboards, service health views, dependency maps, and operational runbooks.
Analyze infrastructure performance trends and identify optimization opportunities.
Collaborate with operations, platform engineering, and application teams to improve service reliability.
Implement SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational resilience metrics.
Document monitoring standards, alerting procedures, and operational best practices.
Provide mentoring and enablement to internal operations teams.
Required Qualifications
7+ years of experience in Infrastructure Operations or Systems Engineering.
5+ years of experience implementing enterprise monitoring solutions.
3+ years of hands-on experience with Dynatrace administration and configuration.
Strong experience supporting Windows and Linux environments.
Experience with Active Directory, DNS, VMware vSphere/vCenter, and enterprise backup platforms.
Understanding of infrastructure architecture and service dependencies.
Experience developing actionable alerts and reducing alert fatigue.
Strong troubleshooting and root-cause analysis skills.
Experience with ITSM tools and operational processes.
Preferred Qualifications
Site Reliability Engineering (SRE) experience.
Experience with cloud platforms including AWS and Azure.
Knowledge of automation tools such as Ansible, PowerShell, or Terraform.
Experience integrating Dynatrace with ServiceNow or other ITSM platforms.