Description
Key Skills: AWS, Dynatrace, Splunk, Datadog, OpenTelemetry, Terraform, Ansible, Jenkins, Kubernetes, CI/CD
Good to Have Skills: Experience with cloud operations including EC2, EKS/ECS, RDS, S3, Lambda, CloudWatch, IAM, VPC. Knowledge of incident management, root cause analysis, SLIs/SLOs implementation, AIOps, disaster recovery planning, and reliability engineering practices. Strong collaboration and communication skills with technical and non-technical stakeholders.
Roles & Responsibilities:
- Design and implement monitoring and observability solutions using Dynatrace and Splunk along with Datadog and OpenTelemetry to build scalable automated platforms.
- Develop reusable patterns templates and automation scripts to drive consistency across observability practices and reduce manual effort in telemetry onboarding.
- Build and maintain dashboards that deliver actionable insights into system performance reliability and user experience across the organization.
- Integrate observability into CI/CD workflows using Jenkins and related tooling to enable continuous feedback and faster incident detection.
- Automate infrastructure provisioning and deployment using Terraform and Ansible to support observability solutions at enterprise scale.
- Implement and manage OpenTelemetry pipelines for standardized collection of traces metrics and logs supporting vendor-agnostic ingestion strategies.
- Collaborate with Business Units Developers and Platform Engineers to embed observability into software delivery lifecycle and improve developer experience.
- Define and implement SLIs SLOs and error budgets with Business Units to support reliability engineering and improve service health visibility.
- Enhance operational excellence by enabling proactive monitoring reducing customer pain points and streamlining incident response workflows effectively.
- Manage and operate systems hosted on AWS including EC2 EKS ECS RDS S3 Lambda CloudWatch IAM and VPC services.
- Participate in production incident response troubleshooting service restoration and perform comprehensive root cause analysis for post-incident reviews.
- Support high availability scalability and performance of production systems while implementing and maintaining SLIs SLOs and SLAs.
Experience Required: Overall 5+ years of experience in production support and building scalable automated and developer-friendly observability platforms. Minimum 3+ years of hands-on experience with AWS environments. Over 3 years of hands-on experience with Dynatrace with expertise in creating dashboards leveraging logs metrics and traces