You will design and implement observability solutions across cloud environments.
Responsibilities
Build and support the modernization and integration of observability tools in private and public cloud offerings (GCP, AKS, EKS).
Design, implement, and automate telemetry, logging, and monitoring solutions, including dashboards, alerts, and CI/CD integration.
Enable teams to leverage observability data for reliability, performance, and security use cases; provide actionable recommendations.
Collaborate with DevOps, SRE, and security teams to share best practices and support adoption of observability standards.
Mentor and upskill client teams through knowledge transfer and participate in on-call activities as required.
Required Skills
At least 5 years of relevant experience in Observability, Logging, and Monitoring in enterprise environments.
Hands-on experience with observability tools such as Grafana, Prometheus, Loki, Cortex, Tempo, ElasticSearch, Datadog, Splunk, or equivalents.
Experience working with container technologies (Docker, Kubernetes) and orchestration platforms (GKE or similar).
Proficiency in setting up and configuring dashboards, alerts, and alarms on Grafana and/or GCP Monitoring.
Experience in integrating observability tools with CI/CD pipelines and automating through scripting (Python, Bash, JSON, YAML, Terraform or similar).
Proficiency with Linux operating systems and databases (MySQL, DB2, MSSQL, or similar).
Solid understanding of how enterprise service delivery components interact (web servers, application servers, databases, web services, storage, security).