← Back to jobs
Morristown, NJ, USA
No related jobs found
Key Responsibilities
Design| develop| and maintain Grafana dashboards| visualizations| and reports for infrastructure| applications| and business metrics. Configure and manage Grafana Alerting for proactive monitoring and incident response. Integrate Grafana with various data sources such as: Prometheus Lokio Elasticsearch/OpenSearch Influx DB SQL databases Cloud monitoring platforms (AWS CloudWatch| Azure Monitor| Google Cloud Operations)Develop observability solutions covering metrics| logs| and traces. Implement monitoring strategies for Kubernetes| containers| virtual machines| cloud infrastructure| and enterprise applications. Collaborate with SRE| DevOps| Application Support| and Platform Engineering teams to define monitoring requirements. Automate dashboard deployment and configuration using Infrastructure as Code (IaC) tools. Tune monitoring systems to minimize alert fatigue and improve operational efficiency. Perform root cause analysis using monitoring and logging data. Create and maintain technical documentation| monitoring standards| and operational runbooks. Support capacity planning| performance analysis| and system optimization initiatives. Ensure security and governance for monitoring platforms and data access. Required Skills Strong experience with Grafana dashboard development and administration. Expertise in Grafana Alerting| notification channels| and alert rule management. Experience with Grafana Enterprise is a plus.Experience with: Prometheus|Loki|Tempo|OpenTelemetry|Elasticsearch/OpenSearch|Splunk (preferred)Understanding of Metrics| Logs| and Distributed Tracing concepts. Experience with one or more cloud platforms: AWS| Azure and GCP Familiarity with Kubernetes and container orchestration.Knowledge of CI/CD pipelines and DevOps practices. Proficiency in scripting languages such as: Python Bash/Shell PowerShell Experience with Terraform| Ansible| or similar automation tools. Experience querying and analyzing monitoring data. Strong analytical and troubleshooting skills. Excellent communication and stakeholder management .Ability to work independently and collaboratively. Problem-solving mindset with attention to detail. Strong documentation and knowledge-sharing capabilities. Grafana Enterprise deployment experience. Open Telemetry implementation experience. Experience with AIOps and observability platforms . Exposure to application performance monitoring (APM) tools such as Dynatrace| AppDynamics| Datadog| or New Relic. Keywords: Grafana| Prometheus| Loki| Tempo| Open Telemetry| Kubernetes| Cloud Monitoring| Observability| SRE| DevOps| Terraform| AWS| Azure| Monitoring| Alerting| PromQL| LogQL
Bachelor's degree
No related jobs found
← Back to jobs