Description
You will design and maintain observability architectures, focusing on Splunk integration for application telemetry and reliability engineering.
This role is on-site.
Responsibilities
- Define and implement strategies for ingesting application-specific telemetry (logs, metrics, traces) into the Splunk environment.
- Design monitoring frameworks to provide actionable insights into application health, performance bottlenecks, and user experience.
- Map complex application-level parameters to Splunk dashboards to ensure granular visibility into microservices and workflows.
- Automate incident response, reducing MTTD and MTTR through proactive observability and SRE principles.
- Analyze and optimize Splunk queries and data ingestion pipelines for speed, cost, and relevance.
Required Skills
- Deep expertise in Splunk, including SPL, data modeling, dashboard creation, and Splunk App for Infrastructure/APM.
- Strong understanding of distributed systems, microservices architectures, and application instrumentation (OpenTelemetry, agents, SDKs).
- Scripting experience in Python, Bash, or Go to automate configuration tasks and data pipelines.
- Familiarity with cloud-native technologies (AWS, Azure, GCP) and container orchestration (Kubernetes, Docker).
- Proven experience implementing Error Budgets, SLOs, and SLIs.
- Bachelor's degree in Computer Science or related field.
- 5+ years of experience in Site Reliability Engineering or related roles.
Preferred Skills
- Experience partnering with CIS teams to align application configurations with enterprise standards for performance, security, and data retention.