Description
You will own production support and observability to ensure system stability and performance.
Responsibilities
- Identify and triage issues proactively by correlating data from logs, observability dashboards, and recent infrastructure or application changes.
- Debug failures across the full tech stack, including application, database, container platforms, and network layers.
- Lead incident triage calls and direct technical teams on necessary corrective actions during high-visibility outages.
- Configure observability dashboards and monitor system performance using Splunk, AppDynamics, Grafana, and synthetic monitoring.
- Perform deep-dive analysis including heap dumps, memory leak detection, and resource optimization.
Required Skills
- 5+ years of experience in production support and SRE observability.
- Proficiency with Splunk (including Splunk APM and Splunk O11y), AppDynamics, Grafana, RedMetrics, and 1000Eyes.
- Hands-on experience debugging VMs, load balancers, firewalls, API gateways, and Linux/Unix environments.
- Experience managing containerized workloads using Docker and Kubernetes.
- Experience with cloud platforms including AWS, Azure, and PCF.
- Technical ability to analyze network traffic via Wireshark and use APM/NMON tools.
- Competency in database performance monitoring and analysis.
- Knowledge of Docker, Kubernetes, AWS, PCF, Azure, Java, Python, Oracle, Cassandra, and SQL Server.
- Ability to lead technical triage sessions involving executive leadership.
Preferred Skills
- Development experience with Java, Python, AWS, Azure, Oracle, Cassandra, SQL Server, MySQL, or MongoDB.
- Experience with ServiceNow, including AIOps, self-heal tools, and automated playbooks.