Description
You will own production support and observability to ensure system stability and performance.
This role is on-site.
Responsibilities
- Identify and triage issues proactively by correlating data from logs, observability dashboards, and recent infrastructure or application changes.
- Debug failures across the full tech stack, including application, database, container platforms, and network layers.
- Lead incident triage calls and direct technical teams on necessary corrective actions during high-visibility outages.
- Perform deep-dive analysis including heap dumps, memory leak detection, and resource optimization.
Required Skills
- 5+ years of experience in production support and SRE observability.
- Proficiency with Splunk (APM and O11y), AppDynamics, Grafana, RedMetrics, and 1000Eyes.
- Hands-on experience debugging VMs, load balancers, firewalls, API gateways, and Linux/Unix environments.
- Experience managing containerized workloads using Docker and Kubernetes.
- Experience with cloud platforms including AWS, Azure, and PCF.
- Ability to analyze network traffic via Wireshark and use APM/NMON tools.
- Competency in database performance monitoring and analysis for Oracle, Cassandra, and SQL Server.
Preferred Skills
- Development experience with Java, Python, AWS, Azure, Oracle, Cassandra, SQL Server, MySQL, or MongoDB.
- Experience with ServiceNow, including AIOps, self-heal tools, and automated playbooks.