← Back to jobs
Naperville, IL, USA
No related jobs found
Key Responsibilities:
Reliability Engineering & Operations
Ensure high availability and reliability of enterprise applications in a 24×7 production environment.
Monitor applications, batch jobs, and workflows to maintain operational continuity.
Incident, Problem & Change Management
Lead and manage major incidents (P1/P2) and drive resolution to minimize business impact.
Perform root cause analysis (RCA) and implement preventive measures.
Ensure adherence to SLA/SLO and ITIL-based incident, problem, and change management processes.
Monitoring & Observability
Design and maintain monitoring dashboards.
Implement proactive alerting and improve system observability.
Troubleshooting & Support
Diagnose and resolve application and data-related issues using SQL queries and log analysis.
Provide backend validation and technical support across distributed environments.
Release & Deployment Support
Support release deployments, change validation, and post-deployment activities.
Participate in disaster recovery testing and release readiness validation.
Collaboration & Documentation
Collaborate with infrastructure, DBA, and development teams to resolve technical issues.
Create and maintain operational documentation, runbooks, and knowledge base articles.
Required Skills & Qualifications:
Core Skills
Site Reliability Engineering (SRE) and Application Support Incident & Problem Management Root Cause Analysis (RCA) SLA / SLO Compliance Batch Monitoring & Scheduling ITIL Framework
Technical Skills
CI/CD Tools: GitHub
Cloud Platforms: AWS (EC2, S3, VPC)
Databases: Oracle, SQL Server
Languages: SQL, SQR, Basic Java
Ticketing Tools: ServiceNow, Jira
Operating Systems: UNIX, Linux, Windows
Bachelor's degree
No related jobs found
← Back to jobs