← Back to jobs

Openkyber Logo
BizOps SRE

Openkyber

 

St. Louis, MO, USA

Posted On: 30+ days ago
Experience: 10+ years
Availability: Onsite
Openings: 1
Category: BizOps Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

We are seeking an experienced BizOps SRE / Production Engineer to support critical production services and drive reliability, automation, and operational excellence across a global environment. The ideal candidate will have strong hands-on experience with SRE, DevOps, production operations, cloud infrastructure, Kubernetes, Terraform, observability, incident management, and automation. This role will focus on improving service stability, reducing MTTR and operational toil, strengthening production readiness, and partnering closely with Engineering, SRE, Operations, and Product teams.

Key Responsibilities:

  • Own production operations, service stability, availability, and continuity for critical services.
  • Lead Major Incidents / Technical Response Team (TRT) activities and drive timely resolution and reduced MTTR.
  • Provide clear and timely communication to stakeholders, customers, and leadership during production incidents.
  • Drive operational readiness for application releases, infrastructure migrations, and peak business events.
  • Implement and improve monitoring, observability, alerting, and service health dashboards.
  • Lead Root Cause Analysis (RCA) and problem management activities.
  • Identify reliability risks and implement proactive solutions to improve system resilience.
  • Support capacity planning, performance improvement, and resilience initiatives.
  • Automate repetitive operational activities and reduce manual operational toil.
  • Develop and maintain reusable runbooks, SOPs, operational procedures, and readiness standards.
  • Partner with Engineering, Infrastructure SRE, Operations, Product, and other technical teams.
  • Convert production insights and incident learnings into engineering improvements and roadmap initiatives.
  • Participate in global production support and operational activities as required.

Required Skills:

  • Strong experience in SRE, DevOps, Production Engineering, or Site Reliability Engineering.
  • Hands-on experience supporting production environments.
  • Strong knowledge of Linux administration and troubleshooting.
  • Experience with at least one major cloud platform: AWS, Azure, or Google Cloud Platform.
  • Hands-on experience with Kubernetes.
  • Experience with Terraform and Infrastructure as Code.
  • Strong scripting/programming experience with one or more of: Python, Go, Java, Bash.
  • Experience with observability and monitoring tools such as: Splunk, Dynatrace, Prometheus, Grafana, Datadog.
  • Strong experience with Incident Management, Major Incident Management, RCA, Problem Management, and Change/Release Management.
  • Experience with CI/CD pipelines and automation.
  • Strong troubleshooting and production support skills.
  • Excellent communication and stakeholder management skills.

Preferred Qualifications:

  • Experience supporting 24x7/global production environments.
  • Experience working with distributed/global teams.
  • Experience in financial services, banking, payments, or other high-availability environments.
  • Experience with production readiness reviews and release/migration readiness.
  • Experience developing operational runbooks, SOPs, and reliability standards.
  • Demonstrated experience reducing incidents, MTTR, or operational toil through automation and process improvements

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs