Description
You will design and implement architectural frameworks to improve enterprise observability, performance, and system reliability.
Responsibilities
- Design, prototype, and document technical solutions that improve system resilience and operational predictability.
- Publish technology strategies and observability standards to support AI/MLOps maturity.
- Develop Observability Driven Development procedures using open standards like OTel and MELTS.
- Implement full-stack reliability patterns and monitoring standards for scalability and availability.
- Create AI-augmented testing strategies to enable federated execution and enterprise governance.
Required Skills
- 5+ years of experience in engineering or architectural roles.
- Deep expertise in AI-Ops and AI/MLOps practices.
- Strong background in Observability and SRE engineering.
- Experience with open standards including OTel and MELTS.
- Proven ability to translate business requirements into technical diagnostic and descriptive designs.
- Experience establishing monitoring and alerting standards for large-scale systems.
- Ability to develop training and documentation for enterprise-wide knowledge transfer.
Preferred Skills
- Experience in designing integration patterns for prescriptive disruption response.