Description
ROLE OVERVIEW
Seeking a highly experienced Technical Lead SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of enterprise platforms and modern web applications. This role is ideal for a senior engineer, technical lead, or architect who enjoys solving complex technical challenges, mentoring engineers, and influencing technical direction while remaining hands-on. The position offers a clear growth path into a future Technical Manager SRE and Platform Engineering role as organizational needs and leadership responsibilities expand.
The Technical Lead will serve as a senior technical leader for Site Reliability Engineering (SRE), monitoring and observability, domain portfolio management, cloud platform operations, infrastructure automation, web application security, and application availability. The role will work closely with engineering teams in Dallas, Europe, and Asia Pacific to establish consistent operational standards, improve platform reliability, and drive technology modernization across global technology landscape.
The successful candidate will play a critical role in ensuring the reliability, performance, security, and operational health of Mary Kay's externally facing digital platforms through ownership of key platform services including observability, domain services, DNS, certificate management, CDN technologies, web application firewalls, cloud platform infrastructure, and Infrastructure-as-Code (IaC) solutions.
KEY RESPONSIBILITIES
- Provide technical leadership across Platform Engineering and Site Reliability Engineering functions.
- Establish engineering standards, operational best practices, and reliability objectives.
- Lead technical decision-making for cloud infrastructure, observability platforms, domain services, infrastructure automation, application security, and operational tooling.
- Mentor engineers and provide technical coaching across multiple disciplines.
- Drive technical roadmaps and continuous improvement initiatives.
- Evaluate, promote, and help operationalize emerging engineering capabilities, including AI-assisted development tools, coding agents, Infrastructure-as-Code automation, and other technologies that improve engineering productivity, quality, and speed of delivery.
- Lead enterprise reliability initiatives focused on availability, scalability, resiliency, performance, and operational excellence.
- Define and drive adoption of SLOs, SLIs, Error Budgets, Incident Management, and Root Cause Analysis.
- Drive automation initiatives that reduce operational overhead and improve service reliability.
- Serve as the technical owner for the enterprise monitoring and observability platform supporting APM, infrastructure monitoring, synthetic monitoring, Real User Monitoring (RUM), centralized logging, and distributed tracing.
- Define dashboards, alerting standards, operational metrics, and reporting.
- Lead governance of the global domain portfolio including registrations, renewals, DNS services, certificate lifecycle management, and vendor relationships ensuring security and compliance.
- Provide technical leadership for AWS infrastructure, Kubernetes/EKS, CDN, DNS, SSL/TLS, load balancing, Web Application Firewalls (WAF), and edge security services.
- Design and support Infrastructure-as-Code solutions using Terraform, establish standards for cloud provisioning, automation, and environment consistency.
- Promote automated provisioning, version control, testing, and infrastructure governance.
- Administer and optimize AWS WAF and related WAF technologies managing rules, rate limiting, bot protection, IP reputation controls, and application-layer threat mitigation.
- Partner with Information Security to improve web application protection capabilities.
- Serve as a senior escalation point for complex production issues, lead troubleshooting across client-side and server-side technologies, diagnose issues involving browser behavior, APIs, DNS, CDN, WAF, load balancing, networking, cloud infrastructure, and application performance.
- Drive reliability, resiliency, and end-user experience improvements.
- Work closely with engineering teams across North America, Europe, and Asia Pacific to participate in technical reviews, architecture discussions, operational planning, and knowledge sharing to evolve a follow-the-sun operating model.
- Participate in scheduled on-call rotations supporting critical platforms and services, provide leadership during major incidents and after-hours escalations.
- Support maintenance, upgrades, deployments, and disaster recovery activities.
REQUIRED QUALIFICATIONS
- 8+ years supporting enterprise applications, cloud platforms, infrastructure services, or web technologies.
- 3+ years serving as a Technical Lead, Senior Engineer, Architect, or equivalent.
- Strong experience with AWS, Terraform, Kubernetes/EKS, DNS, CDN, AWS WAF, SSL/TLS, observability platforms, CI/CD, and SRE practices.
- Experience troubleshooting large-scale customer-facing web applications.
- Experience managing global domain portfolios.
- Experience with enterprise observability platforms and global support models.
- Experience developing enterprise Infrastructure-as-Code frameworks and reusable Terraform modules.
- Experience leveraging AI-assisted development tools and coding agents such as GitHub Copilot, Microsoft Copilot, Claude Code, Cursor, Amazon Q Developer or similar technologies to accelerate software delivery, infrastructure automation, troubleshooting, and operational efficiency.
- AWS, Terraform, Kubernetes, SRE, networking, security, or cloud certifications