← Back to jobs
Boston, MA, USA
No related jobs found
Job Description
looking for an experienced Senior Azure Site Reliability Engineer to ensure the continuous uptime, data integrity, and compliance of healthcare software applications. Operating within an Azure cloud environment, this individual will serve as the primary line of defense for critical production issues, helping minimize disruption to clinical and administrative workflows. The role functions as the primary on-call responder and requires a proactive technical leader who can rapidly resolve critical incidents, automate system recovery processes, and maintain strict HIPAA security standards.
Title: Senior Azure Site Reliability Engineer
Remote / Hybrid Schedule: 100% remote
Must Haves
• 5+ years of experience in Production Support, Site Reliability Engineering (SRE), or a related support engineering function, including 2+ years supporting healthcare environments.
• Hands-on experience managing and supporting production systems within Microsoft Azure.
• Strong expertise in Azure SQL, T-SQL, and healthcare data exchange standards such as HL7 and FHIR.
• Advanced scripting and automation skills using PowerShell, Azure CLI, Bash, or Python.
• Ability to serve as the primary point of contact for high-severity production incidents and critical escalations.
• Experience leading incident response efforts, coordinating bridge calls, and driving rapid resolution within strict SLA requirements.
• Proven ability to perform root cause analysis (RCA) across infrastructure, application, and database layers.
• Commitment to participating in a 24/7 on-call rotation and supporting mission-critical healthcare applications.
• Working knowledge of HIPAA, HITRUST, and healthcare data privacy and security requirements.
Pluses
• Experience supporting Epic EHR environments, including Interconnect, Bridges, Cogito, or related Epic modules.
• Familiarity with Azure OpenAI Services, AI-powered applications, and Large Language Model (LLM) operations.
• Knowledge of monitoring AI application performance, including prompt/response latency and operational health metrics.
• Experience working with vector databases or AI-driven search and retrieval platforms.
• Exposure to enterprise disaster recovery planning, compliance audits, and healthcare regulatory environments.
Day-to-Day
• Act as the primary support lead for high-severity production outages and critical business escalations.
• Coordinate incident response activities and bridge calls to restore service for mission-critical healthcare applications.
• Perform troubleshooting and root cause analysis across Azure infrastructure, databases, and application services.
• Execute recovery and remediation activities to minimize disruption to patient care and business operations.
• Monitor and maintain the health, availability, and performance of Azure App Services, Virtual Machines, and Azure Functions.
• Analyze telemetry and system performance data using Azure Monitor, Log Analytics, and Application Insights.
• Support and optimize Azure SQL databases, enterprise data pipelines, and cloud storage solutions.
• Deploy system patches, hotfixes, and updates in accordance with security and operational standards.
• Audit system logs, user access, and security controls to ensure ongoing compliance and data protection.
• Develop and maintain automation scripts to reduce manual support activities and improve operational efficiency.
• Create and maintain technical documentation, support runbooks, disaster recovery procedures, and compliance artifacts.
• Collaborate with engineering, infrastructure, security, and application teams to drive service reliability and continuous improvement
Any Graduate
No related jobs found
← Back to jobs