← Back to jobs
Boston, MA, USA
No related jobs found
Must Haves
• 5+ years of experience in Production Support, Site Reliability Engineering (SRE), or a related support engineering function, including 2+ years supporting healthcare environments.
• Hands-on experience managing and supporting production systems within Microsoft Azure.
• Strong expertise in Azure SQL, T-SQL, and healthcare data exchange standards such as HL7 and FHIR.
• Advanced scripting and automation skills using PowerShell, Azure CLI, Bash, or Python.
• Ability to serve as the primary point of contact for high-severity production incidents and critical escalations.
• Experience leading incident response efforts, coordinating bridge calls, and driving rapid resolution within strict SLA requirements.
• Proven ability to perform root cause analysis (RCA) across infrastructure, application, and database layers.
• Commitment to participating in a 24/7 on-call rotation and supporting mission-critical healthcare applications.
• Working knowledge of HIPAA, HITRUST, and healthcare data privacy and security requirements.
Pluses
• Experience supporting Epic EHR environments, including Interconnect, Bridges, Cogito, or related Epic modules.
• Familiarity with Azure OpenAI Services, AI-powered applications, and Large Language Model (LLM) operations.
• Knowledge of monitoring AI application performance, including prompt/response latency and operational health metrics.
• Experience working with vector databases or AI-driven search and retrieval platforms.
• Exposure to enterprise disaster recovery planning, compliance audits, and healthcare regulatory environments.
Day-to-Day
• Act as the primary support lead for high-severity production outages and critical business escalations.
• Coordinate incident response activities and bridge calls to restore service for mission-critical healthcare applications.
• Perform troubleshooting and root cause analysis across Azure infrastructure, databases, and application services.
• Execute recovery and remediation activities to minimize disruption to patient care and business operations.
• Monitor and maintain the health, availability, and performance of Azure App Services, Virtual Machines, and Azure Functions.
• Analyze telemetry and system performance data using Azure Monitor, Log Analytics, and Application Insights.
• Support and optimize Azure SQL databases, enterprise data pipelines, and cloud storage solutions.
• Deploy system patches, hotfixes, and updates in accordance with security and operational standards.
• Audit system logs, user access, and security controls to ensure ongoing compliance and data protection.
• Develop and maintain automation scripts to reduce manual support activities and improve operational efficiency.
• Create and maintain technical documentation, support runbooks, disaster recovery procedures, and compliance artifacts.
• Collaborate with engineering, infrastructure, security, and application teams to drive service reliability and continuous improvement
Any Graduate
No related jobs found
← Back to jobs