← Back to jobs
Irvine, CA, USA
No related jobs found
• Administer, monitor, and support hybrid cloud infrastructure and data platform services across AWS and Microsoft Azure. • Deploy, configure, monitor, and troubleshoot containerized workloads running on Kubernetes and Docker. • Support platform components including Databricks, dbt, Apache Airflow, AutoSys, LASR, and Caspian, including scheduled jobs, dependencies, and integrations. • Investigate incidents and service degradation across cloud, network, compute, storage, container, orchestration, and data platform layers. • Manage incidents, service requests, problems, changes, and engineering backlog items through ServiceNow and Jira in accordance with established SLAs and change controls. • Automate provisioning, configuration, deployment, health checks, and recurring support activities using scripting, Infrastructure as Code, and CI/CD practices. • Monitor platform availability, performance, capacity, job execution, alerts, and logs; improve dashboards and operational visibility. • Support upgrades, patching, vulnerability remediation, access controls, backup, disaster recovery readiness, and production release activities. • Create and maintain runbooks, technical documentation, troubleshooting guides, and knowledge articles. • Collaborate with application, data engineering, security, network, DevOps, and service management teams to resolve dependencies and improve platform reliability. • Participate in on-call or production-support rotation as required by the engagement. Generic Managerial Skills, If any" "• Cloud Platforms: Hands-on administration and support experience across AWS and Microsoft Azure, basic knowledge on compute, storage, networking, identity, access, and monitoring. • Containers: Practical experience with Kubernetes and Docker, including deployments, services, configuration, logs, health checks, and basic cluster troubleshooting. • Data Platforms: Working knowledge of Databricks, dbt, Apache Airflow, AutoSys, and enterprise data platform operations. Exposure to LASR and Caspian is preferred. • Infrastructure as Code and CI/CD: Experience with Terraform and CI/CD tools such as Harness, Azure DevOps, GitHub Actions, Jenkins, or equivalent. • Scripting: Ability to automate operational tasks using Python, Shell, Bash, or PowerShell. • Observability: Experience with cloud-native monitoring and log analysis using Azure Monitor, AWS CloudWatch, Datadog, Splunk, or similar tools. • IT Service Management: Proficiency with ServiceNow and Jira for incident, problem, change, service request, and backlog management. • Security and Reliability: Understanding of IAM/RBAC, secrets management, vulnerability remediation, patching, platform resilience, and disaster recovery controls. • Professional Skills: Strong troubleshooting, documentation, collaboration, and written and verbal communication skills in an enterprise environment." "• Cloud Platforms: Hands-on administration and support experience across AWS and Microsoft Azure, basic knowledge on compute, storage, networking, identity, access, and monitoring. • Containers: Practical experience with Kubernetes and Docker, including deployments, services, configuration, logs, health checks, and basic cluster troubleshooting. • Data Platforms: Working knowledge of Databricks, dbt, Apache Airflow, AutoSys, and enterprise data platform operations. Exposure to LASR and Caspian is preferred. • Infrastructure as Code and CI/CD: Experience with Terraform and CI/CD tools such as Harness, Azure DevOps, GitHub Actions, Jenkins, or equivalent. • Scripting: Ability to automate operational tasks using Python, Shell, Bash, or PowerShell. • Observability: Experience with cloud-native monitoring and log analysis using Azure Monitor, AWS CloudWatch, Datadog, Splunk, or similar tools. • IT Service Management: Proficiency with ServiceNow and Jira for incident, problem, change, service request, and backlog management. • Security and Reliability: Understanding of IAM/RBAC, secrets management, vulnerability remediation, patching, platform resilience, and disaster recovery controls. • • Professional Skills: Strong troubleshooting, documentation, collaboration, and written and verbal communication skills in an enterprise environment
Bachelor's degree
No related jobs found
← Back to jobs