← Back to jobs
Alpharetta, GA, USA
No related jobs found
Job Responsibilities include:
Monitor, maintain, and improve the reliability, availability, and performance of cloud-based applications and infrastructure on Google Cloud Platform (GCP).
Manage and resolve high-priority production incidents, perform root cause analysis, and implement preventive measures to minimize recurring issues.
Support and optimize GCP services including Pub/Sub, BigQuery, Spanner, Dataflow, and Firestore to ensure stable and efficient operations.
Deploy, manage, and troubleshoot containerized applications using Kubernetes, GKE, Docker, and Helm.
Implement and enhance monitoring, alerting, and observability solutions while reducing alert noise and improving operational efficiency.
Collaborate with engineering and AI teams to support Vertex AI model training, pipeline deployments, and Agentic AI solutions in production environments.
Required Qualifications:
Bachelor’s degree in computer science, Information Technology, Engineering, or a related technical field.
6–8 years of experience in Site Reliability Engineering (SRE), Cloud Operations, or Infrastructure Engineering.
Strong hands-on experience with Google Cloud Platform (GCP), including Pub/Sub, BigQuery, Spanner, Dataflow, Firestore, Kubernetes/GKE, Docker, and Helm.
Proven expertise in incident management, production support, monitoring optimization, and experience with Vertex AI and Agentic AI solutions
Any Graduate
No related jobs found
← Back to jobs