← Back to jobs
Dallas, TX, USA
No related jobs found
Seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Google Cloud Platform (GCP). Design, implement, and maintain highly available, scalable, and reliable applications and infrastructure on GCP. Manage and support GCP services, cloud infrastructure, networking, compute, storage, and security. Develop and maintain CI/CD pipelines and automate deployment and operational processes. Implement infrastructure as code using Terraform or similar tools. Monitor system health, performance, availability, and reliability using tools such as Prometheus, Grafana, Cloud Monitoring, and Cloud Logging. Troubleshoot production issues, perform root-cause analysis, and drive permanent resolutions. Implement automation to reduce manual operational activities and improve system reliability. Work closely with development, DevOps, cloud, and application teams to support releases and production environments. Participate in incident management, capacity planning, disaster recovery, and performance optimization. Follow best practices for security, reliability, scalability, and observability in GCP environments. Required Skills: Strong experience with GCP Site Reliability Engineering / DevOps Kubernetes / GKE Terraform / Infrastructure as Code CI/CD Linux Monitoring & Observability Prometheus / Grafana Cloud Logging & Monitoring Production troubleshooting and incident management Scripting/automation
Bachelor's degree
No related jobs found
← Back to jobs