← Back to jobs

Vision Square Inc Logo
Senior Site Reliability Engineer

Vision Square Inc

 

Dallas, TX, USA

Posted On: 1 day ago
Experience: 10+ years
Availability: Hybrid
Openings: 1
Category: Senior Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

Job Summary
We are seeking a highly experienced Senior Site Reliability Engineer (SRE) with strong hands-on expertise in GCP, AWS, Azure, Kubernetes, Terraform, CI/CD, observability, and cloud automation.
The candidate will be responsible for improving platform reliability, scalability, availability, security, disaster recovery, monitoring, incident response, and operational efficiency across enterprise cloud environments.
Key Responsibilities
Design, operate, and optimize highly available GCP, AWS, and Azure cloud environments.
Manage Kubernetes platforms including GKE, EKS, AKS, and ECS, along with Docker containers.
Develop Infrastructure as Code using Terraform, CloudFormation, ARM Templates, and Bicep.
Build and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, Azure DevOps, and AWS CodePipeline.
Implement blue-green, canary, rolling deployments, automated rollback, and deployment validation.
Develop observability solutions using Prometheus, Grafana, ELK, OpenSearch, CloudWatch, Azure Monitor, and GCP Cloud Monitoring.
Define and manage SLIs, SLOs, SLAs, error budgets, MTTD, and MTTR.
Design and test disaster recovery, high availability, backup, and multi-region failover solutions.
Automate infrastructure and operational processes using Python, Bash, PowerShell, and Ansible.
Troubleshoot Linux, cloud infrastructure, networking, and production issues.
Participate in incident response, blameless postmortems, and reliability improvement initiatives.
Partner with development, security, infrastructure, and platform teams on cloud modernization and governance.
Required Qualifications
10+ years of experience in SRE, DevOps, cloud infrastructure, platform engineering, or production operations.
Strong hands-on experience with GCP, AWS, and Azure.
Deep Kubernetes and Docker experience, especially GKE, EKS, and AKS.
Strong Terraform / Infrastructure as Code experience.
Strong CI/CD and DevOps experience.
Experience with Prometheus, Grafana, ELK/OpenSearch, and cloud monitoring tools.
Strong experience with DR, HA, multi-region failover, and production reliability.
Strong scripting skills in Python, Bash, PowerShell, and/or Ansible.
Strong Linux administration and troubleshooting skills.
Good understanding of cloud networking including VPC/VNet, subnets, VPN, NAT, load balancers, Direct Connect, ExpressRoute, PrivateLink, Route 53, and Transit Gateway.
Strong understanding of SRE principles, SLI/SLO, error budgets, incident management, and postmortems.
Preferred
Experience with enterprise cloud migration and modernization.
Experience establishing SRE/reliability governance.
Cloud or Kubernetes certifications.
Experience with security, compliance, and cloud governance

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs