← Back to jobs

E-Solutions Logo
Site Reliability Engineer

E-Solutions

 

Dallas, TX, USA

Posted On: 2 days ago
Experience: 3+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

Key Responsibilities

 

Reliability & Operations

•       Ensure the availability, performance, scalability, and reliability of data and intelligence platforms.

•       Manage production environments supporting data ingestion, processing, transformation, storage, analytics, and AI/ML workloads.

•       Participate in on-call rotations and incident response activities.

•       Lead troubleshooting efforts for complex production issues and drive root cause analysis (RCA).

•       Develop and implement service level indicators (SLIs), service level objectives (SLOs), and error budgets.

 

Automation & Engineering

•       Design and develop automation to improve operational efficiency and system reliability.

•       Build self-healing solutions and automate routine operational tasks.

•       Create tools and scripts to monitor, deploy, and manage large-scale distributed systems.

•       Improve deployment processes through CI/CD pipelines and Infrastructure as Code (IaC).

 

Observability & Monitoring

•       Design and maintain monitoring, logging, tracing, and alerting solutions.

•       Create dashboards and actionable alerts to proactively identify service degradation.

•       Analyze system performance metrics and recommend optimization opportunities.

•       Drive observability standards across data services and platforms.

 

Platform & Infrastructure Management

•       Support cloud-based infrastructure and platform services across Azure, AWS, or GCP environments.

•       Optimize compute, storage, networking, and data platform resources.

•       Work with containerized and Kubernetes-based workloads.

•       Ensure high availability and disaster recovery capabilities are implemented and tested.

 

Data Platform Reliability

•       Support modern data ecosystems including data lakes, warehouses, streaming platforms, and analytics environments.

•       Monitor ETL/ELT pipelines, batch processing, real-time streaming, and data orchestration services.

•       Partner with Data Engineers to improve pipeline reliability and data quality monitoring.

•       Ensure platform scalability for growing data volumes and user demands.

 

Security & Compliance

•       Implement security best practices and operational controls.

•       Support compliance requirements related to data governance and privacy.

•       Collaborate with security teams to remediate vulnerabilities and improve platform security posture.

 

Continuous Improvement

•       Conduct post-incident reviews and drive corrective and preventive actions.

•       Identify reliability risks and implement long-term improvements.

•       Promote a culture of operational excellence, resilience, and automation.

•       Contribute to engineering standards, runbooks, knowledge sharing, and best practices.

 

Required Qualifications

Education

•       Bachelor’s degree in computer science, Information Technology, Engineering, or related field, or equivalent practical experience.

Experience

•       3+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related role.

•       Experience supporting production environments with high availability requirements.

•       Experience managing cloud infrastructure and distributed systems.

 

Technical Skills

•       Strong knowledge of Linux systems administration.

•       Proficiency in one or more programming languages such as Python, Java, Go, C#, or JavaScript.

•       Experience with CI/CD tools and deployment automation.

•       Experience with Infrastructure as Code tools such as Terraform, ARM, or Bicep.

•       Experience with Kubernetes and container technologies.

•       Knowledge of monitoring and observability technologies such as Prometheus, Grafana, Datadog, Azure Monitor, Splunk, or OpenTelemetry.

•       Understanding of networking, DNS, load balancing, and distributed systems concepts.

 

Experience supporting data platforms such as:

o   Azure Data Lake

o   Azure Synapse Analytics

o   Databricks

o   Snowflake

o   Kafka

o   SQL/NoSQL databases

o   Data orchestration platforms

 

Preferred Qualifications

•       Experience supporting AI/ML platforms and MLOps environments.

•       Experience with Azure cloud-native services.

•       Familiarity with data governance and data quality frameworks.

•       Knowledge of reliability engineering best practices and SRE methodologies.

•       Experience implementing SLOs, SLIs, and error budgets.

•       Experience supporting large-scale analytics and business intelligence environments.

•       Azure, AWS, Kubernetes, Terraform, or DevOps certifications.

 

Key Competencies

•       Problem-solving and analytical thinking

•       Incident management and troubleshooting

•       Automation-first mindset

•       Collaboration and stakeholder management

•       Strong communication skills

•       Continuous learning and innovation

•       Customer-focused approach

•       Operational excellence and accountability Success Measures

•       Maintaining high availability and reliability targets for critical services.

•       Reducing operational toil through automation.

•       Improving platform observability and incident response effectiveness.

•       Meeting service-level objectives and performance goals.

•       Enhancing deployment reliability and operational efficiency.

•       Driving measurable improvements in system resilience, scalability, and customer experience

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs