← Back to jobs
Dallas, TX, USA
No related jobs found
Key Responsibilities
Reliability & Operations
• Ensure the availability, performance, scalability, and reliability of data and intelligence platforms.
• Manage production environments supporting data ingestion, processing, transformation, storage, analytics, and AI/ML workloads.
• Participate in on-call rotations and incident response activities.
• Lead troubleshooting efforts for complex production issues and drive root cause analysis (RCA).
• Develop and implement service level indicators (SLIs), service level objectives (SLOs), and error budgets.
Automation & Engineering
• Design and develop automation to improve operational efficiency and system reliability.
• Build self-healing solutions and automate routine operational tasks.
• Create tools and scripts to monitor, deploy, and manage large-scale distributed systems.
• Improve deployment processes through CI/CD pipelines and Infrastructure as Code (IaC).
Observability & Monitoring
• Design and maintain monitoring, logging, tracing, and alerting solutions.
• Create dashboards and actionable alerts to proactively identify service degradation.
• Analyze system performance metrics and recommend optimization opportunities.
• Drive observability standards across data services and platforms.
Platform & Infrastructure Management
• Support cloud-based infrastructure and platform services across Azure, AWS, or GCP environments.
• Optimize compute, storage, networking, and data platform resources.
• Work with containerized and Kubernetes-based workloads.
• Ensure high availability and disaster recovery capabilities are implemented and tested.
Data Platform Reliability
• Support modern data ecosystems including data lakes, warehouses, streaming platforms, and analytics environments.
• Monitor ETL/ELT pipelines, batch processing, real-time streaming, and data orchestration services.
• Partner with Data Engineers to improve pipeline reliability and data quality monitoring.
• Ensure platform scalability for growing data volumes and user demands.
Security & Compliance
• Implement security best practices and operational controls.
• Support compliance requirements related to data governance and privacy.
• Collaborate with security teams to remediate vulnerabilities and improve platform security posture.
Continuous Improvement
• Conduct post-incident reviews and drive corrective and preventive actions.
• Identify reliability risks and implement long-term improvements.
• Promote a culture of operational excellence, resilience, and automation.
• Contribute to engineering standards, runbooks, knowledge sharing, and best practices.
Required Qualifications
Education
• Bachelor’s degree in computer science, Information Technology, Engineering, or related field, or equivalent practical experience.
Experience
• 3+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related role.
• Experience supporting production environments with high availability requirements.
• Experience managing cloud infrastructure and distributed systems.
Technical Skills
• Strong knowledge of Linux systems administration.
• Proficiency in one or more programming languages such as Python, Java, Go, C#, or JavaScript.
• Experience with CI/CD tools and deployment automation.
• Experience with Infrastructure as Code tools such as Terraform, ARM, or Bicep.
• Experience with Kubernetes and container technologies.
• Knowledge of monitoring and observability technologies such as Prometheus, Grafana, Datadog, Azure Monitor, Splunk, or OpenTelemetry.
• Understanding of networking, DNS, load balancing, and distributed systems concepts.
Experience supporting data platforms such as:
o Azure Data Lake
o Azure Synapse Analytics
o Databricks
o Snowflake
o Kafka
o SQL/NoSQL databases
o Data orchestration platforms
Preferred Qualifications
• Experience supporting AI/ML platforms and MLOps environments.
• Experience with Azure cloud-native services.
• Familiarity with data governance and data quality frameworks.
• Knowledge of reliability engineering best practices and SRE methodologies.
• Experience implementing SLOs, SLIs, and error budgets.
• Experience supporting large-scale analytics and business intelligence environments.
• Azure, AWS, Kubernetes, Terraform, or DevOps certifications.
Key Competencies
• Problem-solving and analytical thinking
• Incident management and troubleshooting
• Automation-first mindset
• Collaboration and stakeholder management
• Strong communication skills
• Continuous learning and innovation
• Customer-focused approach
• Operational excellence and accountability Success Measures
• Maintaining high availability and reliability targets for critical services.
• Reducing operational toil through automation.
• Improving platform observability and incident response effectiveness.
• Meeting service-level objectives and performance goals.
• Enhancing deployment reliability and operational efficiency.
• Driving measurable improvements in system resilience, scalability, and customer experience
Bachelor's degree
No related jobs found
← Back to jobs