← Back to jobs

Xforia Global Talent Solutions Logo
Site Reliability Engineer
Posted On: 1 day ago
Experience: 8+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will own the reliability, deployment, and scaling of machine learning and data services.

This role is on-site.

Responsibilities

  • Develop and deploy scalable tools for ML training and inference.
  • Build MLOps pipelines on AWS and implement cloud solutions.
  • Automate continuous deployment using GitHub Actions and Flux.
  • Document operational processes and system designs.
  • Communicate technical requirements between data scientists, engineers, and architects.

Required Skills

  • 8+ years of professional experience in MLOps environments.
  • Expert proficiency with Kubernetes and Docker.
  • Strong command of Python and Linux administration.
  • Experience designing cloud solutions using AWS.
  • Familiarity with MongoDB and Microservices architecture.
  • Ability to build custom integrations between cloud systems via APIs.
  • Experience with CI/CD practices and test automation.
  • Knowledge of ML models, LLMs, and data-oriented workflow orchestration (e.g., Airflow).
  • Experience with Site Reliability principles and Platform engineering.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs