← Back to jobs

Xforia Global Talent Solutions Logo
Site Reliability Engineer

Xforia Global Talent Solutions

 

Sunnyvale, CA, USA

Posted On: 15+ days ago
Experience: 8+ years
Availability: Onsite
Openings: 1
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will build and maintain scalable tools and services for machine learning training and inference.

This role is on-site.

Responsibilities

  • Implement continuous deployment using GitHub Actions and Flux.
  • Develop and deploy cloud solutions, including MLOps pipelines on AWS.
  • Containerize data science models using Docker and deploy them via Kubernetes.
  • Communicate technical requirements between data scientists, engineers, and architects.

Required Skills

  • 8+ years of experience in MLOps with strong knowledge of Kubernetes, Python, MongoDB, and AWS.
  • Experience building MLOps pipelines on cloud solutions (AWS).
  • Proficiency with Linux administration and experience with software development and test automation.
  • Experience developing containers and Kubernetes in cloud computing environments.
  • Familiarity with data-oriented workflow orchestration frameworks (Kubeflow, Airflow, Argo, etc.).
  • Ability to design and implement cloud solutions and build custom integrations between cloud systems using APIs.
  • Strong understanding of software testing, benchmarking, and continuous integration.

Preferred Skills

  • Experience developing and maintaining ML systems built with open-source tools.
  • Knowledge of ML models and LLM.

Education

Any Gradute

Related Jobs

No related jobs found

← Back to jobs