← Back to jobs

Saransh Inc Logo
Site Reliability Engineer with ML Platform

Saransh Inc

 

Sunnyvale, CA, USA

Posted On: 1 day ago
Experience: 6+ years
Availability: Hybrid
Openings: 1
Category: Site Reliability Engineer
Tenure: Full-time Only
Related Jobs

No related jobs found

Description

You will build and maintain scalable MLOps infrastructure and deployment pipelines for machine learning training and inference.

This role is on-site.

Responsibilities

  • Build MLOps pipelines and cloud solutions on AWS.
  • Manage continuous deployment using GitHub Actions, Flux, and Kustomize.
  • Containerize and deploy data science models using Docker and Kubernetes.
  • Develop and deploy scalable tools for machine learning training and inference.
  • Collaborate with data scientists, engineers, and architects to document processes.

Required Skills

  • 6+ years of experience in MLOps.
  • Strong proficiency with Kubernetes, Python, MongoDB, and AWS.
  • Experience with Docker and container orchestration in cloud environments.
  • Hands-on experience with Linux administration.
  • Knowledge of ML models and Large Language Models (LLM).
  • Experience building custom integrations between cloud systems using APIs.
  • Understanding of software testing, benchmarking, and continuous integration.
  • Ability to translate business needs into technical requirements.
  • Experience with software development and test automation.

Preferred Skills

  • Familiarity with MLOps frameworks like Kubeflow, MLFlow, DataRobot, or Airflow.
  • Experience with data-oriented workflow orchestration such as Argo or Airflow.
  • Knowledge of Apache SOLR.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs