← Back to jobs

Meraki7 Logo
LLMOps Lead

Meraki7

 

Toronto, ON, Canada

Posted On: 30+ days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: LLMOps Engineer
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will architect, manage, and scale the operational infrastructure for large language model workflows.

Responsibilities

  • Design, build, and maintain LLM infrastructure including model serving, versioning, rollout, and inference pipelines.
  • Automate model training, fine-tuning, evaluation, and retraining workflows.
  • Implement observability, monitoring, logging, and alerting for model performance, drift, latency, and reliability.
  • Optimize cost, scaling, caching, batching, and hardware utilization for inference.
  • Lead MLOps best practices covering reproducibility, infrastructure as code, and model lineage.

Required Skills

  • 5+ years of backend, MLOps, or infrastructure experience, including 2+ years with large language models or deep learning production systems.
  • Deep familiarity with PyTorch, TensorFlow, Transformers, and Hugging Face.
  • Experience with model serving platforms such as Triton, TorchServe, KFServing, or TensorFlow Serving.
  • Strong knowledge of distributed systems, Docker, Kubernetes, and cloud infrastructure (AWS, GCP, Azure).
  • Hands-on experience with orchestration tools like Airflow, Kubeflow, or MLFlow.
  • Proven track record building scalable, low-latency inference systems.
  • Familiarity with cost optimization strategies for model inference (quantization, pruning, batching, caching).
  • Master’s or PhD in Computer Science, Machine Learning, or related field.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs