← Back to jobs

Thomson Reuters Logo
Applied Scientist

Thomson Reuters

 

Zug, Zug, Switzerland

Posted On: 1 day ago
Experience: 10+ years
Availability: Onsite
Openings: 1
Category: Applied Scientist
Tenure: No Preference/Any
Related Jobs

No related jobs found

Description

You will design and execute evaluation pipelines for LLMs and agentic systems, assessing reasoning, factual accuracy, and alignment.

This role is hybrid.

Responsibilities

  • Develop tools and frameworks for automatic evaluation, including synthetic dataset creation and LLM-as-a-judge workflows.
  • Partner with scientists, ML engineers, and product managers to translate evaluation results into model improvements.
  • Prototype new evaluation metrics and contribute to internal reports on evaluation methods.
  • Champion reproducibility, transparency, and ethical AI evaluation practices across the team.

Required Skills

  • PhD in Computer Science, AI, ML, or related field (Exceptional Master’s candidates considered).
  • 10+ years of relevant experience in research or hands-on work with large language models, NLP evaluation, or agent-based AI systems.
  • Strong understanding of LLM performance measurement, prompt evaluation, and reliability testing.
  • Proficiency in Python and familiarity with PyTorch, Transformers, and LangChain.
  • Experience with experimental design and data analysis.
  • Familiarity with cloud platforms including AWS, Azure, or GCP.

Preferred Skills

  • Experience with LLM evaluation frameworks (e.g., HELM, LM Harness, or custom auto-eval tools).
  • Familiarity with retrieval-augmented generation (RAG) or tool-using agents.
  • Record of publications in top-tier venues or equivalent research contributions.

Education

Any Graduate

Related Jobs

No related jobs found

← Back to jobs