← Back to jobs

Vaspire Technologies Inc Logo
AI Evaluation Engineer
Posted On: 8 days ago
Experience: 5+ years
Availability: Onsite
Openings: 1
Category: AI Evaluation Engineer
Tenure: Contract - Corp-to-Corp
Related Jobs

No related jobs found

Description

Key Responsibilities

• Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.

• Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.

• Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.

• Support execution-based benchmarking across quality, productivity, and e iciency measures, including cost and latency.

• Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.

• Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

͏

Required Skills & Experience 

• Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.

• Coding-agent evaluation: Evaluating AI coding agents that modify code repositories, including validating generated code changes against expected outcomes.

• Evaluation harnesses: Building automated, reproducible evaluation workflows, including test execution, environment setup, and result validation.

• Benchmarking and reliability: Establishing baselines, measuring run-to-run variance, analyzing failures, and ensuring consistent evaluation results.

• Git and CI/CD integration: Working with repository history, branches, pull requests, and automated testing within CI/CD pipelines.

• LLM-as-a-Judge: Using LLMs to evaluate coding-agent outputs, including calibration against human assessments.

• Proficient in at least one general-purpose language such as Python, Java, or JavaScript — the specific language background is flexible.

• Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.

• Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.

• Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.

• Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.

• Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively

Education

Bachelor's degree

Related Jobs

No related jobs found

← Back to jobs