← Back to jobs
San Diego, CA, USA
No related jobs found
Key Responsibilities
• Build and integrate evaluation harnesses and automation for software development use
cases, including turning real engineering artifacts like merged pull requests into repeatable
benchmark tasks.
• Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with
reproducible run environments (pinned dependencies, containerized runs, isolated worktrees)
so results stay comparable over time.
• Validate and calibrate evaluation approaches against human judgment, so scores are
consistent and correct rather than just repeatable.
• Support execution-based benchmarking across quality, productivity, and efficiency
measures, including cost and latency.
• Analyze results across repeated runs, looking at variance, failure patterns, and cost per
outcome, and find ways to make the workflows more reliable and more automated.
• Work with engineering and data teams to improve the tooling, and document how the
evaluations work and what they found for both technical and leadership audiences.
Required Skills & Experience
• Strong software engineering background, with real experience building automation,
developer tooling, or test and validation systems.
• Proficient in Python, and comfortable in at least one of Java, JavaScript, or a similar
language.
• Solid working knowledge of Git, including how branches, history, and working trees
behave, and of containerization with Docker.
• Experience with APIs, development environments, CI/CD pipelines, and typical
engineering workflows.
• Some familiarity with how AI, LLM, or agent evaluation works and where it goes wrong,
such as why a judge can be consistent but still wrong, why a single run can mislead, and how
benchmark contamination happens.
• Able to troubleshoot technical problems, think clearly about whether a measurement is
valid, and analyze results carefully.
Preferred Experience
• Hands-on work with AI-powered coding tools and agentic applications, such as Claude
Code, Devin, or OpenCode.
• Experience designing benchmarks or evaluations for software systems, especially
execution-based grading that verifies against tests.
• Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them
against human raters.
• Experience with build-system-aware test selection, such as Bazel or mapping changed
files to the tests that cover them.
• Experience building reproducible test environments and managing versioned evaluation
datasets.
• Comfortable writing up methodology and results for engineering leadership
Bachelor’s degree
No related jobs found
← Back to jobs