hirly

Build AI

Member of Technical Staff, Evals Lead

San Francisco

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Build AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Build AI

Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.

Job Summary

We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability. More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard.

Key Responsibilities

Design and run benchmarks for video and world models, with Build data and without it, under the same protocol

Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers

Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes

Work with research and dataset/quality so collection and evals inform each other

Push evals that are consequential enough that people change plans when the number moves

Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries

You may be a good fit if you have (Must-have qualifications)

You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions

You understand eval philosophy well enough to argue about validity, not only to plot a curve

Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain

You will not confuse a pretty dashboard with an eval that is allowed to decide things

Strong candidates may also have experience with (Nice-to-have qualifications)

Video, robotics, world models, or multimodal evals

You have designed evals used in a paper, a product launch, or a data decision

Background in construct validity, contamination, or leakage

Exposure to evaluation operations: throughput, failure taxonomy, repeatability

Benefits

Competitive pay

Medical, dental, and vision packages with generous premium coverage

$500 per month credit for waiving medical benefits

Housing subsidy of $2k per month for those living within walking distance of the office

Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)

Various wellness benefits covering fitness, mental health, and more

Daily lunch and dinner in our office

Unlimited compute budget subject to ROI justification

Unlimited Codex and Claude credits

Travel

How we're different

Build believes in the Bitter Lesson . By taking a general approach of learning from humans, our addressable market is all physical labor.

We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: [email protected]

Original posting on Build AI's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job