Build AI
ML Engineer, Inference Optimization
San Francisco
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Data & ML
- Seniority
- Mid level
- Stated salary
- $200,000 – $320,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Build AI
Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.
Job Summary
Inference is about 90% of compute spend. Economics are heavily driven by inference optimization. We’re hiring someone to make inference cheaper, faster, and good enough that we can scale the data engine and the product without the GPU bill eating the company.
Key Responsibilities
Own inference performance: latency, throughput, and cost per unit of work (tokens, frames, or jobs)
Cut the 90% compute line: kernels, batching, quantization, compilation, serving, and hardware utilization
Profile pipelines (Nsight, PyTorch Profiler, or equivalent), find the real bottleneck, and ship the fix
Work with research and product so models that are accurate are also affordable to run at scale
Build the serving and eval path so experiments don’t hide the inference bill
Measure cost as a first-class metric, not an afterthought once quality is “done”
You may be a good fit if you have (Must-have qualifications)
Strong ML / systems engineer with real inference optimization experience (serving, compilers, CUDA/kernels, quantization, or similar)
Comfortable in Python and in C++ or Rust for performance-critical paths
You think in dollars and tokens/frames per second, not only in accuracy tables
Familiarity with PyTorch (or JAX) and with profiling tools
Comfortable in a small research team shipping under cost pressure
Strong candidates may also have experience with (Nice-to-have qualifications)
CUDA, kernels, compilers (TVM, MLIR, TensorRT), or quantization in production
You have owned GPU/accelerator cost as a first-class metric
Serving stacks for video or large models
Understanding of memory hierarchy, data movement, and low-precision compute
Benefits
Competitive pay
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month for those living within walking distance of the office
Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
Unlimited Codex and Claude credits
Travel
How we're different
Build believes in the Bitter Lesson . By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: [email protected]
Similar jobs
- ML Engineer, ProductMach9 · San FranciscoFirst seen yesterday
- ML EngineerUniversalAGI · San FranciscoFirst seen yesterday
- AI and ML EngineerBah · 2 LocationsFirst seen today
- 2027 Internship Behavior ML Engineer, Learned Manipulation PoliciesBedrock Robotics · San Francisco, CAFirst seen today
- AI / ML Engineer – Advanced ConceptsBETA Technologies · Plattsburgh, New York, United States; Raleigh, North Carolina, United States; South Burlington, Vermont, United StatesFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job