hirly

MakerMaker

RESEARCHER, EFFICIENT INFERENCE

San Francisco

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at MakerMaker first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run.

This is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).

WHAT YOU'LL DO

Research and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic

Design and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding

Investigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning

Run controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks

Co-design methods with the inference engineering team: push results to where they actually run, not stop at the paper

Read deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack

Publish when the work warrants it; share findings internally

Partner with model and training researchers so efficiency choices align with model architecture and post-training decisions

WHAT WE'RE LOOKING FOR

Strong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent

5+ years of hands-on research experience

Deep familiarity with both training and inference performance characteristics

Fluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving-framework level when methods require it

Track record of moving efficiency research from prototype to production

Strong statistical expertise: you'd notice a flawed comparison before someone else points it out

Strong written communication

Published research at NeurIPS, ICML, ICLR, MLSys, or comparable venues

NICE TO HAVE

PhD in ML, systems, or related field

Open-source contributions to quantization, speculative-decoding, or efficient-inference libraries

Experience with hardware-aware optimization and accelerator-specific tooling

Background in numerical methods, low-precision arithmetic, or

approximate computation

THIS ROLE IS PROBABLY NOT FOR YOU IF

You want to focus on pretraining large models from scratch (that's a different role)

You prefer abstract algorithmic research without hands-on implementation

You want a fixed benchmark with stable targets (our targets shift with what our models actually need to do)

Original posting on MakerMaker's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
RESEARCHER, EFFICIENT INFERENCE – MakerMaker | hirly.me