hirly

Heymilo

Research Engineer, Applied AI

San Francisco

Apply through hirly

hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
Remote-friendly
First seen by hirly
1 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

HeyMilo is building AI interviewers that automate and improve hiring through conversational AI. We work closely with companies to bring AI into real hiring workflows. We also run an applied AI team that studies where today's models succeed and fail across industries.

The Role

We're hiring a Research Engineer to join our Applied AI team in San Francisco. You'll design and build reinforcement learning environments with verifiable rewards for specific real-world use cases: the simulators, reward functions, and evaluation harnesses that let us measure and improve how models perform on real work.

The work is equal parts ML research and systems engineering. You'll work directly with the founders and our research advisors, taking a use case from problem definition to a reproducible environment that models can be evaluated and trained against.

What you'll do

Design and build RL environments for specific real-world use cases: realistic simulators, tool interfaces, and episodic task generation with proper isolation and reproducibility

Design verifiable reward functions that score correct intermediate actions as well as end states, and hold up against reward hacking

Build and maintain evaluation harnesses that run task suites across frontier and open-weight models, with clean scoring and cost tracking

Run post-training experiments (e.g. RLVR-style fine-tuning of open models) to validate that your environments produce a learnable signal

Package environments and results for reproducibility, and contribute to research write-ups and published evaluations

Help define which use cases we pursue next, informed by where models are weakest

What we're looking for

Master's or PhD in AI, Machine Learning, Computer Science, or a closely related field

Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation methodology)

Strong software engineering skills in Python, with code that others can run and build on

Hands-on experience with LLMs: running evaluations, building agentic loops, tool calling, fine-tuning

Comfortable with containers and infrastructure (Docker, Linux, cloud environments) for reproducible experiment setups

Ability to operate in ambiguity and move quickly; comfortable owning a problem end to end

Based in (or willing to relocate to) the San Francisco Bay Area

Bonus

Published research or open-source contributions in ML, RL, or evaluation

Experience with RL/eval frameworks and simulated or sandboxed environments

Experience training or fine-tuning open-weight models at any scale

Why join

Ground-floor role on a new applied AI team with real influence over how we build

Work directly with the founders and experienced research advisors, with your name on published work

High visibility, fast-paced, execution-driven environment

Competitive pay, equity, and benefits

Is this role actually a fit for you?

hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.

Score it against my resume
Research Engineer, Applied AI at Heymilo — hirly