hirly

Moonlake

Member of Technical Staff — Agent Post-Training

San Francisco, CA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Moonlake first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Country
US
Work mode
Remote-friendly
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Introducing Moonlake, AI for creating world simulations.

About Moonlake

Moonlake is building the frontier of interactive world models: systems that generate, simulate, and reason over 3D environments for robotics, embodied AI, and interactive applications.

We develop the infrastructure that enables intelligent systems to learn, evaluate, and interact within realistic virtual environments before operating in the physical world.

Our work sits at the intersection of:

Robotics

Embodied AI

Interactive 3D Worlds

World Models

Simulation Infrastructure

Physical AI

Moonlake is building the next generation of AI infrastructure for interactive digital worlds. Our mission is to enable anyone to create, simulate, and interact with rich environments using natural language and multimodal inputs, turning simple ideas into worlds with structure, physics, and intelligent behavior.

Our team has raised $50M in seed funding from NVIDIA Ventures, Threshold Ventures, AIX Ventures, and notable angels including Naval Ravikant and Jeff Dean to build the foundational layer for the future of AI—powering everything from robotics training and simulation to digital twins and interactive environments.

We are looking for exceptional engineers to help build the simulation systems that will power the next generation of robotics and embodied intelligence.

The Role

We are hiring a Member of Technical Staff to lead reinforcement learning infrastructure and model post-training.

You will work closely with Qi and the research team to improve large vision-language and code-generating agents through fine-tuning, reinforcement learning, trajectory data, and scalable evaluation.

Moonlake already has deep expertise in 3D and world-building. This role adds the model-training experience needed to systematically improve agent performance and prepare the company for larger-scale RL across both digital and physical environments.

We are looking for a full-stack researcher and engineer who understands the complete training system and can make strong judgments about when training is necessary, which methods are likely to work, and what not to pursue.

What You’ll Do

Build RL and post-training pipelines for multimodal, vision-language, and code-generating agents

Develop infrastructure for supervised fine-tuning, preference optimization, reward modeling, and reinforcement learning

Create systems for collecting, filtering, replaying, and learning from agent trajectories

Design rewards, verifiers, and evaluations for long-horizon agent tasks

Improve agents’ ability to plan, write and execute code, use tools, recover from errors, and complete complex workflows

Scale distributed training and high-throughput rollout generation across multi-GPU environments

Improve training reliability, reproducibility, observability, and cost efficiency

Help define Moonlake’s long-term strategy for agent, robotics, and embodied-model training

What We’re Looking For

Real-world experience training large language, vision-language, multimodal, or code models

Strong experience in reinforcement learning, post-training, or large-scale fine-tuning

Experience building distributed training or high-throughput inference systems

Familiarity with supervised fine-tuning, preference optimization, reward modeling, and agentic RL

Experience with code-generation agents, long-horizon evaluation, or tool-using systems

Strong Python skills and experience with PyTorch, JAX, or similar frameworks

Ability to work across data, models, environments, rewards, evaluation, and infrastructure

Strong research judgment and a bias toward building reliable systems

Preferred Experience

Experience at a frontier AI lab or organization operating large-scale training systems

Experience with code-model post-training or autonomous coding agents

Experience with multimodal models, robotics, simulation, or embodied AI

Experience designing verifiable rewards or outcome-based training systems

Experience scaling RL workloads across large GPU clusters

What Success Looks Like

Within your first year, you will have:

Built Moonlake’s core post-training and RL infrastructure

Created scalable systems for learning from agent trajectories

Delivered measurable improvements in agent quality and task completion

Helped the team determine which problems require training and which do not

Established a reliable foundation for larger-scale agent and embodied-model training

Why This Role Matters

Moonlake’s agents must do more than generate content. They must understand complex requests, reason across vision and language, write and execute code, operate tools, build interactive worlds, and recover from mistakes.

This role will build the training systems that allow those agents to continuously improve.

We are committed to being an on-site, in-person team currently based in San Francisco.

Original posting on Moonlake's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Member of Technical Staff – Moonlake | hirly.me