Epsilon Health
Research Scientist - Post-training / RL
San Francisco, CA
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Data & ML
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 28 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Us
We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.
Role Overview
We're seeking a Research Scientist with deep expertise in post-training and reinforcement learning to join our ML Research team . You'll be at the forefront of developing and deploying state-of-the-art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine-tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool-use training, and inference-time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state-of-the-art in grounded report generation, reward design, and inference-time reasoning while maintaining the clinical rigor required for healthcare deployment.
Key Responsibilities
Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity-relation matching, grounding IoU, measurement accuracy, and reporting schema compliance.
Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines ( RLHF ) built on expert preferences and report edits.
Run GRPO-family algorithms with complex multi-reward objectives , tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss.
Train explicit reward models , including multimodal reward models conditioned on the image, with both outcome and process supervision.
Train chain-of-thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors.
Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi-turn trajectories.
Develop inference-time methods including best-of-N sampling against reward models and grounding-aware decoding, and distill the resulting gains back into the policy.
Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards.
Stay current with cutting-edge research in reinforcement learning, reward modeling, and multimodal post-training.
Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post-training medical VLMs at scale.
Qualifications
6+ years of academia/industry experience in reinforcement learning, post-training, or multimodal machine learning
Deep expertise in post-training large language or vision-language models (e.g., Qwen-VL, InternVL, LLaVA, or similar architectures)
Strong foundation in modern post-training and reinforcement learning techniques including:
Group-relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi-reward objectives
Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals
Reward model training: pairwise and generative reward models, outcome and process supervision
Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF
Inference-time compute scaling, including best-of-N sampling and verifier-guided decoding
Practical experience diagnosing and mitigating reward hacking and reward over-optimization
Track record of implementing complex models from research papers and adapting them to new domains
Proficiency in PyTorch or JAX, with experience training large models on multi-GPU/distributed systems
Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF
Experience with autoregressive language modeling and instruction tuning
Strong software engineering skills and ability to write production-quality code
Preferred Qualifications
Publications at top-tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)
Hands-on experience with medical imaging applications, particularly radiology report generation
Experience with agentic or multi-turn reinforcement learning, including credit assignment over tool-use trajectories
Experience with grounded generation tasks (visual grounding, referring expression comprehension)
Knowledge of evaluation methodologies for long-form generation, including factuality assessment and hallucination detection
Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning
Familiarity with clinical NLP and medical knowledge representation
Experience with model interpretability, explainability, and uncertainty quantification in safety-critical applications
Similar jobs
- Research Scientist - Bioanalytical LCMS Research and DevelopmentThermofisher · Richmond, Virginia, USAFirst seen today
- Assistant Research Scientist-Calcium Carbonate Digital TwinUmd · UMCES Horn Point LaboratoryFirst seen today
- Research Scientist, Fundamental Generative AI - New College Grad 2026Nvidia · US, CA, Santa ClaraFirst seen today
- Research Scientist level 3/4Ngc · United States-West Virginia-Rocket CenterFirst seen today
- Research ScientistJobgether · USFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job