hirly

Epsilon Health

Software Engineer - ML Infrastructure

San Francisco, CA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Epsilon Health first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

Role Overview

We’re looking for an ML infrastructure engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Epsilon Health fast and reliable to ensure our research teams can focus on science rather than system bottlenecks.

Sitting in the Engineering team and working closely with research, you'll own the distributed training and reinforcement learning infrastructure our foundation-model and post-training work runs on, and the inference and evaluation systems that carry models from experimentation into production.

Key Responsibilities

Partner directly with researchers to deeply understand their workflows, then anticipate and design for how those needs will change

Build a distributed training infrastructure for foundation models on large-scale medical imaging, including the long-context parallelism and checkpointing that volumetric CT/MR training demands.

Build high-throughput data loading and preprocessing that keeps GPUs saturated on large volumetric and multimodal datasets.

Partner with researchers to prototype new ideas and translate them into production-ready code , owning end-to-end delivery from experimentation through deployment and monitoring.

Contribute to production serving and deployment pipelines (model rollout, canary deployments, and monitoring) alongside the backend team.

Build the reinforcement learning training stack (high-throughput rollout generation, reward-model serving, and experience collection), enabling the research team to run online, multi-reward RL at scale.

Qualifications

6+ years of experience designing, building, and operating large-scale distributed systems or infrastructure in production

Have 2+ years of experience building ML infrastructure or systems in production

Strong Python skills and expertise in PyTorch or JAX

Experience and familiarity with the compute, tooling, and workflow needs of large-scale machine learning research

Experience building infrastructure or platforms specifically for research or machine learning workflows

Deep experience building and operating Kubernetes and cloud infrastructure at scale

Experience with distributed training at scale (FSDP, DeepSpeed, or Megatron-style parallelism) and the systems concerns of keeping large GPU jobs efficient

Prior experience as a technical lead or mentor for other engineers

Preferred Qualifications

Experience operating in a startup or startup-like environment, i.e. a small, fast-moving team with high autonomy

Experience building reinforcement learning training infrastructure: rollout generation, reward-model serving, or online/off-policy learning systems

Experience with high-performance inference and serving (vLLM, SGLang, TensorRT, or Triton) for both training-time rollouts and production

Experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.

Experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.

Experience building internal training or experimentation platforms used by research teams, supporting A/B testing and experimentation workflows

Familiarity with vision-language models (VLMs) or multimodal architectures

Original posting on Epsilon Health's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Software Engineer – Epsilon Health | hirly.me