hirly

Npv

Multimodal ML Engineer

Paris

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Npv first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included
Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Country
FR
Work mode
Remote-friendly
First seen by hirly
9 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

We're looking for a Multimodal ML Engineer to join White Circle , an AI Safety company building the safety, reliability, and optimization layer for AI systems through natural-language policies it automatically tests, enforces, and improves at scale. Backed by $70M (Series A) from top funds and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, and others, White Circle processes 100M+ API calls monthly and fine-tunes and trains its own LLMs to run faster and cheaper than open or proprietary models.

You will

Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.

Design experiments, build multimodal data pipelines, and train MoE architectures.

Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.

Define evaluation metrics that actually matter for the product.

Requirements

3+ years training large-scale multimodal models.

Strong PyTorch and distributed training experience (DeepSpeed, FSDP).

Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.

Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).

Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.

Relocation to Paris or London (hybrid) required.

Bonus

Audio signal processing fundamentals – spectrograms, mel features, noise reduction.

MoE architecture experience.

We offer

Competitive salary + equity.

Official employment, visa and relocation help.

Find more English Speaking Jobs in France on Arbeitnow

Original posting on Npv's site ↗

Listed on hirly, a job board. hirly is not the employer: Npv is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job