hirly

Cantina

Research Scientist, Video Foundation Models

California

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Cantina first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Stated salary
$200,000 – $320,000 per year
Country
US
Work mode
Remote-friendly
First seen by hirly
11 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation.

In this role, you will work on foundational research and large-scale model development across the full model lifecycle, including architecture, data, training, evaluation, post-training, training systems, and efficient inference. You will have the opportunity to shape both the technical direction and the team from an early stage.

What You’ll Do:

Research, develop, and scale native video and multimodal foundation models, from early prototypes through large-scale pre-training, continued training, and post-training.

Explore new model architectures, training objectives, and conditioning mechanisms for video generation, reference- and memory-based generation, multimodal interaction, and joint audio-video generation.

Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines for high-quality and controllable generation.

Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency.

Collaborate closely with researchers, engineers, and product teams to help shape the technical roadmap and, where appropriate, translate model advances into real-world capabilities.

Contribute to research publications and open-source releases when appropriate.

What You’ll Bring:

Strong research and engineering experience in generative modeling, including areas such as diffusion models, flow matching, DiTs, video generation, multimodal models, world models, or related fields.

Hands-on experience training and evaluating large-scale image, video, or unified multimodal models using modern deep learning frameworks and distributed training systems.

A strong track record of developing impactful models or systems, demonstrated through research publications, open-source contributions, production impact, or other significant technical work.

Ability to independently own ambiguous research problems, move effectively from ideas to experiments, and work well in a highly collaborative environment.

Specialized depth in one or more areas across the foundation model lifecycle, such as model architecture, data curation, controllable generation, multimodal understanding and conditioning, post-training and reward modeling, model acceleration, inference systems, or deployment.

Compensation:

The anticipated annual base salary range for this role is between $200,000-$320,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:

Competitive salary and generous company equity

Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

42 days of paid time off, including:

15 PTO days

10 sick days

15 company holidays

2 floating holidays

Generous parental leave & fertility support

401(k) retirement savings plan

Lifestyle spending account – $500/month to use however you’d like

Complimentary lunch and snacks for in-office employees

One Medical membership, and more!

Original posting on Cantina's site ↗

Listed on hirly, a job board. hirly is not the employer: Cantina is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job