← All jobs
S/

Research Scientist, Video Foundation Models

Type
Full-time
Work setup
On-site
Experience
Mid
Posted
11 days ago
AIResearch
$200k - $320k
midpoint above market median

What research roles in crypto pay

45 salaries · our own data
this role$101kmedian$232k

This role pays $200k-$320k, above the $145k median for research roles in crypto on this board.

As a Research Scientist focused on Video Foundation Models at Cantina Labs, you will work on foundational research and large-scale model development for next-generation native video and omni foundation models. You'll shape both the technical direction and the team from an early stage, spanning the full model lifecycle including architecture, data, training, evaluation, post-training, training systems, and efficient inference.

What you'll do

  • Research, develop, and scale native video and multimodal foundation models, from early prototypes through large-scale pre-training, continued training, and post-training.
  • Explore new model architectures, training objectives, and conditioning mechanisms for video generation, reference- and memory-based generation, multimodal interaction, and joint audio-video generation.
  • Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines for high-quality and controllable generation.
  • Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency.
  • Collaborate closely with researchers, engineers, and product teams to help shape the technical roadmap and translate model advances into real-world capabilities.
  • Contribute to research publications and open-source releases when appropriate.

What you bring

  • Strong research and engineering experience in generative modeling, including diffusion models, flow matching, DiTs, video generation, multimodal models, world models, or related fields.
  • Hands-on experience training and evaluating large-scale image, video, or unified multimodal models using modern deep learning frameworks and distributed training systems.
  • A strong track record of developing impactful models or systems, demonstrated through research publications, open-source contributions, production impact, or other significant technical work.
  • Ability to independently own ambiguous research problems, move effectively from ideas to experiments, and work well in a highly collaborative environment.
  • Specialized depth in one or more areas across the foundation model lifecycle, such as model architecture, data curation, controllable generation, multimodal understanding and conditioning, post-training and reward modeling, model acceleration, inference systems, or deployment.

What we offer

  • Annual base salary of $200,000-$320,000
  • Generous company equity
  • Medical, dental, and vision insurance with 99.99% of premiums covered by Cantina
  • 42 days of paid time off: 15 PTO days, 10 sick days, 15 company holidays, and 2 floating holidays
  • Generous parental leave and fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account of $500 per month
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership

About Cantina Labs

Cantina Labs is a social AI company developing advanced real-time models that push the boundaries of expression, personality, and realism. The company brings characters to life and transforms how people tell stories, connect, and create through its flagship Cantina platform and broader ecosystem.

Research Scientist, Video Foundation Models | CryptoJobsHQ