Research Scientist, Post-Training — Video Generation
At a glance
Mid-level Video Production role at Pika. US remote · full-time · $185,000–$400,000 base.
Growth Roles summary, based on the employer's posting.
What you'll do
- You'll run RL post-training for video diffusion and flow-matching models across multi-node systems
- You'll build reward models by defining targets, collecting preferences, and training learned judges
- You'll lead evaluation that compares human preference studies with automated measures
- You'll distill aligned models into efficient few-step samplers while preserving their gains
What you bring
- You have at least two years of hands-on research in post-training or generative modeling
- You've improved generative models through reinforcement learning or preference optimization
- You understand diffusion or flow-matching models, PyTorch, and distributed multi-node training
Who this fits
This role suits a research scientist focused on RL alignment and generative video modeling. It is framed as a staff- and lead-level opportunity and includes close work with engineering and product teams. The position is based in Palo Alto with a flexible on-site or remote hybrid arrangement.
From the employer
About the Role
At Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking Research Scientists with expertise in RL post-training and generative modeling for large-scale video generation. The focus is on refining Pika's video generation models using RL alignment and building robust video reward models. This is a staff and lead-level opportunity.
As a key member of our research team, you will own RL-based post-training for video diffusion/flow-matching models, develop state-of-the-art reward models, and lead post-training evaluation across human and automated metrics. You will collaborate closely with engineering and product teams, shaping the frontier of real-time creative and agentic video platforms.
Scope
RL alignment of Pika's video generation models and the reward models that drive them.
Distillation of RL-tuned models is a secondary focus.
Responsibilities
Run RL post-training (preference optimization, online RL against learned rewards) for video diffusion/flow-matching models at multi-node scale.
Build video reward models: define target evaluation dimensions, design/configure preference data collection workflows, train and validate learned judges, and safeguard against reward hacking.
Own post-training evaluation, including human preference studies and their correlation with automated metrics.
Distill RL-tuned models to efficient few-step samplers while preserving alignment gains (secondary focus).
What We’re Looking For
Required
2+ years hands-on research experience in post-training or generative modeling.
RL or preference-optimization experience on generative models with evidence of model improvement.
Strong grounding in diffusion or flow-matching models, PyTorch, and multi-node distributed training.
Preferred
Experience developing reward models for visual generation, including VLM-as-judge or large-scale preference data collection.
Distillation expertise (distribution matching, consistency, adversarial approaches), ideally for video models.
Familiarity with video-specific failure modes: temporal drift, motion and physics realism.
What We Offer
Competitive salary and substantial equity in a high-growth startup
Full health benefits + 401k matching and more
Collaborative, mission-driven team environment with significant growth opportunities
Flexible on-site/remote hybrid (HQ in Palo Alto, CA)
About Pika
Pika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us to shape the next evolution of creative technology!
If you are passionate about advancing RL alignment and generative modeling for video, and want to scale real-time multimodal foundation models, we want to hear from you.
Not this one either?
Claude or ChatGPT reads the other 90,537 for you.
Get better matches →