Video generation has advanced rapidly for short clips, yet generating long, multi-shot videos that remain coherent, controllable, and reliable is still an open challenge. Across minutes of generation, current systems often suffer from identity drift, scene inconsistency, narrative breakdown, and weak responsiveness to user intent. These challenges make long-horizon video generation a compelling testbed for long-context multimodal modeling, structured generation, interactive systems, and evaluation.
F2S brings together researchers working on the core scientific and practical questions behind this transition from frames to stories. We are broadly interested in methods that maintain consistency over time, support richer forms of control and revision, and enable rigorous evaluation of long-form generation. The workshop welcomes work spanning model design, memory and state tracking, planning, editing, multimodal interaction, datasets, benchmarks, and real-world systems for long-horizon video creation.
Our goal is to foster a shared research agenda around reliable, controllable, and trustworthy long-horizon video generation, while creating space for perspectives from generative modeling, multimodal learning, interactive machine learning, and creative applications.
We invite submissions on all aspects of long-horizon video generation, with a focus on reliability, controllability, and evaluation. Topics include but are not limited to:
Submission URL: OpenReview
Format: All submissions must be in PDF format and anonymized. Submissions are limited to four content pages, including all figures and tables; unlimited additional pages containing references and supplementary materials are allowed. Reviewers may choose to read the supplementary materials but will not be required to. Camera-ready versions may go up to five content pages.
Style file: You must format your submission using the ICML 2026 LaTeX style file. Please include the references and supplementary materials in the same PDF as the main paper. The maximum file size for submissions is 50MB. Submissions that violate the ICML style (e.g., by decreasing margins or font sizes) or page limits may be rejected without further review.
Dual-submission policy: We welcome ongoing and unpublished work. We will also accept papers that are under review at the time of submission, or that have been recently accepted without published proceedings.
Non-archival: The workshop is a non-archival venue and will not have official proceedings. Workshop submissions can be subsequently or concurrently submitted to other venues.
Visibility: Submissions and reviews will not be public. Only accepted papers will be made public.
| Submission deadline | May 10, 2026 (AoE) |
| Notification to authors | May 15, 2026 (AoE) |
| Camera-ready deadline | June 8, 2026 (AoE) |
| Workshop date | July 10, 2026 (Friday; Seoul, South Korea) |
All times are in Korea Standard Time (KST, GMT+9).
| Time (KST) | Session | Speaker / Affiliation |
|---|---|---|
| 08:00 – 08:10 | Opening | Organizers |
| 08:10 – 08:55 | Invited Talk 1 / Keynote | Vincent Sitzmann, MIT |
| 08:55 – 09:25 | Invited Talk 2 | Yukang Chen, NVIDIA Research |
| 09:25 – 09:45 | Coffee Break | — |
| 09:45 – 10:15 | Invited Talk 3 | Alexandre Alahi, EPFL |
| 10:15 – 10:45 | Invited Talk 4 | Bohyung Han, Seoul National University |
| 10:45 – 11:15 | Invited Talk 5 | Xihui Liu, The University of Hong Kong |
| 11:15 – 12:00 | Oral Presentations I 3 papers, 15 min each | Selected Papers 1–3 |
| 12:00 – 12:10 | Poster Setup / Lunch Transition | — |
| 12:10 – 13:40 | Lunch | — |
| 13:40 – 14:10 | Invited Talk 6 | Pinar Yanardag, Virginia Tech |
| 14:10 – 14:40 | Oral Presentations II 2 papers, 15 min each | Selected Papers 4–5 |
| 14:40 – 15:50 | Poster Session | Accepted papers |
| 15:50 – 16:05 | Coffee Break | — |
| 16:05 – 16:35 | Invited Talk 7 | Daquan Zhou, Peking University |
| 16:35 – 17:00 | Best Paper Award & Closing Remarks | Organizers |
Cantina Labs builds lifelike AI characters and the video generation models behind them. They are growing their global research team, especially their Singapore lab, to work on consistency, control, and trustworthy evaluation for long-horizon generation.
Featured hiring areas include research and engineering roles across video generation, image models, memory systems, model evaluation, and AI product engineering.
Contact (Google Group): f2s-workshop@googlegroups.com