Abstract
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot planning. To address this challenge, we propose ShotPlan, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model. Our method introduces learnable planning tokens that capture shot-level transition cues and can be seamlessly integrated with the original video generation tokens to control transition timestamps. Unlike standard video generation tokens, the proposed planning tokens are equipped with Fractional Temporal Rotary Position Embedding (FRoPE), enabling shot transitions to be modeled at the frame level. Experiments demonstrate that ShotPlan significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.
Community
ShotPlan enables frame-accurate control over when shots change in text-to-video generation. A single learnable planning token, replicated per requested transition and placed at fractional temporal RoPE coordinates (FRoPE), tells a pre-trained DiT exactly where to cut without attention masks or architecture changes. Hard cuts land within <1 frame of the requested index, and the same mechanism handles soft transitions and temporally localized camera moves.Models and training data are available.
๐ Project page with 30+ video demos: https://pensioner-11.github.io/ShotPlan/
๐ค Models: https://huggingface.co/Pensioner/ShotPlan-Wan2.2-T2V-A14B-HighNoise
๐๏ธ Data: https://huggingface.co/datasets/Pensioner/shotplan
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- SmartDirector: Keyframe-Conditioned Cinematic Video Generation with Narrative Pacing Control (2026)
- CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation (2026)
- MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control (2026)
- CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation (2026)
- Memento: Reconstruct to Remember for Consistent Long Video Generation (2026)
- TriMotion: Modality-Agnostic Camera Control for Video Generation (2026)
- PE-Field 4D: Video Generation Models as Canvas (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.17675 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
Pensioner/ShotPlan-Wan2.1-T2V-14B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper