Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Paper β’ 2606.09076 β’ Published Jun 8 β’ 66
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies Paper β’ 2606.06601 β’ Published Jun 4 β’ 26
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Paper β’ 2604.25819 β’ Published Apr 28 β’ 17 β’ 3
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Paper β’ 2604.25819 β’ Published Apr 28 β’ 17
Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation Paper β’ 2604.25819 β’ Published Apr 28 β’ 17
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment Paper β’ 2511.20614 β’ Published Nov 25, 2025 β’ 38
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs Paper β’ 2509.18056 β’ Published Sep 22, 2025 β’ 27
Emerging Properties in Unified Multimodal Pretraining Paper β’ 2505.14683 β’ Published May 20, 2025 β’ 136
Running on Zero Agents Featured 620 StoryDiffusion π 620 Generate consistent image sequences from text and photos
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation Paper β’ 2405.01434 β’ Published May 2, 2024 β’ 56
Running on Zero Agents Featured 620 StoryDiffusion π 620 Generate consistent image sequences from text and photos
Running on Zero Agents Featured 1.96k PhotoMaker π· 1.96k Generate personalized photos of a person from a prompt