ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 13 days ago • 18
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 17 days ago • 49
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 13 days ago • 18
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation Paper • 2608.24138 • Published 26 days ago • 13
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published 30 days ago • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published 30 days ago • 4
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 176
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 91
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding Paper • 2605.19846 • Published May 20 • 3
Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them Paper • 2606.06361 • Published Jun 4 • 16
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Paper • 2502.05178 • Published Feb 7, 2025 • 10
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Paper • 2503.14734 • Published Mar 18, 2025 • 9
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models Paper • 2504.03624 • Published Apr 4, 2025 • 20