Training-Free Speech-Centric Omni Understanding with Frozen VLMs
Paper • 2609.04242 • Published • 5
Natural Language Processing, Machine Learning, and Computer Vision
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Training-Free Speech-Centric Omni Understanding with Frozen VLMs