CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation Paper • 2609.04083 • Published 6 days ago • 25
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments Paper • 2609.00048 • Published 10 days ago • 8
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments Paper • 2608.24099 • Published 15 days ago • 12
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains Paper • 2608.09873 • Published about 1 month ago • 29
Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing Paper • 2606.05172 • Published Apr 16 • 2
VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding Paper • 2606.05259 • Published Jun 3 • 40
OpenComputer: Verifiable Software Worlds for Computer-Use Agents Paper • 2605.19769 • Published May 19 • 89
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems Paper • 2605.04018 • Published May 5 • 41
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction Paper • 2604.22880 • Published Apr 24 • 10
Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells Paper • 2603.25240 • Published Mar 26 • 79