Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 4 days ago • 27
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization Paper • 2609.11682 • Published 4 days ago • 16
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 3 days ago • 6
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Paper • 2609.12641 • Published 3 days ago • 25
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 5 days ago • 37
Memory as Plans: World-Action Modeling with Memory-Grounded Planning Paper • 2609.11561 • Published 4 days ago • 36
Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs Paper • 2609.11499 • Published 4 days ago • 28
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published 9 days ago • 59
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 4 days ago • 244
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 5 days ago • 302
PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents Paper • 2609.06702 • Published 8 days ago • 23
Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs Paper • 2609.10355 • Published 5 days ago • 12
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 4 days ago • 56
SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators Paper • 2609.09155 • Published 6 days ago • 16
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents Paper • 2609.08149 • Published 6 days ago • 25
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data Paper • 2609.05405 • Published 10 days ago • 38
SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research? Paper • 2609.09113 • Published 6 days ago • 20
Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails Paper • 2609.09134 • Published 6 days ago • 8