StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 14 days ago • 445
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report Paper • 2608.15763 • Published 7 days ago • 43
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published 5 days ago • 60
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published 4 days ago • 23
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published 5 days ago • 41
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 5 days ago • 200
Fara-1.5: Scalable Learning Environments for Computer Use Agents Paper • 2606.20785 • Published Jun 18 • 9
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published 11 days ago • 24
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design Paper • 2608.13560 • Published 16 days ago • 55
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published 29 days ago • 113
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 22 days ago • 109
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 28 days ago • 263
Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design Paper • 2608.10299 • Published 19 days ago • 135
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published 17 days ago • 113
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 94
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 154
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning Paper • 2608.09888 • Published 19 days ago • 761
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 19 days ago • 342