GMorgulis/Phi-3-mini-4k-instruct-penguin_lora_sgd3e1-STEER0.584375-ft4.43 Updated about 5 hours ago • 1
Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts Paper • 2606.05922 • Published Jun 4 • 70
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages Paper • 2605.27901 • Published May 27 • 14
OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Paper • 2605.28158 • Published May 27 • 6
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws Paper • 2605.21803 • Published May 20 • 5
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL Paper • 2605.18703 • Published May 18 • 50