IDRBench: Benchmarking the Interactive Capabilities of Deep Research Agents
arXiv:2601.06676v3 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based deep research agents perform multi-step reasoning, web exploration, and long-form report generation. In these long-horizon workflows, early deviations from user intent can misdirect research and propagate…