WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks

arXiv:2507.00938v3 Announce Type: replace-cross Abstract: Foundation models now enable autonomous agents to interact with real-world websites, but existing benchmarks emphasize general-purpose browsing, underrepresent research-oriented environments and scholarly discovery workflows, and often…

science

Sources