Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

arXiv:2609.00949v2 Announce Type: replace-cross Abstract: Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However…

aiscience

Sources