World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
arXiv:2609.29964v1 Announce Type: cross Abstract: General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than…