Learning Dynamic Evidence Routes for Vision Transformer Probing
arXiv:2605.00915v2 Announce Type: replace-cross Abstract: Probing frozen vision transformers typically uses permutation-invariant aggregation (GAP or $\texttt{[CLS]}$), treating patch tokens as an unstructured set. Content-dependent probes such as self-attention are useful accuracy controls, but…