WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace
TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in…
Sources
- T2WorkspaceBench: Evaluating Interpretability Methods for the Global WorkspaceAI Alignment Forum / LessWrong (curated)