Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons

arXiv:2609.29781v1 Announce Type: new Abstract: Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and…

aiscience

Sources