Four LLM loss functions → four flavors of LLM misalignment
It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here’s the summary table, and then we’ll go through the rows separately. Training stage Loss function Flavor of misalignment[1] Famous examples Pretraining & SFT Imitative…
Sources
- T2Four LLM loss functions → four flavors of LLM misalignmentAI Alignment Forum / LessWrong (curated)