It seems to me that, for every loss function that we use to train LLMs, we get a very distinct flavor of LLM misalignment. Here's the summary table, and then we’ll go through the rows separately.
Training stage
Loss function
Flavor of misalignment
Famous examples
Pretraining & SFT
Imitative learning (next-token prediction)
“Seven deadly sins” misalignment
Bing-Sydney, “Emergent misalignment”
RLHF & DPO
Human approval
“Glazing” misalignment
GPT-4o
RLVR
Automatic verifier
“Literal genie” misalignment
HuggingFace hacking
RLAIF
Approval from another LLM
“Trickster” misalignment
“Current AIs seem pretty misaligned to me”
Warning: I’m not an LLM power-user myself, but rather relying on reports I’ve read. Also, I don’t consider LLM alignment to be my primary area of expertise. I’m open to feedback!
In imitative learning, the LLM tries to predict what the next token of text will be. Then those predictions magically turn into its outputs. See my earlier discussion: “LLM pretraining magically transmutes observations into behavior, in a way that is profoundly disanalogous to how brains work”.
Podden och tillhörande omslagsbild på den här sidan tillhör
LessWrong. Innehållet i podden är skapat av LessWrong och inte av,
eller tillsammans med, Poddtoppen.