I make no claims to originality for any of this, but some people told me it'd be useful to write it up.
If an AI model acts smart on its training data, it'll usually keep acting pretty smart outside of its training data, unless you screw something up rather badly. I expect this fact to only become more true over time as the AIs we train become more and more capable.
I think many people have an intuition that the same is true of acting aligned. That if a model acts aligned with human values in training, it'll keep acting aligned with human values outside of training unless we screw something up rather badly, and that this will only become more true as the AIs we train become more and more capable, for all the same reasons that make this work with capabilities.
I think this is false. The inductive bias of neural network training toward simplicity that makes the property of 'acting smart' likely to generalise does not, to the same extent, make the property of 'acting aligned with human values' likely to generalise. The main blockers to AI alignment generalising aren't AIs overfitting to the training data [...]
---
Outline:
(01:30) General capabilities generally make the loss go down; alignment doesn't
(07:07) Smart agents pretty automatically self-correct their capabilities, but not their alignment
Podden och tillhörande omslagsbild på den här sidan tillhör
LessWrong. Innehållet i podden är skapat av LessWrong och inte av,
eller tillsammans med, Poddtoppen.