Subtitle: When is "increasing safety budget" a useful concept?
I often use what I’ll call the “safety-usefulness tradeoff model”, which is: developers face a tradeoff between “safety” and “usefulness” of an AI deployment, and the developer has only limited willingness or ability to sacrifice usefulness for the sake of safety. This model assumes that developers choose whether to take safety-relevant actions based on their cost efficiency, i.e., the marginal safety gain relative to the cost. However, that is not necessarily true. In this post, I spell out different stories for how developers choose what safety-relevant actions to take, in order to clarify when this model is relevant and how strategies for reducing AI risk are affected when its assumptions don’t hold.
The model suggests two ways a safety-concerned person can increase safety:
Safety tech improvements: push out the Pareto frontier, so that any given level of usefulness reduction buys more safety than it would have previously.
Safety budget increase: increase the extent to which the developer sacrifices usefulness for safety. On the cheaper end, this means implementing safety measures; on the more expensive end, it might mean refraining from training or deploying models whose risks [...]
---
Outline:
(04:17) Rushed reasonable developers
(06:08) Limited political will
(08:42) This model is unhelpful if developers don't trade efficiently between safety and usefulness
(12:59) Overall thoughts
(14:26) Appendix: Definitions of safety and usefulness in the rushed reasonable developer model
Podden och tillhörande omslagsbild på den här sidan tillhör
Redwood Research. Innehållet i podden är skapat av Redwood Research och inte av,
eller tillsammans med, Poddtoppen.