BLUF:

  • To determine whether AI is ‘improving exponentially’, ‘hitting the wall’, or any other claim which involves a quantity or magnitude (e.g. ‘This model was a big leap/small increment’). We need a good y-axis: an interval scale of AI capability which means +1 unit always represents the same degree of ‘how much better’, in the same way +1 degree Celsius is always the same amount of ‘how much hotter’.
  • Yet there is no good y-axis for AI capability. All our measures are of something related-to but clearly not identical-with it, thus ‘true’ AI capability can be a funhouse-mirror reflection of whatever was measured. Specifically:
    • Benchmark score: One small step in benchmark score can be a giant leap in capability, or the opposite, or whatever else. (My 6/10 vs. your 4/10 ≠ I’m 50% better at maths than you).
    • Elo et al: Can give a real y-axis in terms of winning chances, but doesn’t translate outside of beating others. (Going from 50% to 73% to 88% chance to get a higher score than you on a maths test ≠ gaining 0 → 1 → 2 units of maths ability over you)
    • Epoch Capabilities Index: Analogous to IQ, so [...]

---

Outline:

(03:05) Introduction

(04:26) Both a poor reflection and a dark glass

(07:51) Human benchmarking also has a y-axis problem

(12:27) A metrological elegy

(12:49) The base case: benchmarks (cf. exams)

(13:30) Elo et al.

(16:25) (And maybe not quite 'game ability ≡ winning games', after all?)

(18:28) ECI (cf. IQ)

(20:56) Intervals Rarely True

(25:01) Measure endogeneity

(31:12) Forking IRT

(34:45) (Dimensions of being, and beating, a bat)

(39:40) Prediction (cf. chronometry)

(41:58) Time horizons

(44:34) Human "capability" is also exponential in time horizon

(47:38) Intuitive/interpretative prelude

(52:56) Time horizons and ECI share an axis kink

(57:13) Perhaps money, as a measure, stinks the least

(01:01:07) Finale: AI as normal epistemics

(01:06:48) Acknowledgements

The original text contained 33 footnotes which were omitted from this narration.

---

First published:
August 4th, 2026

Source:
https://forum.effectivealtruism.org/posts/CQvdadxjCpd7i7kjA/general-capability-and-capabilities-generally-have-no-good-y

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Scatter plot titledScatter plot titledLine graph comparingA line graph showing an S-shaped sigmoid curve on grid.Scatter plot titledScatter plot titledScatter plot titledThree line graphs comparing capability versus time with benchmark ranges.Scatter plot titledScatter plot titledScatter plot titledScatter plotScatter plot comparing log(TH) values against average score percentage with trend lines.Scatter plot comparing task duration versus average score for two thresholds.Scatter plot comparing log(TH) values against average score percentage with trend lines.Scatter plotLine graph titled

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör EA Forum Team. Innehållet i podden är skapat av EA Forum Team och inte av, eller tillsammans med, Poddtoppen.