BeyondQuality
Avsnitt

Episode 11: Can LLM Judges go beyond automated testing?

Dela

LLM outputs are hard to automatically test. You can’t use assertions for verifying a software that produces varied outputs for the same inputs that may nevertheless be acceptable.

Evaluation of AI systems using human domain experts might work for a few cases, but this doesn’t scale. We need automation of judgement and not merely string comparison. But once you do automate judgement at scale, can you go from verification to validation? Can you move from from testing an AI product to improving it?

This episode centers around Anupam’s enquiry on LLM judges. Hosts Maryia, Vitaly and Anupam discuss the two hypotheses he presented in this enquiry:

1. An LLM-Judge must be calibrated by one or more domain-experts for it to be useful

2. A calibrated LLM judge could go beyond being a means for catching defects, and could actually be used as a calibrated instrument for continuous improvement of an AI-System

They further talk about why calibration is needed, how teams should perform it, and what factors trigger recalibration of am LLM Judge.

Links:

- Discussion on BeyondQuality: https://github.com/BeyondQuality/beyondquality/discussions/40

- Recording of Anupam’s talk on LLM judge calibration: https://www.youtube.com/watch?v=E6B7lrrhrnM

- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena:

https://arxiv.org/pdf/2306.05685

Podden och tillhörande omslagsbild på den här sidan tillhör Vitaly Sharovatov. Innehållet i podden är skapat av Vitaly Sharovatov och inte av, eller tillsammans med, Poddtoppen.