Do big frontier models outperform narrow AI tools? Here we look into the healthcare domain, where a paper titled "General-purpose large language models outperform specialized clinical AI tools on medical benchmarks" was making headlines. The paper seemed to suggest that frontier LLMs from OpenAI and Gemini appeared to outperform more specialised clinical AI tools. However, there is more than meets the eye here. In this episode Dev and Doc deep dive into the fascinating debate of whether generalised LLMs do indeed outperform smaller specialised models.

👋 Hey! If you are enjoying our conversations, reach out, share your thoughts and journey with us. Don't forget to subscribe whilst you're here :)

2:25 intro - Does the convention still hold?
7:05 what about coding? How general is a task
8:38 data availability - Healthcare's data dichotomy
11:14 OpenEvidence Nature paper start
12:33 Context from a Doctor before AI
21:58 paper break down - benchmarks methodology
31:59 how to design clinical rubrics
34:22 OpenEvidence rebuttal
37:20 Conclusion & discussion
46:59 how to improve the research

To support us buymeacoffee.com/devanddoc

👨🏻‍⚕️Doc - Dr. Joshua Au Yeung - https://www.linkedin.com/in/dr-joshua-auyeung/
🤖Dev - Zeljko Kraljevic https://twitter.com/zeljkokr

YT - https://youtube.com/@DevAndDoc
Spotify - https://podcasters.spotify.com/pod/show/devanddoc
Apple - https://podcasts.apple.com/gb/podcast/dev-and-doc-ai-for-healthcare-podcast/id1751495120
Substack - https://aiforhealthcare.substack.com/

For enquiries - 📧Devanddoc@gmail.com

🎞️Editor- Dragan Kraljević https://www.instagram.com/dragan_kraljevic/
🎨Brand design and art direction - Ana Grigorovici https://www.behance.net/anagrigorovici027d

Podden och tillhörande omslagsbild på den här sidan tillhör Dev and Doc. Innehållet i podden är skapat av Dev and Doc och inte av, eller tillsammans med, Poddtoppen.