Mixture-of-Experts and Trends in Large-Scale Language Modeling with Irwan Bello - #569

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

コンテンツは TWIML and Sam Charrington によって提供されます。エピソード、グラフィック、ポッドキャストの説明を含むすべてのポッドキャストコンテンツは、TWIML and Sam Charrington またはそのポッドキャストプラットフォームパートナーによって直接アップロードされ、提供されます。誰かがあなたの著作権で保護された作品をあなたの許可なく使用していると思われる場合は、ここで概説されているプロセスに従うことができますhttps://ja.player.fm/legal。

2y ago 46:22

MP3•エピソードのホーム

Today we’re joined by Irwan Bello, formerly a research scientist at Google Brain, and now on the founding team at a stealth AI startup. We begin our conversation with an exploration of Irwan’s recent paper, Designing Effective Sparse Expert Models, which acts as a design guide for building sparse large language model architectures. We discuss mixture of experts as a technique, the scalability of this method, and it's applicability beyond NLP tasks the data sets this experiment was benchmarked against. We also explore Irwan’s interest in the research areas of alignment and retrieval, talking through interesting lines of work for each area including instruction tuning and direct alignment.

The complete show notes for this episode can be found at twimlai.com/go/569

700 つのエピソード