Building Kanana-o: A Unified Multimodal Language Model that Sees, Hears, and Speaks
| 구분 | 응용수학 |
|---|---|
| 일정 | 2026-07-29(수) 14:00~15:30 |
| 세미나실 | 27동 116호 |
| 강연자 | 노병석 (카카오) |
| 담당교수 | 홍영준 |
| 기타 |
Generative AI is rapidly evolving beyond text-only language models toward multimodal AI capable of understanding and generating information across multiple modalities, including images and speech. More recently, omni-modal language models have emerged as a promising direction for enabling AI systems that can naturally see, hear, and speak, bringing human-like interaction closer to reality.
In this talk, I will present the development journey of Kanana-o, Kakao’s unified multimodal language model, and share practical experiences in building large-scale multimodal foundation models for real-world applications. The talk begins with Kanana-v, a Vision Language Model, covering its architecture, training strategy, and key techniques such as high-resolution image processing, visual encoders, multimodal projectors, knowledge distillation, and reinforcement learning for visual reasoning. It then introduces Kanana-a, an Audio Language Model capable of both speech understanding and speech generation, discussing audio encoders, speech tokenization, streaming speech synthesis, and architectural innovations for reducing latency in real-time speech.
*본 강연은 카카오 AI 모델 카나나 책임자인 노병석 이사님이 최신 기술을 설명해주는 강연입니다. 산업계에 관심있는 학생 및 포닥분들 모두 참석 가능하니 부담없이 오셔서 들어주세요.