Building Kanana-o: A Unified Multimodal Language Model that Sees, Hears, and Speaks

모드선택 :              
세미나 신청은 모드에서 세미나실 사용여부를 먼저 확인하세요

Building Kanana-o: A Unified Multimodal Language Model that Sees, Hears, and Speaks

홍영준 0 819
구분 응용수학
일정 2026-07-29(수) 14:00~15:30
세미나실 27동 116호
강연자 노병석 (카카오)
담당교수 홍영준
기타

Generative AI is rapidly evolving beyond text-only language models toward multimodal AI capable of understanding and generating information across multiple modalities, including images and speech. More recently, omni-modal language models have emerged as a promising direction for enabling AI systems that can naturally see, hear, and speak, bringing human-like interaction closer to reality.

In this talk, I will present the development journey of Kanana-o, Kakao’s unified multimodal language model, and share practical experiences in building large-scale multimodal foundation models for real-world applications. The talk begins with Kanana-v, a Vision Language Model, covering its architecture, training strategy, and key techniques such as high-resolution image processing, visual encoders, multimodal projectors, knowledge distillation, and reinforcement learning for visual reasoning. It then introduces Kanana-a, an Audio Language Model capable of both speech understanding and speech generation, discussing audio encoders, speech tokenization, streaming speech synthesis, and architectural innovations for reducing latency in real-time speech.


*본 강연은 카카오 AI 모델 카나나 책임자인 노병석 이사님이 최신 기술을 설명해주는 강연입니다. 산업계에 관심있는 학생 및 포닥분들 모두 참석 가능하니 부담없이 오셔서 들어주세요.

    정원 :
    부속시설 :
세미나명