中国人民大学:2024年迈向可解释和可理解的多模态大规模语言模型(英文版)(33页).pdf
1、JOURNAL OF LATEX CLASS FILES,VOL.14,NO.8,OCTOBER 20241Towards Explainable and Interpretable MultimodalLarge Language Models:A Comprehensive SurveyYunkai Dang1,*Kaichen Huang1,*Jiahao Huo1,*Yibo Yan1,2Sirui Huang1Dongrui Liu3Mengxi Gao1Jie Zhang3Chen Qian3Kun Wang4Yong Liu5Jing Shao3Hui Xiong1,2Xumin
2、g Hu1,21The Hong Kong University of Science and Technology(Guangzhou)2The Hong Kong University of Science and Technology3Shanghai AI Laboratory4Nanyang Technological University5Renmin University of ChinaAbstractThe rapid development of Artificial Intelligence(AI)has revolutionized numerous fields,wi
3、th large languagemodels(LLMs)and computer vision(CV)systems drivingadvancements in natural language understanding and visualprocessing,respectively.The convergence of these technologieshas catalyzed the rise of multimodal AI,enabling richer,cross-modal understanding that spans text,vision,audio,and
4、videomodalities.Multimodal large language models(MLLMs),in par-ticular,have emerged as a powerful framework,demonstratingimpressive capabilities in tasks like image-text generation,visualquestion answering,and cross-modal retrieval.Despite theseadvancements,the complexity and scale of MLLMs introduc
5、e sig-nificant challenges in interpretability and explainability,essentialfor establishing transparency,trustworthiness,and reliability inhigh-stakes applications.This paper provides a comprehensivesurvey on the interpretability and explainability of MLLMs,proposing a novel framework that categorize
6、s existing researchacross three perspectives:(I)Data,(II)Model,(III)Training&Inference.We systematically analyze interpretability fromtoken-level to embedding-level representations,assess approachesrelated to both architecture analysis and design,and exploretraining and inference strategies that enh





点击查看更多