跳到主要导航 跳到搜索 跳到主要内容

How Vision-Language Tasks Benefit From Large Pre-Trained Models: A Survey

  • Yayun Qi
  • , Hongxi Li
  • , Yiqi Song
  • , Xinxiao Wu*
  • , Jiebo Luo
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Shenzhen MSU-BIT University
  • University of Rochester

科研成果: 期刊稿件文献综述同行评审

摘要

The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the attention of the research community. Despite improvements in overall performance, classic challenges still exist in vision-language tasks and hinder the development of this area. In recent years, the rise of pre-trained models is driving the research on vision-language tasks. Because of the massive scale of training data and model parameters, pre-trained models have exhibited excellent performance in numerous downstream tasks. Inspired by the powerful capabilities of pre-trained models, new paradigms have emerged to solve the classic challenges. Such methods have become mainstream in current research with increasing attention and rapid advances. In this paper, we present a comprehensive overview of how vision-language tasks benefit from pre-trained models. First, we review several main challenges in vision-language tasks and discuss the limitations of previous solutions before the era of pre-training. Next, we summarize the recent advances in incorporating pre-trained models to address the challenges in vision-language tasks. Finally, we analyze the potential risks associated with the inherent limitations of pre-trained models, discuss possible solutions, and attempt to provide future research directions.

源语言英语
页(从-至)1188-1210
页数23
期刊IEEE Transactions on Multimedia
28
DOI
出版状态已出版 - 2026
已对外发布

学术指纹

探究 'How Vision-Language Tasks Benefit From Large Pre-Trained Models: A Survey' 的科研主题。它们共同构成独一无二的学术指纹。

引用此