跳到主要导航 跳到搜索 跳到主要内容

VideoEvent: Hierarchical and adaptive event modelling for complex video understanding

  • Feng Zhang
  • , Xin Sun*
  • , Jianfei Zhao
  • , Yuming Shang
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Applications
  • Beijing University of Posts and Telecommunications
  • Ministry of Education in China

科研成果: 期刊稿件文章同行评审

摘要

While Large Language Models (LLMs) have accelerated the development of Video Large Language Models (VideoLLMs), existing methods struggle with complex video understanding. These models often treat videos holistically, resulting in coarse-grained spatiotemporal modelling and necessitating low sampling rates to manage computational overload. This approach leads to significant information loss and produces lower-quality visual tokens, rendering such models inadequate for tasks requiring detailed analysis or extended temporal understanding. In this paper, we propose VideoEvent, which reshapes video feature modelling from an event-centric perspective. We conceptualize a video as a collection of long-, medium-, and short-term events and introduce hierarchical encoding layers to model event features at these distinct time scales. Furthermore, we propose an event feature perceiver that adaptively filters and contextualizes these event features based on the user query. This combined hierarchical and adaptive framework strikes a critical balance between fine-grained modelling and computational efficiency. Consequently, VideoEvent generates a compact set of high-quality event features that capture the most relevant spatiotemporal information, enabling it to excel in complex scenarios, including fine-grained and long video understanding. Experiments show that VideoEvent improves accuracy by 3.7%, 6.9%, and 2.5% on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, respectively, and achieves state-of-the-art performance across all five evaluation metrics on VCGBench. Notably, it maintains low inference latency and memory usage, even when processing hour-long videos.

源语言英语
期刊论文编号116502
期刊Knowledge-Based Systems
350
DOI
出版状态已出版 - 27 9月 2026
已对外发布

学术指纹

探究 'VideoEvent: Hierarchical and adaptive event modelling for complex video understanding' 的科研主题。它们共同构成独一无二的学术指纹。

引用此