Skip to main navigation Skip to search Skip to main content

VideoEvent: Hierarchical and adaptive event modelling for complex video understanding

  • Feng Zhang
  • , Xin Sun*
  • , Jianfei Zhao
  • , Yuming Shang
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Beijing Engineering Research Center of High Volume Language Information Processing and Cloud Computing Applications
  • Beijing University of Posts and Telecommunications
  • Ministry of Education in China

Research output: Contribution to journalArticlepeer-review

Abstract

While Large Language Models (LLMs) have accelerated the development of Video Large Language Models (VideoLLMs), existing methods struggle with complex video understanding. These models often treat videos holistically, resulting in coarse-grained spatiotemporal modelling and necessitating low sampling rates to manage computational overload. This approach leads to significant information loss and produces lower-quality visual tokens, rendering such models inadequate for tasks requiring detailed analysis or extended temporal understanding. In this paper, we propose VideoEvent, which reshapes video feature modelling from an event-centric perspective. We conceptualize a video as a collection of long-, medium-, and short-term events and introduce hierarchical encoding layers to model event features at these distinct time scales. Furthermore, we propose an event feature perceiver that adaptively filters and contextualizes these event features based on the user query. This combined hierarchical and adaptive framework strikes a critical balance between fine-grained modelling and computational efficiency. Consequently, VideoEvent generates a compact set of high-quality event features that capture the most relevant spatiotemporal information, enabling it to excel in complex scenarios, including fine-grained and long video understanding. Experiments show that VideoEvent improves accuracy by 3.7%, 6.9%, and 2.5% on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, respectively, and achieves state-of-the-art performance across all five evaluation metrics on VCGBench. Notably, it maintains low inference latency and memory usage, even when processing hour-long videos.

Original languageEnglish
Article number116502
JournalKnowledge-Based Systems
Volume350
DOIs
Publication statusPublished - 27 Sept 2026
Externally publishedYes

Keywords

  • Event features
  • Large Language Model
  • Multimodal
  • Video understanding

Fingerprint

Dive into the research topics of 'VideoEvent: Hierarchical and adaptive event modelling for complex video understanding'. Together they form a unique fingerprint.

Cite this