TY - JOUR
T1 - VideoEvent
T2 - Hierarchical and adaptive event modelling for complex video understanding
AU - Zhang, Feng
AU - Sun, Xin
AU - Zhao, Jianfei
AU - Shang, Yuming
N1 - Publisher Copyright:
© 2026
PY - 2026/9/27
Y1 - 2026/9/27
N2 - While Large Language Models (LLMs) have accelerated the development of Video Large Language Models (VideoLLMs), existing methods struggle with complex video understanding. These models often treat videos holistically, resulting in coarse-grained spatiotemporal modelling and necessitating low sampling rates to manage computational overload. This approach leads to significant information loss and produces lower-quality visual tokens, rendering such models inadequate for tasks requiring detailed analysis or extended temporal understanding. In this paper, we propose VideoEvent, which reshapes video feature modelling from an event-centric perspective. We conceptualize a video as a collection of long-, medium-, and short-term events and introduce hierarchical encoding layers to model event features at these distinct time scales. Furthermore, we propose an event feature perceiver that adaptively filters and contextualizes these event features based on the user query. This combined hierarchical and adaptive framework strikes a critical balance between fine-grained modelling and computational efficiency. Consequently, VideoEvent generates a compact set of high-quality event features that capture the most relevant spatiotemporal information, enabling it to excel in complex scenarios, including fine-grained and long video understanding. Experiments show that VideoEvent improves accuracy by 3.7%, 6.9%, and 2.5% on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, respectively, and achieves state-of-the-art performance across all five evaluation metrics on VCGBench. Notably, it maintains low inference latency and memory usage, even when processing hour-long videos.
AB - While Large Language Models (LLMs) have accelerated the development of Video Large Language Models (VideoLLMs), existing methods struggle with complex video understanding. These models often treat videos holistically, resulting in coarse-grained spatiotemporal modelling and necessitating low sampling rates to manage computational overload. This approach leads to significant information loss and produces lower-quality visual tokens, rendering such models inadequate for tasks requiring detailed analysis or extended temporal understanding. In this paper, we propose VideoEvent, which reshapes video feature modelling from an event-centric perspective. We conceptualize a video as a collection of long-, medium-, and short-term events and introduce hierarchical encoding layers to model event features at these distinct time scales. Furthermore, we propose an event feature perceiver that adaptively filters and contextualizes these event features based on the user query. This combined hierarchical and adaptive framework strikes a critical balance between fine-grained modelling and computational efficiency. Consequently, VideoEvent generates a compact set of high-quality event features that capture the most relevant spatiotemporal information, enabling it to excel in complex scenarios, including fine-grained and long video understanding. Experiments show that VideoEvent improves accuracy by 3.7%, 6.9%, and 2.5% on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, respectively, and achieves state-of-the-art performance across all five evaluation metrics on VCGBench. Notably, it maintains low inference latency and memory usage, even when processing hour-long videos.
KW - Event features
KW - Large Language Model
KW - Multimodal
KW - Video understanding
UR - https://www.scopus.com/pages/publications/105043563618
U2 - 10.1016/j.knosys.2026.116502
DO - 10.1016/j.knosys.2026.116502
M3 - Article
AN - SCOPUS:105043563618
SN - 0950-7051
VL - 350
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116502
ER -