Abstract
While Large Language Models (LLMs) have accelerated the development of Video Large Language Models (VideoLLMs), existing methods struggle with complex video understanding. These models often treat videos holistically, resulting in coarse-grained spatiotemporal modelling and necessitating low sampling rates to manage computational overload. This approach leads to significant information loss and produces lower-quality visual tokens, rendering such models inadequate for tasks requiring detailed analysis or extended temporal understanding. In this paper, we propose VideoEvent, which reshapes video feature modelling from an event-centric perspective. We conceptualize a video as a collection of long-, medium-, and short-term events and introduce hierarchical encoding layers to model event features at these distinct time scales. Furthermore, we propose an event feature perceiver that adaptively filters and contextualizes these event features based on the user query. This combined hierarchical and adaptive framework strikes a critical balance between fine-grained modelling and computational efficiency. Consequently, VideoEvent generates a compact set of high-quality event features that capture the most relevant spatiotemporal information, enabling it to excel in complex scenarios, including fine-grained and long video understanding. Experiments show that VideoEvent improves accuracy by 3.7%, 6.9%, and 2.5% on MSVD-QA, MSRVTT-QA, and ActivityNet-QA, respectively, and achieves state-of-the-art performance across all five evaluation metrics on VCGBench. Notably, it maintains low inference latency and memory usage, even when processing hour-long videos.
| Original language | English |
|---|---|
| Article number | 116502 |
| Journal | Knowledge-Based Systems |
| Volume | 350 |
| DOIs | |
| Publication status | Published - 27 Sept 2026 |
| Externally published | Yes |
Keywords
- Event features
- Large Language Model
- Multimodal
- Video understanding
Fingerprint
Dive into the research topics of 'VideoEvent: Hierarchical and adaptive event modelling for complex video understanding'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver