跳到主要导航 跳到搜索 跳到主要内容

Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding

  • Beijing Institute of Technology
  • Beijing University of Technology

科研成果: 期刊稿件文章同行评审

摘要

Existing video-language large models (video-LLMs) have made notable progress in understanding global video content but still struggle to capture fine-grained temporal dynamics and accurately align visual events with textual descriptions. In this work, we introduce DAPE-VLLM (Density-Aware Position Encoding enhanced Video-Language Large Model), a time-perception enhanced video grounding framework that integrates density-aware position encoding with boundary perception, to address these limitations and bridge the temporal discrepancy between video and language modalities. We propose a suite of boundary-aware tasks that explicitly model event durations, locations, and transitions, enabling large language models to better understand temporal structure within videos. In addition, we develop a novel density-aware temporal encoding mechanism that dynamically adapts position representations to varying temporal information density, enhancing the model’s ability to localize events precisely across diverse video content. Extensive experiments across ActivityNet, Charades, and DiDeMo datasets demonstrate the effectiveness of our approach, yielding consistent performance improvements and achieving up to 11.2% relative gain in R@0.3. These results validate the importance of integrating density-aware modeling and boundary perception to advance fine-grained video grounding capabilities in large language models.

源语言英语
期刊IEEE Transactions on Circuits and Systems for Video Technology
DOI
出版状态已接受/待刊 - 2026
已对外发布

学术指纹

探究 'Time-Perception Enhanced Video Grounding Framework with Density-Aware Position Encoding' 的科研主题。它们共同构成独一无二的学术指纹。

引用此