TY - GEN
T1 - MoVi
T2 - 2026 IEEE Conference on Computer Communications, INFOCOM 2026
AU - Li, Zekai
AU - Fan, Xiaoyi
AU - Hu, Xiping
AU - Zhu, Yifei
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.
AB - Large multimodal models (LMMs) are transforming traditional mobile video analytics into interactive services, where wearable cameras stream live videos and users pose free-form queries and receive responses in natural language. Due to the high deployment costs, LMMs are primarily deployed on cloud servers. Real-time video streaming under dynamic network conditions and the afterward inference thus significantly affect quality of experience (QoE). However, existing video analytics systems are designed for task-specific, frame-independent, single-modal tasks, making them unsuitable for interaction with LMMs over multimodal, contextual inputs. Even worse, the autoregressive decoding of LMMs delays visual token processing during text generation, affecting real-time response. To bridge these gaps, we present MoVi, the first collaborative system for real-time LMM-driven interactive video analytics over dynamic networks. MoVi first establishes an interaction-oriented QoE model tailored to emerging interactive video analytics applications based on real-world user studies. It then jointly designs the video streaming and inference stages for QoE optimization. Specifically, MoVi employs a causal-aware streaming controller to adapt video configurations under dynamic networks. A query-assisted token manager further reduces response latency by dynamically pruning buffered tokens. Extensive experiments on real-world datasets and user studies demonstrate that MoVi achieves 40.1% higher QoE and 38.9% higher user opinion scores than existing state-of-the-art systems.
KW - edge-cloud collaboration
KW - large multimodal model
KW - video analytics
KW - video streaming
UR - https://www.scopus.com/pages/publications/105044492956
U2 - 10.1109/INFOCOM59046.2026.11571484
DO - 10.1109/INFOCOM59046.2026.11571484
M3 - Conference contribution
AN - SCOPUS:105044492956
T3 - Proceedings - IEEE INFOCOM
BT - INFOCOM 2026 - IEEE Conference on Computer Communications
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 May 2026 through 21 May 2026
ER -