TY - JOUR
T1 - Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics
AU - Peng, Haosong
AU - Feng, Wei
AU - Li, Hao
AU - Zhan, Yufeng
AU - Jin, Ren
AU - Xia, Yuanqing
N1 - Publisher Copyright:
© 1968-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - The advent of edge computing has made real-time intelligent video analytics feasible. Previous works, based on traditional model architecture (e.g., CNN, RNN, etc.), employ various strategies to filter out non-region-of-interest content to minimize bandwidth and computation consumption but show inferior performance in adverse environments. Recently, visual foundation models based on transformers have shown great performance in adverse environments due to their amazing generalization capability. However, they require a large amount of computation power, which limits their applications in realtime intelligent video analytics. In this paper, we find visual foundation models like Vision Transformer (ViT) also have a dedicated acceleration mechanism for video analytics. To this end, we introduce Arena, an end-to-end edge-assisted video inference acceleration system based on ViT. We leverage the capability of ViT that can be accelerated through token pruning by only offloading and feeding Patches-of-Interest to the downstream models. Additionally, we design an adaptive keyframe inference switching algorithm tailored to different videos, capable of adapting to the current video content to jointly optimize accuracy and bandwidth. Through extensive experiments, our findings reveal that Arena can boost inference speeds by up to 1.58×, 1.82× and 1.98× on average while consuming only 47%, 31% and 27% of the bandwidth, respectively, all with high inference accuracy.
AB - The advent of edge computing has made real-time intelligent video analytics feasible. Previous works, based on traditional model architecture (e.g., CNN, RNN, etc.), employ various strategies to filter out non-region-of-interest content to minimize bandwidth and computation consumption but show inferior performance in adverse environments. Recently, visual foundation models based on transformers have shown great performance in adverse environments due to their amazing generalization capability. However, they require a large amount of computation power, which limits their applications in realtime intelligent video analytics. In this paper, we find visual foundation models like Vision Transformer (ViT) also have a dedicated acceleration mechanism for video analytics. To this end, we introduce Arena, an end-to-end edge-assisted video inference acceleration system based on ViT. We leverage the capability of ViT that can be accelerated through token pruning by only offloading and feeding Patches-of-Interest to the downstream models. Additionally, we design an adaptive keyframe inference switching algorithm tailored to different videos, capable of adapting to the current video content to jointly optimize accuracy and bandwidth. Through extensive experiments, our findings reveal that Arena can boost inference speeds by up to 1.58×, 1.82× and 1.98× on average while consuming only 47%, 31% and 27% of the bandwidth, respectively, all with high inference accuracy.
KW - Edge Computing
KW - Video Analytics System
KW - Vision Transformer
UR - https://www.scopus.com/pages/publications/105040946732
U2 - 10.1109/TC.2026.3699450
DO - 10.1109/TC.2026.3699450
M3 - Article
AN - SCOPUS:105040946732
SN - 0018-9340
JO - IEEE Transactions on Computers
JF - IEEE Transactions on Computers
ER -