跳到主要导航 跳到搜索 跳到主要内容

SpikeCV: a survey on the hierarchical modeling of continuous-time spike representations and systematic progress in neuromorphic vision

投稿的翻译标题: SpikeCV脉冲视觉综述: 连续时间脉冲表征的层级建模与系统化进展
  • Yajing Zheng
  • , Rui Zhao
  • , Lin Zhu
  • , Yujia Liu
  • , Tiejun Huang*
  • *此作品的通讯作者
  • Peking University
  • Nanyang Technological University
  • Beijing Normal University

科研成果: 期刊稿件文章同行评审

摘要

With the rapid development of neuromorphic vision sensors, spike cameras have emerged as a promising paradigm for continuous-time visual perception. In contrast with conventional frame-based cameras that sample scenes at fixed frame rates, spike cameras encode luminance variations as asynchronous binary spike streams triggered by intensity accumulation at each pixel. This sensing mechanism changes the manner in which visual information is acquired and represented. Instead of producing discrete image frames, spike cameras generate continuous spike events that directly reflect temporal changes in scene intensity. Consequently, spike cameras provide several unique advantages, including extremely high temporal resolution, wide dynamic range, low motion blur, and sparse event-driven representations. These properties make spike cameras particularly suitable for challenging visual environments, such as high-speed motion analysis, extreme illumination conditions, and subtle temporal change detection, where traditional imaging systems frequently encounter limitations. Spike vision also introduces new challenges for visual computing. The statistical characteristics and data structures of spike streams differ significantly from those of conventional images or videos. Spike cameras produce sparse binary spike sequences in continuous time rather than dense intensity frames. Therefore, many established computer vision algorithms cannot be directly applied to spike data without modifications. Effective processing of spike streams requires new modeling strategies that explicitly consider the temporal dynamics and sparsity of signals. This survey reviews recent progress in spike vision research from the perspective of the hierarchical modeling of continuous-time spike representations. Existing methods are organized into several levels that reflect the progressive expansion of spike vision capabilities, ranging from signal modeling and reconstruction to semantic perception and system deployment. At the lowest level, the physically consistent modeling of spike generation and sensor noise provides a foundation for understanding spike data statistics. Studies in this direction analyze pixel triggering mechanisms, noise characteristics, and spike accumulation processes, forming the basis for reliable signal processing and algorithm design. Low-level visual reconstruction methods aim to recover stable visual signals from spike streams. Representative tasks include intensity reconstruction, high dynamic range imaging, motion deblurring, super-resolution, and low-light enhancement. These approaches convert spike sequences into interpretable intensity representations while preserving the temporal information contained in the spike data. The next level focuses on spatiotemporal modeling. The continuous-time nature of spike streams enables joint modeling of spatial structure and temporal motion. Research in this area addresses problems, such as optical flow estimation, motion segmentation, and dynamic scene analysis. Compared with frame-based methods, spike-based models provide improved temporal fidelity in fast motion scenarios. At the semantic perception level, spike representations are increasingly applied to tasks, such as object detection, recognition, and multi-object tracking. Continuous spike streams are integrated with deep neural networks, Transformer architectures, or spiking neural networks to perform higher-level visual reasoning. These methods utilize the temporal sparsity of spike data while maintaining low-latency processing. Spike cameras have also been introduced into 3D scene modeling. Recent studies combine spike streams with neural implicit representations to reconstruct static or dynamic 3D scenes. Continuous spike measurements provide detailed temporal information that can benefit dynamic scene reconstruction and neural rendering. System-level considerations play an important role in practical spike vision deployment. The evaluation of spike-based methods not only involves accuracy but also system metrics, such as latency, throughput, and energy consumption. These metrics become particularly important in real-time perception systems and edge computing scenarios. Progress in spike vision has also been supported by the development of datasets, simulation tools, and open-source platforms. Collecting spike camera data can be costly, and thus, spike simulators are widely used in algorithm development and validation. Simulation methods attempt to reproduce sensor physics and temporal spike generation processes. Public datasets and benchmarking protocols further support reproducible research. The SpikeCV platform provides a unified open-source framework for spike vision research. It integrates datasets, algorithm implementations, hardware interfaces, and evaluation tools, allowing researchers to prototype and evaluate spike-based algorithms rapidly. The platform has helped facilitate collaborative development and reproducible experiments within the community. Research activity in spike vision has grown rapidly in recent years. Publications, open-source resources, and benchmark datasets have increased steadily. Two international competitions organized in 2025 attracted wide participation from academic institutions and industry teams. These competitions encouraged standardized task definitions and stimulated methodological progress in the field. However, continuous-time representation learning for spike streams is still an active research area. Large-scale self-supervised learning for spike data remains largely unexplored. Multimodal fusion with complementary sensors introduces additional challenges into temporal alignment and noise modeling. System-level optimization that involves latency, throughput, and energy consumption also requires further investigation. Hardware-algorithm codesign with neuromorphic processors may provide new opportunities for efficient spike-based computation. Spike vision represents an emerging direction that reconsiders visual perception from a continuous-time perspective. Advances in sensor technology, representation learning, and system integration are gradually forming a new framework for visual computing. Continued progress in spike-based sensing, modeling, and deployment may enable high-speed, energy-efficient visual intelligence for future perception systems. Overall, this work provides a structured overview of spike vision research from the perspectives of sensing principles, representation modeling, perception algorithms, and system-level infrastructure. By organizing existing studies within a hierarchical framework of continuous-time spike representations, this study highlights how spike vision methods evolve from low-level signal reconstruction to high-level semantic perception and 3D scene understanding. The discussion of datasets, simulators, evaluation protocols, and open-source platforms further reflects the growing research ecosystem that surrounds spike-based vision. Through this synthesis, we aim to clarify the relationships among different research directions, summarize the current development status of the field, and provide a reference for future work on continuous-time visual computing and neuromorphic perception systems.

投稿的翻译标题SpikeCV脉冲视觉综述: 连续时间脉冲表征的层级建模与系统化进展
源语言英语
页(从-至)2045-2069
页数25
期刊Journal of Image and Graphics
31
6
DOI
出版状态已出版 - 2026
已对外发布

指纹

探究 'SpikeCV脉冲视觉综述: 连续时间脉冲表征的层级建模与系统化进展' 的科研主题。它们共同构成独一无二的指纹。

引用此