TY - JOUR
T1 - SpikeCV
T2 - a survey on the hierarchical modeling of continuous-time spike representations and systematic progress in neuromorphic vision
AU - Zheng, Yajing
AU - Zhao, Rui
AU - Zhu, Lin
AU - Liu, Yujia
AU - Huang, Tiejun
N1 - Publisher Copyright:
© 2026, Editorial and Publishing Board of JIG. All rights reserved.
PY - 2026
Y1 - 2026
N2 - With the rapid development of neuromorphic vision sensors, spike cameras have emerged as a promising paradigm for continuous-time visual perception. In contrast with conventional frame-based cameras that sample scenes at fixed frame rates, spike cameras encode luminance variations as asynchronous binary spike streams triggered by intensity accumulation at each pixel. This sensing mechanism changes the manner in which visual information is acquired and represented. Instead of producing discrete image frames, spike cameras generate continuous spike events that directly reflect temporal changes in scene intensity. Consequently, spike cameras provide several unique advantages, including extremely high temporal resolution, wide dynamic range, low motion blur, and sparse event-driven representations. These properties make spike cameras particularly suitable for challenging visual environments, such as high-speed motion analysis, extreme illumination conditions, and subtle temporal change detection, where traditional imaging systems frequently encounter limitations. Spike vision also introduces new challenges for visual computing. The statistical characteristics and data structures of spike streams differ significantly from those of conventional images or videos. Spike cameras produce sparse binary spike sequences in continuous time rather than dense intensity frames. Therefore, many established computer vision algorithms cannot be directly applied to spike data without modifications. Effective processing of spike streams requires new modeling strategies that explicitly consider the temporal dynamics and sparsity of signals. This survey reviews recent progress in spike vision research from the perspective of the hierarchical modeling of continuous-time spike representations. Existing methods are organized into several levels that reflect the progressive expansion of spike vision capabilities, ranging from signal modeling and reconstruction to semantic perception and system deployment. At the lowest level, the physically consistent modeling of spike generation and sensor noise provides a foundation for understanding spike data statistics. Studies in this direction analyze pixel triggering mechanisms, noise characteristics, and spike accumulation processes, forming the basis for reliable signal processing and algorithm design. Low-level visual reconstruction methods aim to recover stable visual signals from spike streams. Representative tasks include intensity reconstruction, high dynamic range imaging, motion deblurring, super-resolution, and low-light enhancement. These approaches convert spike sequences into interpretable intensity representations while preserving the temporal information contained in the spike data. The next level focuses on spatiotemporal modeling. The continuous-time nature of spike streams enables joint modeling of spatial structure and temporal motion. Research in this area addresses problems, such as optical flow estimation, motion segmentation, and dynamic scene analysis. Compared with frame-based methods, spike-based models provide improved temporal fidelity in fast motion scenarios. At the semantic perception level, spike representations are increasingly applied to tasks, such as object detection, recognition, and multi-object tracking. Continuous spike streams are integrated with deep neural networks, Transformer architectures, or spiking neural networks to perform higher-level visual reasoning. These methods utilize the temporal sparsity of spike data while maintaining low-latency processing. Spike cameras have also been introduced into 3D scene modeling. Recent studies combine spike streams with neural implicit representations to reconstruct static or dynamic 3D scenes. Continuous spike measurements provide detailed temporal information that can benefit dynamic scene reconstruction and neural rendering. System-level considerations play an important role in practical spike vision deployment. The evaluation of spike-based methods not only involves accuracy but also system metrics, such as latency, throughput, and energy consumption. These metrics become particularly important in real-time perception systems and edge computing scenarios. Progress in spike vision has also been supported by the development of datasets, simulation tools, and open-source platforms. Collecting spike camera data can be costly, and thus, spike simulators are widely used in algorithm development and validation. Simulation methods attempt to reproduce sensor physics and temporal spike generation processes. Public datasets and benchmarking protocols further support reproducible research. The SpikeCV platform provides a unified open-source framework for spike vision research. It integrates datasets, algorithm implementations, hardware interfaces, and evaluation tools, allowing researchers to prototype and evaluate spike-based algorithms rapidly. The platform has helped facilitate collaborative development and reproducible experiments within the community. Research activity in spike vision has grown rapidly in recent years. Publications, open-source resources, and benchmark datasets have increased steadily. Two international competitions organized in 2025 attracted wide participation from academic institutions and industry teams. These competitions encouraged standardized task definitions and stimulated methodological progress in the field. However, continuous-time representation learning for spike streams is still an active research area. Large-scale self-supervised learning for spike data remains largely unexplored. Multimodal fusion with complementary sensors introduces additional challenges into temporal alignment and noise modeling. System-level optimization that involves latency, throughput, and energy consumption also requires further investigation. Hardware-algorithm codesign with neuromorphic processors may provide new opportunities for efficient spike-based computation. Spike vision represents an emerging direction that reconsiders visual perception from a continuous-time perspective. Advances in sensor technology, representation learning, and system integration are gradually forming a new framework for visual computing. Continued progress in spike-based sensing, modeling, and deployment may enable high-speed, energy-efficient visual intelligence for future perception systems. Overall, this work provides a structured overview of spike vision research from the perspectives of sensing principles, representation modeling, perception algorithms, and system-level infrastructure. By organizing existing studies within a hierarchical framework of continuous-time spike representations, this study highlights how spike vision methods evolve from low-level signal reconstruction to high-level semantic perception and 3D scene understanding. The discussion of datasets, simulators, evaluation protocols, and open-source platforms further reflects the growing research ecosystem that surrounds spike-based vision. Through this synthesis, we aim to clarify the relationships among different research directions, summarize the current development status of the field, and provide a reference for future work on continuous-time visual computing and neuromorphic perception systems.
AB - With the rapid development of neuromorphic vision sensors, spike cameras have emerged as a promising paradigm for continuous-time visual perception. In contrast with conventional frame-based cameras that sample scenes at fixed frame rates, spike cameras encode luminance variations as asynchronous binary spike streams triggered by intensity accumulation at each pixel. This sensing mechanism changes the manner in which visual information is acquired and represented. Instead of producing discrete image frames, spike cameras generate continuous spike events that directly reflect temporal changes in scene intensity. Consequently, spike cameras provide several unique advantages, including extremely high temporal resolution, wide dynamic range, low motion blur, and sparse event-driven representations. These properties make spike cameras particularly suitable for challenging visual environments, such as high-speed motion analysis, extreme illumination conditions, and subtle temporal change detection, where traditional imaging systems frequently encounter limitations. Spike vision also introduces new challenges for visual computing. The statistical characteristics and data structures of spike streams differ significantly from those of conventional images or videos. Spike cameras produce sparse binary spike sequences in continuous time rather than dense intensity frames. Therefore, many established computer vision algorithms cannot be directly applied to spike data without modifications. Effective processing of spike streams requires new modeling strategies that explicitly consider the temporal dynamics and sparsity of signals. This survey reviews recent progress in spike vision research from the perspective of the hierarchical modeling of continuous-time spike representations. Existing methods are organized into several levels that reflect the progressive expansion of spike vision capabilities, ranging from signal modeling and reconstruction to semantic perception and system deployment. At the lowest level, the physically consistent modeling of spike generation and sensor noise provides a foundation for understanding spike data statistics. Studies in this direction analyze pixel triggering mechanisms, noise characteristics, and spike accumulation processes, forming the basis for reliable signal processing and algorithm design. Low-level visual reconstruction methods aim to recover stable visual signals from spike streams. Representative tasks include intensity reconstruction, high dynamic range imaging, motion deblurring, super-resolution, and low-light enhancement. These approaches convert spike sequences into interpretable intensity representations while preserving the temporal information contained in the spike data. The next level focuses on spatiotemporal modeling. The continuous-time nature of spike streams enables joint modeling of spatial structure and temporal motion. Research in this area addresses problems, such as optical flow estimation, motion segmentation, and dynamic scene analysis. Compared with frame-based methods, spike-based models provide improved temporal fidelity in fast motion scenarios. At the semantic perception level, spike representations are increasingly applied to tasks, such as object detection, recognition, and multi-object tracking. Continuous spike streams are integrated with deep neural networks, Transformer architectures, or spiking neural networks to perform higher-level visual reasoning. These methods utilize the temporal sparsity of spike data while maintaining low-latency processing. Spike cameras have also been introduced into 3D scene modeling. Recent studies combine spike streams with neural implicit representations to reconstruct static or dynamic 3D scenes. Continuous spike measurements provide detailed temporal information that can benefit dynamic scene reconstruction and neural rendering. System-level considerations play an important role in practical spike vision deployment. The evaluation of spike-based methods not only involves accuracy but also system metrics, such as latency, throughput, and energy consumption. These metrics become particularly important in real-time perception systems and edge computing scenarios. Progress in spike vision has also been supported by the development of datasets, simulation tools, and open-source platforms. Collecting spike camera data can be costly, and thus, spike simulators are widely used in algorithm development and validation. Simulation methods attempt to reproduce sensor physics and temporal spike generation processes. Public datasets and benchmarking protocols further support reproducible research. The SpikeCV platform provides a unified open-source framework for spike vision research. It integrates datasets, algorithm implementations, hardware interfaces, and evaluation tools, allowing researchers to prototype and evaluate spike-based algorithms rapidly. The platform has helped facilitate collaborative development and reproducible experiments within the community. Research activity in spike vision has grown rapidly in recent years. Publications, open-source resources, and benchmark datasets have increased steadily. Two international competitions organized in 2025 attracted wide participation from academic institutions and industry teams. These competitions encouraged standardized task definitions and stimulated methodological progress in the field. However, continuous-time representation learning for spike streams is still an active research area. Large-scale self-supervised learning for spike data remains largely unexplored. Multimodal fusion with complementary sensors introduces additional challenges into temporal alignment and noise modeling. System-level optimization that involves latency, throughput, and energy consumption also requires further investigation. Hardware-algorithm codesign with neuromorphic processors may provide new opportunities for efficient spike-based computation. Spike vision represents an emerging direction that reconsiders visual perception from a continuous-time perspective. Advances in sensor technology, representation learning, and system integration are gradually forming a new framework for visual computing. Continued progress in spike-based sensing, modeling, and deployment may enable high-speed, energy-efficient visual intelligence for future perception systems. Overall, this work provides a structured overview of spike vision research from the perspectives of sensing principles, representation modeling, perception algorithms, and system-level infrastructure. By organizing existing studies within a hierarchical framework of continuous-time spike representations, this study highlights how spike vision methods evolve from low-level signal reconstruction to high-level semantic perception and 3D scene understanding. The discussion of datasets, simulators, evaluation protocols, and open-source platforms further reflects the growing research ecosystem that surrounds spike-based vision. Through this synthesis, we aim to clarify the relationships among different research directions, summarize the current development status of the field, and provide a reference for future work on continuous-time visual computing and neuromorphic perception systems.
KW - SpikeCV
KW - continuous-time representation
KW - high-speed motion
KW - neuromorphic vision
KW - open-source ecosystem
KW - spatiotemporal modeling
KW - spike vision
UR - https://www.scopus.com/pages/publications/105042865056
U2 - 10.11834/jig.260128
DO - 10.11834/jig.260128
M3 - Article
AN - SCOPUS:105042865056
SN - 1006-8961
VL - 31
SP - 2045
EP - 2069
JO - Journal of Image and Graphics
JF - Journal of Image and Graphics
IS - 6
ER -