Abstract
A hybrid multi-head attention transformer-CNN (AttnConvNeXt) model for computational ghost imaging (CGI) that can recognize targets both with and without images is presented in this paper. This unified architecture runs across several resolutions (128 × 128, 64 × 64, and 32 × 32) and directly analyzes raw 1D bucket measurements without reconstruction, in contrast to previous GI classifiers that were restricted to either reconstructed images or pre-processed signals. AttnConvNeXt achieves robust classification under low sampling ratios (SR = 0.8) where traditional approaches fail by combining multi-head attention with convolutional layers to capture both local features and global dependencies. Our model achieves 99%–100% recognition over resolutions when used for reconstructed images, providing a high-accuracy baseline. It outperforms a 12-layer CNN by 65% when processing solely on bucket signals in image-free mode achieving 84% accuracy at SR < 1. Real-time viability is demonstrated by the recognition time scaling effectively with resolution from 0.017 s/image (128 × 128) to 0.00056 s/signal (raw measurements). By creating the first multi-resolution, dual-mode GI recognition framework, to the best of our knowledge, this study removes the need for Fourier transforms for reconstruction-based recognition and makes it possible to use it for medical diagnostics, low-light surveillance, and scattering media.
| Original language | English |
|---|---|
| Pages (from-to) | 5564-5574 |
| Number of pages | 11 |
| Journal | Applied Optics |
| Volume | 65 |
| Issue number | 16 |
| DOIs | |
| Publication status | Published - 1 Jun 2026 |
Fingerprint
Dive into the research topics of 'Real-time target recognition with a hybrid multi-attention transformer-CNN (AttnConvNeXt) for computational ghost imaging at low sampling ratio'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver