Skip to main navigation Skip to search Skip to main content

厄米-高斯模式缺陷态空域并行检测与信息解码(特 邀)

Translated title of the contribution: Spatial Parallel Detection and Information Decoding of Defect States in Hermite-Gaussian Modes (Invited)
  • Lianghaoyue Zhang
  • , Yunfei Ma
  • , Yetong Hu
  • , Lingyu Kong
  • , Amna Ali
  • , Xiangyang Pan
  • , Xiaoyu Chang
  • , Feiyue Xia
  • , Haiyang Zhang
  • , Changming Zhao
  • , Zilong Zhang*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Ministry of Education in China
  • Ministry of Industry and Information Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Objective Free-space optical communication provides a promising technical route for high-speed, low-latency, and anti-electromagnetic-interference information transmission. However, conventional free-space optical links usually rely on temporal intensity modulation and symbol-by-symbol decision. Under atmospheric disturbance, platform jitter, pointing deviation, and geometric mismatch, the received optical field may suffer from random rotation, scale variation, intensity fluctuation, local distortion, and reduced signal-to-noise ratio, which restrict both information capacity and decoding reliability. Structured light offers additional spatial degrees of freedom for optical information transmission, but many existing mode-division or mode-keying schemes mainly improve the capacity by enlarging the recognizable mode set, making the receiver vulnerable to turbulence-induced mode coupling and classification errors. Hermite‒Gaussian (HG) modes have a regular Cartesian sub-spot lattice and are naturally compatible with pixelated image sensors. By introducing local defects into the HG lattice, the presence and absence of sub-spots can be used as binary occupancy states. This work aims to develop a spatial-domain parallel detection and information decoding method based on defective HG modes, so that multi-bit information can be recovered from a single intensity frame with improved robustness under complex receiving conditions. Methods In the proposed scheme, each defective HG mode is regarded as an independent spatial coding unit. A reference-assisted encoding strategy is adopted to reserve the leftmost column and the bottom row of the HG sub-spot lattice as reference structures for grid localization and geometric normalization, while the inner lattice region is used as the payload area. In this way, the defective HG mode can be represented as a two-dimensional binary occupancy matrix, and the payload matrix can be serialized into a one-dimensional bitstream in row-major order. At the receiver, a two-stage vision-based decoding framework is constructed. The first stage uses a rotation-aware object detector to locate multiple defective HG modes in a single-frame intensity image, estimate their orientations, and perform rotated region cropping and affine normalization. This step converts randomly distributed and rotated HG patterns into normalized regions of interest. The second stage detects small sub-spots inside each normalized HG region, maps the detected centroids onto a standard lattice, reconstructs the binary occupancy matrix, and finally outputs the decoded bitstream. To improve the recognition of weak, dense, and low-contrast sub-spots, the second-stage detector incorporates a VMamba backbone for long-range spatial dependency modeling, a high-frequency and spatial-aware feature pyramid network for multi-scale detail preservation, and a density-focused extractor for enhancing physically meaningful high-density optical regions. The dataset is constructed by combining experimental acquisition and numerical simulation, covering random rotation, scale variation, Gaussian noise, Poisson noise, background texture disturbance, and phase-screen-induced local distortion. Detection metrics, bit accuracy, bit error rate, and inference speed are jointly used to evaluate the decoding performance. Results and Discussions The first-stage detector shows stable global localization and geometric correction capability in both regular-array and random-tiling scenarios. As shown in Fig. 4, the detector can distinguish HG coding units under different target sizes, densities, and orientations, providing a unified coordinate basis for subsequent sub-spot decoding. Quantitatively, the first-stage rotation-aware detector achieves an average precision for IoU threshold of 0.50 (AP50) of 93.23%, an orientation estimation mean absolute error of 2.42°, and an inference speed of 112 frames per second. After global normalization, the second-stage detector is evaluated on the sub-spot recognition task. Compared with representative CNN- and Transformer-based detectors, the proposed model achieves better accuracy for weak and dense optical spots, with an average precision of 76.8%, an AP50 of 95.2%, an AP75 of 80.1%, and a small-object AP of 66.2%, while maintaining 10.33 million parameters and an inference speed of 92 frames/s. The qualitative comparison further explains the source of this improvement. YOLOv12-S tends to over-respond to background noise and introduces false positive spots, while RT-DETRv2-S may miss low-resolution and low-contrast sub-spots. In contrast, the proposed detector maintains a more complete occupancy structure and suppresses background-induced false responses. The ablation experiment confirms the contribution of each module. Replacing the baseline backbone with VMamba increases the average precision from 74.2% to 75.3%; introducing the density-focused extractor further improves it to 76.1%; and adding the high-frequency and spatial-aware feature pyramid network raises the final average precision to 76.8% and the small-object AP to 66.2%. More importantly, the improvement in small-spot detection is directly transferred to communication decoding reliability. With the same global geometric correction process, the proposed second-stage detector achieves a bit accuracy of 96.4% and a bit error rate of 0.036, outperforming Faster R-CNN, YOLOv12-S, and RT-DETRv2-S. The relationship among dataset complexity, accuracy-speed trade-off, ablation performance, and bit-level decoding quality is summarized in Fig. 6. For a high-order defective HG mode with 49 payload bits and a single frame containing up to 49 coding units, the effective spatial payload can reach 2401 bit/frame. At a cascaded decoding speed of approximately 92 frames per second, the corresponding theoretical transmission rate is about 220.9 kbit/s. Conclusions This work demonstrates that the spatial occupancy structure of defective HG modes can serve as a high-capacity information carrier for free-space optical communication. By combining reference-assisted occupancy encoding with two-stage vision-based decoding, the proposed method transforms the receiver-side decision process from temporal symbol judgment into spatial occupancy matrix recovery. The experimental results verify that the method can achieve robust multi-bit parallel decoding under random rotation, scale variation, low signal-to-noise ratio, and local distortion. This provides a feasible receiver-side solution for high-capacity, low-complexity, and robust spatial-domain free-space optical communication.

Translated title of the contributionSpatial Parallel Detection and Information Decoding of Defect States in Hermite-Gaussian Modes (Invited)
Original languageChinese (Traditional)
Article number1404019
JournalZhongguo Jiguang/Chinese Journal of Lasers
Volume53
Issue number14
DOIs
Publication statusPublished - Jul 2026
Externally publishedYes

Fingerprint

Dive into the research topics of 'Spatial Parallel Detection and Information Decoding of Defect States in Hermite-Gaussian Modes (Invited)'. Together they form a unique fingerprint.

Cite this