TY - JOUR
T1 - Accelerating Training Convergence for Point Cloud Semantic Segmentation of Large-Scale Urban Scenes with Scene-Ensemble Prototypes
AU - Han, Jiawei
AU - Liu, Kaiqi
AU - Li, Wei
AU - Lin, Musen
AU - Li, Wei
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Point cloud semantic segmentation serves as a vital means of remote sensing, with the existing segmentation networks capable of achieving commendable results. However, the complex network architecture and extensive training data often demand substantial time and computational resources for model convergence. This study proposes a novel method to significantly expedite the convergence of point cloud semantic segmentation networks to save computational resources, which is called rapid convergence with scene-ensemble prototypes (RCSP). This method utilizes the knowledge of point clouds from the temporal ensemble of base segmentation network to supervise the training. The knowledge of the temporal-ensemble network is concretized as scene-ensemble prototypes and soft category prediction probabilities. It provides additional constraints beyond category labels, further narrowing the convergence direction of the segmentation network and reaching the optimal solution earlier. Experimental evaluations on large-scale urban scenes (Toronto-3D, SemanticKITTI, ISPRS) and urban indoor environments (S3DIS, ScanNet v2) demonstrate that RCSP achieves model convergence in approximately 30% of the iterations and 45% of the training time required by the baseline network under equivalent GPU memory constraints. Furthermore, the proposed framework delivers substantial improvements in segmentation performance over the baseline.
AB - Point cloud semantic segmentation serves as a vital means of remote sensing, with the existing segmentation networks capable of achieving commendable results. However, the complex network architecture and extensive training data often demand substantial time and computational resources for model convergence. This study proposes a novel method to significantly expedite the convergence of point cloud semantic segmentation networks to save computational resources, which is called rapid convergence with scene-ensemble prototypes (RCSP). This method utilizes the knowledge of point clouds from the temporal ensemble of base segmentation network to supervise the training. The knowledge of the temporal-ensemble network is concretized as scene-ensemble prototypes and soft category prediction probabilities. It provides additional constraints beyond category labels, further narrowing the convergence direction of the segmentation network and reaching the optimal solution earlier. Experimental evaluations on large-scale urban scenes (Toronto-3D, SemanticKITTI, ISPRS) and urban indoor environments (S3DIS, ScanNet v2) demonstrate that RCSP achieves model convergence in approximately 30% of the iterations and 45% of the training time required by the baseline network under equivalent GPU memory constraints. Furthermore, the proposed framework delivers substantial improvements in segmentation performance over the baseline.
KW - Point cloud semantic segmentation
KW - convergence
KW - large-scale urban scenes
KW - prototypes
UR - https://www.scopus.com/pages/publications/105029684766
U2 - 10.1109/TGRS.2026.3660139
DO - 10.1109/TGRS.2026.3660139
M3 - Article
AN - SCOPUS:105029684766
SN - 0196-2892
VL - 64
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
M1 - 5700814
ER -