TY - JOUR
T1 - Selective multi-scale depth estimation with liquid lens and scanning mirror
AU - Li, Chaohui
AU - Xing, Haoyue
AU - Zheng, Kun
AU - Yang, Wenshi
AU - Hao, Qun
AU - Cheng, Yang
N1 - Publisher Copyright:
© 2025 Elsevier Ltd
PY - 2026/2/24
Y1 - 2026/2/24
N2 - Monocular depth estimation has become a crucial technology in computer vision applications such as autonomous navigation, robotic manipulation, and augmented reality. However, current methods always face inherent limitations in balancing a wide field of view (FOV) coverage with high-resolution depth detail acquisition. To overcome the trade-off between global perception and local detail precision, a selective multi-scale depth estimation framework is proposed. It enables wide-area environmental understanding while adaptively acquiring high-resolution depth in regions of interest. The system integrates a wide-angle camera for global context with a narrow-angle camera equipped with a two-dimensional scanning mirror and liquid lens for high-resolution capture. The scanning mirror enables precise spatial control for region-of-interest selection, while the liquid lens provides rapid optical focus adaptation without mechanical movement. We develop a depth map fusion approach utilizing DepthAnything as the base model and adopting PatchFusion's strategy to combine global depth context with actively acquired high-resolution regional details. Experimental results demonstrate a significant enhancement in depth detail recovery for regions of interest. Compared to the purely passive wide-angle view, our fused depth maps achieve average improvements of 793% in sharpness and 1317% in detail texture. Our fusion algorithm further refines the actively captured high-resolution depth map, boosting sharpness and detail by an additional 10.6% and 10.1% respectively, while effectively suppressing noise and false edges. Ground-truth validation confirms the system's metric accuracy, reducing indoor RMSE by 22.0% and AbsRel by 21.3%, and lowering the mean relative error for long-range outdoor targets from 40.7% to 17.4%. The proposed framework enables accurate and scalable depth perception, supporting tasks like autonomous navigation, robotic manipulation, and augmented reality that require both global context and fine-grained local depth estimation.
AB - Monocular depth estimation has become a crucial technology in computer vision applications such as autonomous navigation, robotic manipulation, and augmented reality. However, current methods always face inherent limitations in balancing a wide field of view (FOV) coverage with high-resolution depth detail acquisition. To overcome the trade-off between global perception and local detail precision, a selective multi-scale depth estimation framework is proposed. It enables wide-area environmental understanding while adaptively acquiring high-resolution depth in regions of interest. The system integrates a wide-angle camera for global context with a narrow-angle camera equipped with a two-dimensional scanning mirror and liquid lens for high-resolution capture. The scanning mirror enables precise spatial control for region-of-interest selection, while the liquid lens provides rapid optical focus adaptation without mechanical movement. We develop a depth map fusion approach utilizing DepthAnything as the base model and adopting PatchFusion's strategy to combine global depth context with actively acquired high-resolution regional details. Experimental results demonstrate a significant enhancement in depth detail recovery for regions of interest. Compared to the purely passive wide-angle view, our fused depth maps achieve average improvements of 793% in sharpness and 1317% in detail texture. Our fusion algorithm further refines the actively captured high-resolution depth map, boosting sharpness and detail by an additional 10.6% and 10.1% respectively, while effectively suppressing noise and false edges. Ground-truth validation confirms the system's metric accuracy, reducing indoor RMSE by 22.0% and AbsRel by 21.3%, and lowering the mean relative error for long-range outdoor targets from 40.7% to 17.4%. The proposed framework enables accurate and scalable depth perception, supporting tasks like autonomous navigation, robotic manipulation, and augmented reality that require both global context and fine-grained local depth estimation.
KW - Active perception
KW - Depth fusion
KW - Liquid lens
KW - Monocular depth estimation
UR - https://www.scopus.com/pages/publications/105024722488
U2 - 10.1016/j.measurement.2025.120063
DO - 10.1016/j.measurement.2025.120063
M3 - Article
AN - SCOPUS:105024722488
SN - 0263-2241
VL - 262
JO - Measurement: Journal of the International Measurement Confederation
JF - Measurement: Journal of the International Measurement Confederation
M1 - 120063
ER -