TY - JOUR
T1 - HMV-Net
T2 - Resource-efficient monocular depth estimation for UAV cargo delivery
AU - Luo, Bingxin
AU - Guo, Hongwei
AU - Wang, Wuhong
AU - Sun, Dongxian
N1 - Publisher Copyright:
Copyright © 2026. Published by Elsevier B.V.
PY - 2026/3
Y1 - 2026/3
N2 - Monocular depth estimation for UAV cargo delivery must balance limited onboard computational resources against high accuracy requirements in safety-critical near-field regions (0–30 meters). Inspired by human myopic vision principles, this work presents HMV-Net, a bio-inspired framework that reallocates computational resources toward operationally critical foreground regions without introducing inference overhead. The framework implements two key components: a training-time attention mechanism that imposes no deployment cost, and a parameter-free refinement module for identified critical regions. Compared with uniform processing baselines, HMV-Net improves depth accuracy by 1.6%, reduces RMSE by 2.1%, and enhances edge accuracy from 0.512 to 0.562, while maintaining identical parameter count (24.79 M) and throughput (187 vs 192 FPS on desktop GPU). Embedded deployment on Jetson Orin NX achieves over 33 FPS with 1.8–5.0% accuracy improvements across NYU-v2, KITTI, and COCO2017, demonstrating its potential for resource-constrained UAV platforms. Code is available at https://github.com/tian1bukewei/HMV-Net .
AB - Monocular depth estimation for UAV cargo delivery must balance limited onboard computational resources against high accuracy requirements in safety-critical near-field regions (0–30 meters). Inspired by human myopic vision principles, this work presents HMV-Net, a bio-inspired framework that reallocates computational resources toward operationally critical foreground regions without introducing inference overhead. The framework implements two key components: a training-time attention mechanism that imposes no deployment cost, and a parameter-free refinement module for identified critical regions. Compared with uniform processing baselines, HMV-Net improves depth accuracy by 1.6%, reduces RMSE by 2.1%, and enhances edge accuracy from 0.512 to 0.562, while maintaining identical parameter count (24.79 M) and throughput (187 vs 192 FPS on desktop GPU). Embedded deployment on Jetson Orin NX achieves over 33 FPS with 1.8–5.0% accuracy improvements across NYU-v2, KITTI, and COCO2017, demonstrating its potential for resource-constrained UAV platforms. Code is available at https://github.com/tian1bukewei/HMV-Net .
KW - Monocular depth estimation
KW - Resource-constrained deployment
KW - Selective attention mechanism
KW - UAV visual perception
UR - https://www.scopus.com/pages/publications/105028860927
U2 - 10.1016/j.rineng.2026.109192
DO - 10.1016/j.rineng.2026.109192
M3 - Article
AN - SCOPUS:105028860927
SN - 2590-1230
VL - 29
JO - Results in Engineering
JF - Results in Engineering
M1 - 109192
ER -