Autonomous driving paper index

Improved DPT-Hybrid for Monocular Depth Estimation with Geometry-Enhanced Encoding and Structure-Aware Gated Fusion

2026-08-05 · Electronics

autonomous drivingdepth estimationmonocular depthkittiperceptionpredictioncontrol

One-line summary

To address these issues, we propose a structure-aware enhanced DPT-Hybrid framework.

Engineering notes

Experiments on NYUv2 and KITTI demonstrate that the proposed method achieves lower single-run error metrics than the controlled DPT-Hybrid baseline under the evaluated settings. On NYUv2, our method achieves an absolute relative error (AbsRel) of 0.099 and an RMSE of 0.334.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。

Original abstract

Monocular depth estimation aims to recover dense 3D scene geometry from a single RGB image and plays an important role in autonomous driving, robotic perception, augmented reality, and 3D reconstruction. Although Transformer-based dense prediction models have achieved strong performance, existing DPT-Hybrid frameworks still suffer from three limitations: insufficient local geometric modeling in shallow stages, inadequate cross-scale fusion for preserving fine structures, and training objectives that only weakly constrain structural consistency. To address these issues, we propose a structure-aware enhanced DPT-Hybrid framework. First, a geometry-enhanced encoder introduces lightweight depth-wise separable convolution branches into shallow Transformer stages to better capture local edge and texture cues while preserving global contextual modeling. Second, a Structure-Aware Cross-Scale Gated Attention Fusion (S-GAF) module is proposed to improve decoder-side feature aggregation by jointly modeling channel-wise and spatial importance with an auxiliary RGB-gradient input. Third, joint structure–geometric consistency loss combines scale-invariant logarithmic loss, gradient consistency loss, and edge-focused loss to improve pixel-level accuracy, geometric plausibility, and boundary sharpness. Experiments on NYUv2 and KITTI demonstrate that the proposed method achieves lower single-run error metrics than the controlled DPT-Hybrid baseline under the evaluated settings. On NYUv2, our method achieves an absolute relative error (AbsRel) of 0.099 and an RMSE of 0.334. On KITTI, it achieves an AbsRel of 0.058 and an RMSE of 2.455. The proposed method introduces only modest additional complexity while producing more accurate and structurally sharper depth predictions.

5.0Engineering value
7.0Research novelty
5.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.

Request B2B research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment