Autonomous driving paper index

Li-ViP3D++: Query-Gated Deformable Camera–LiDAR Fusion for End-to-End Perception and Trajectory Prediction

2026-07-22 · KITopen

autonomous drivingbevend-to-endtrajectory predictiontrajectory forecastinglidarnusceneshd mapperceptionprediction

One-line summary

We propose Li-ViP3D++, a query-based multimodal PnP framework that introduces Query-Gated Deformable Fusion (QGDF) to integrate multi-view RGB and LiDAR in query space.

Engineering notes

Key topics: autonomous driving, bev, end-to-end, trajectory prediction, trajectory forecasting, lidar, nuscenes, hd map, perception, prediction. See the paper for implementation details and experimental results.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。

Original abstract

End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving. Modular pipelines restrict information flow and can amplify upstream errors. Recent query-based, fully differentiable perception-and-prediction (PnP) models mitigate these issues, yet the complementarity of cameras and LiDAR in the query-space has not been sufficiently explored. Models often rely on fusion schemes that introduce heuristic alignment and discrete selection steps which prevent full utilization of available information and can introduce unwanted bias. We propose Li-ViP3D++, a query-based multimodal PnP framework that introduces Query-Gated Deformable Fusion (QGDF) to integrate multi-view RGB and LiDAR in query space. QGDF 1) aggregates image evidence via masked attention across cameras and feature levels, 2) extracts LiDAR context through fully differentiable BEV sampling with learned per-query offsets, and 3) applies query-conditioned gating to adaptively weight visual and geometric cues per agent. The resulting architecture jointly optimizes detection, tracking, and multi-hypothesis trajectory forecasting in a single end-to-end model. On nuScenes, Li-ViP3D++ improves end-to-end behavior and detection quality, achieving higher EPA (0.505) and mAP (0.616) while substantially reducing false positives (FP ratio 0.069), and it is faster than the prior Li-ViP3D variant (139.82 ms vs. 145.91 ms). Additional experiments were performed to evaluate the impact of the reduced resolution of RGB inputs and missing HD maps on the behavior of the model. The results of the experiments indicate that query-space, fully differentiable camera–LiDAR fusion can increase the robustness of end-to-end PnP without sacrificing deployability.

6.0Engineering value
7.0Research novelty
5.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.

Request B2B research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment