Autonomous driving paper index

TempoCross: instance-aware sparse representation for multimodal temporal fusion in 3D detection

2026-07-18 · Scientific Reports

autonomous drivingbev3d detectionlidarnuscenesperception

One-line summary

In this work, we present TempoCross, a 3D detection method based on instance-aware sparse representations for multimodal temporal fusion.

Engineering notes

On the nuScenes test set, TempoCross achieves 74.1% mAP and 75.7% NDS, outperforming mainstream baseline detectors.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。

Original abstract

With rising performance demands for autonomous driving perception systems, bird's-eye-view (BEV) perception that combines heterogeneous sensor data and temporal information has become a key research focus. Recent studies have advanced either multimodal temporal detection through BEV-level feature aggregation or sparse object-centric temporal modeling. However, integrating these two directions while preserving modality-specific temporal states remains less explored for LiDAR-camera detection. In this work, we present TempoCross, a 3D detection method based on instance-aware sparse representations for multimodal temporal fusion. TempoCross encodes features from different timestamps and sensor modalities in a unified instance space, enabling adaptive extraction of critical target information from modality- and time-specific states. Initially, both branches perform a preliminary cross-modal fusion to generate queries for the current frame. In the motion compensation module, a hybrid motion modeling strategy reduces alignment discrepancies between historical and current instances caused by complex object motion. This strategy combines explicit rigid-body transformations with implicit learnable deformation residuals, improving both accuracy and robustness in cross-frame instance association. Next, temporal-aware enhancement integrates the initial queries and current-frame features with motion-compensated historical instances. Finally, a lightweight cross-attention fuses current and historical instances from both branches. This formulation reduces the reliance on repeatedly propagating full-scene fused BEV features and concentrates temporal interaction on target-related instance states. On the nuScenes test set, TempoCross achieves 74.1% mAP and 75.7% NDS, outperforming mainstream baseline detectors. The results support the effectiveness of combining LiDAR-camera fusion with instance-aware sparse temporal modeling.

5.5Engineering value
7.0Research novelty
5.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.

Request B2B research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment