Autonomous driving paper index

<b>Computational Modeling of Soundscape and Visual Imagery in Beijing Old City Communities Based on Multi-modal Data</b>

2026-08-06 · Journal of Digital Frontier

autonomous drivingperceptionprediction

One-line summary

In this paper, we propose a joint embedding model driven by cross-modal contrastive learning to realize the dynamic coupling of acoustic physical-perceptual features and visual global-local semantic features.

Engineering notes

Based on 8,432 groups of measured multimodal data (including day-night differences) from four typical community types in Beijing’s old city, our experiments show that the proposed model achieves a cross-modal retrieval mAP of 76.82%, a perception score prediction RMSE as low as 0.318, and an intraclass correlation coefficient of 0.824 with manual scores, all of which are significantly better than those of the comparison methods (p < 0.001).

Chinese explanation / 中文解读

中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。

Original abstract

As an important carrier of the spirit of place, soundscape lacks systematic quantitative evaluation tools. The fundamental bottleneck lies in the lack of an accurate cross-modal alignment calculation framework between soundscape temporal sequence signals and visual image spatial semantics. In this paper, we propose a joint embedding model driven by cross-modal contrastive learning to realize the dynamic coupling of acoustic physical-perceptual features and visual global-local semantic features. We introduce a cross-modal attention forgetting gate to suppress modality-specific noise, design a hierarchical contrastive loss function to embed spatial structure with both instance-level and category-level constraints, and construct a spatio-temporal adaptive weight regulator to realize spatially differentiated modulation of the modality fusion coefficient. Based on 8,432 groups of measured multimodal data (including day-night differences) from four typical community types in Beijing’s old city, our experiments show that the proposed model achieves a cross-modal retrieval mAP of 76.82%, a perception score prediction RMSE as low as 0.318, and an intraclass correlation coefficient of 0.824 with manual scores, all of which are significantly better than those of the comparison methods (p < 0.001). Ablation experiments verify the independent contributions of the three innovations. Feature response analysis reveals that the partial correlation coefficient between soundscape roughness and visual building density exceeds 0.78, indicating a deep sensory coordination mechanism between auditory texture density and visual spatial envelopment perception. The proposed model provides an effective computational tool for quantitative diagnosis of multi-sensory quality in old city communities.

5.0Engineering value
7.0Research novelty
5.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.

Request B2B research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment