Autonomous driving paper index
Object Detection and Scene Perception for Connected and Autonomous Vehicles Using LM-JEPA
One-line summary
This paper presents the latent model-joint embedding predictive architecture (LM-JEPA), a resource-efficient collaborative perception framework for connected and autonomous vehicles that integrates latent predictive representation learning with lightweight multi-modal reasoning.
Engineering notes
Key topics: autonomous driving, autonomous vehicle, object detection, lidar, sensor fusion, nuscenes, large language model, deployment, radar, perception. See the paper for implementation details and experimental results.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
This paper presents the latent model-joint embedding predictive architecture (LM-JEPA), a resource-efficient collaborative perception framework for connected and autonomous vehicles that integrates latent predictive representation learning with lightweight multi-modal reasoning. Autonomous driving in urban and highway environments requires accurate scene understanding under strict latency, energy, and communication constraints, limiting the practicality of large language model (LLM) and vision–language model (VLM)-based approaches in edge deployments. To address this, LM-JEPA encodes heterogeneous inputs including camera, LiDAR, radar, and map data into a unified latent space using a joint embedding predictive architecture, enabling efficient perception and reasoning without token-level inference. Unlike existing latent-space learning approaches that primarily learn predictive visual embeddings for single-modal perception, the proposed framework integrates multi-modal latent reasoning and adaptive sensor fusion to support collaborative perception under resource-constrained vehicular edge environments. The collaborative perception framework introduces a context-adaptive multi-modal fusion mechanism that dynamically weights sensor and model contributions, along with selective latent transmission and adaptive decoding for resource-aware operation. A lightweight VLM is integrated with an edge-assisted vehicular pipeline to support real-time on-vehicle inference with adaptive offloading based on latency and energy constraints, while a latent-space reasoning module enables cooperative decision-making. Experiments on BDD100K and nuScenes-QA show that LM-JEPA improves perception accuracy by 5% and reduces latency by approximately 7% over LLM and VLM baselines, while achieving up to 25% improvement in scene understanding, 20% higher intersection success rates, improved highway merging, and approximately 15% reduction in the transmitted model parameters.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments