Autonomous driving paper index
Bus-Mounted Vision Sensing for Traffic Object Detection: BFTD and a Local–Global Attention Framework
One-line summary
An autonomous driving research paper: Bus-Mounted Vision Sensing for Traffic Object Detection: BFTD and a Local–Global Attention Framework.
Engineering notes
To support this sensing scenario while avoiding ambiguity with previously used dataset acronyms, we construct the Bus Front-view Traffic Dataset (BFTD), a high-resolution benchmark collected from forward-facing cameras mounted on multiple buses operating on urban routes during real-world service. Extensive experiments on BFTD and public benchmarks show that the proposed framework improves detection accuracy, particularly for small and visually crowded traffic participants, while maintaining a practical accuracy–efficiency trade-off.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
Bus-mounted vision sensing provides a practical and complementary perspective for intelligent transportation systems, but reliable traffic object detection from bus front-view cameras remains challenging because elevated viewpoints induce severe scale skewness, dense interactions around bus stops and intersections, and frequent heterogeneous occlusion. To support this sensing scenario while avoiding ambiguity with previously used dataset acronyms, we construct the Bus Front-view Traffic Dataset (BFTD), a high-resolution benchmark collected from forward-facing cameras mounted on multiple buses operating on urban routes during real-world service. The BFTD contains 8131 images and 56,137 annotated instances across five traffic-participant categories, covering dense pedestrians, mixed-traffic flow, illumination variation, rain, fog, and occlusion-prone scenes. Based on the visual characteristics of bus-mounted cameras, we propose YOLO-M2LA, a local–global attention detection framework in which CBS-SPD preserves fine-grained information during early downsampling and M2LA couples multi-scale local context modeling with efficient global dependency aggregation. Extensive experiments on BFTD and public benchmarks show that the proposed framework improves detection accuracy, particularly for small and visually crowded traffic participants, while maintaining a practical accuracy–efficiency trade-off. Dataset statistics, condition-specific evaluation, ablation analysis, and qualitative visualization further support the effectiveness of BFTD and YOLO-M2LA for vision-based traffic sensing. The dataset and implementation are publicly available online.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments