Autonomous driving paper index
Geospatial Vision-Language Models for Spatial Reasoning and Temporal Change Understanding: A Task-Centered Benchmarking Framework and Evidence Synthesis Discussion
One-line summary
Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work.
Engineering notes
As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
Geospatial vision-language models (VLMs) are increasingly expected to do more than assign scene labels or generate generic captions. In realistic Earth-observation and urban intelligence settings, useful multimodal systems must support fine-grained spatial reasoning, cross-view interpretation, grounded localization, and explicit understanding of change across time. Yet the current literature remains fragmented across remote sensing visual question answering, visual grounding, urban multi-view reasoning, and bi-temporal change captioning. As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. This paper reconstructs the field around a more defensible technical center: geospatial multimodal intelligence as the joint problem of spatial reasoning and temporal change understanding. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work. We first formalize a task-centered problem definition that unifies image-level, region-level, cross-view, and bi-temporal reasoning. We then propose a reference GST-VLM architecture consisting of spatial encoding, temporal difference modeling, multimodal fusion, task-specific decoding, and reliability estimation. Next, we synthesize publicly reported evidence from representative datasets and benchmarks including RSVQA, EarthVQA, VRSBench, GeoChat, LEVIR-CD, LEVIR-CC, SECOND-CC, CHOICE, GEOBench-VLM, CityBench, and UrBench. The synthesis shows that recent models are improving rapidly but remain far from robust geospatial reasoning systems: on GEOBench-VLM, the best public model reported only 41.7% multiple-choice accuracy; on UrBench, even GPT-4o still trails human performance by an average 17.4 percentage points; and while specialized systems such as GeoReasoner, GeoChat, GeoLLaVA, and MModalCC outperform generic baselines on targeted tasks, their gains remain strongly benchmark-dependent. Based on this evidence, we identify the principal bottlenecks as benchmark fragmentation, weak temporal grounding, inadequate calibration, scarce cross-region validation, limited deployment reporting, and insufficient integration of geometry with language-conditioned reasoning. The paper concludes with a concrete research agenda for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments