Autonomous driving paper index
Why Large Language Models Cannot Be Certified for Safety-Critical Systems
One-line summary
Large language models (LLMs) are moving rapidly into domains governed by functional-safety certification: aviation, road vehicles, medicine, and critical infrastructure.
Engineering notes
Impossibility results establish a nonzero error floor for calibrated probabilistic generators under stated conditions; reported benchmark metrics in legal, medical, and agentic tasks are not directly convertible to certification targets, and no published deployment supplies the system-level hazard and exposure model that a compliance demonstration would require; and every major mitigation family (retrieval augmentation, guardrails, formal verification, uncertainty quantification, and neurosymbolic hybrids) is documented to narrow, but not close, the gap.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
Large language models (LLMs) are moving rapidly into domains governed by functional-safety certification: aviation, road vehicles, medicine, and critical infrastructure. Certification regimes such as IEC 61508, DO-178C, and ISO 26262 combine system-level risk targets with process-, traceability-, configuration-, and evidence-based obligations, in mixes that differ by regime but share a demand for verifiable, bounded behavior. This review synthesizes three literatures that rarely meet: statistical learning theory on hallucination, empirical measurements of LLM error rates in high-stakes domains, and the standards and regulatory documents that define certification. The convergent finding is that the mismatch between the two worlds is structural rather than incidental. Impossibility results establish a nonzero error floor for calibrated probabilistic generators under stated conditions; reported benchmark metrics in legal, medical, and agentic tasks are not directly convertible to certification targets, and no published deployment supplies the system-level hazard and exposure model that a compliance demonstration would require; and every major mitigation family (retrieval augmentation, guardrails, formal verification, uncertainty quantification, and neurosymbolic hybrids) is documented to narrow, but not close, the gap. We formalize error compounding over execution horizons, tabulate the documented limits of each mitigation, and examine why plausibility cannot substitute for assurance. Two coherent exits emerge: statistical acceptance criteria in standards for bounded tasks, and architectures that confine the LLM to an untrusted-proposer role inside a deterministic, independently verifiable execution envelope. Certifying the envelope rather than the model is, on the current evidence, the most defensible path consistent with both the mathematics and the standards.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments