Autonomous driving paper index
A decision-making method for autonomous vehicles based on the adaptive control of thought–rational cognitive architecture and deep reinforcement learning
One-line summary
To this end, this paper proposes a hybrid behavior decision-making framework based on the adaptive control of thought—rational (ACT-R) cognitive architecture and deep reinforcement learning.
Engineering notes
Experimental results show that the proposed ACTR–ADPPO method outperforms existing comparison algorithms in terms of convergence speed, reward values, and safety performance, demonstrating the effectiveness and superiority of the hybrid decision-making framework in complex autonomous driving scenarios.
Chinese explanation / 中文解读
中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。
Original abstract
In autonomous driving decision-making systems, conventional single-mode approaches—either knowledge-driven or data-driven—each have inherent strengths and limitations. Knowledge-driven methods provide high interpretability but lack adaptability in complex scenarios, whereas data-driven methods exhibit strong learning capability while relying heavily on large amounts of labeled data and offering limited interpretability. At present, hybrid decision-making frameworks that integrate cognitive modeling and data learning are widely regarded as an effective means to address these limitations. However, the deep integration of knowledge and data still faces challenges in theoretical foundations and practical implementation. To this end, this paper proposes a hybrid behavior decision-making framework based on the adaptive control of thought—rational (ACT-R) cognitive architecture and deep reinforcement learning. First, an optimized decision tree method is used to automatically construct the procedural knowledge module within ACT-R, replacing traditional manual rule construction to improve modeling efficiency and generalization capability. Second, a hybrid decision-making mechanism that combines ACT-R cognitive strategies with a proximal policy optimization (PPO) policy is designed to enhance autonomous learning capabilities while ensuring safety during training. Third, an adaptive clipping strategy is introduced to dynamically adjust the PPO clipping factor according to the source of each strategy, thereby improving training stability and policy performance. Experimental results show that the proposed ACTR–ADPPO method outperforms existing comparison algorithms in terms of convergence speed, reward values, and safety performance, demonstrating the effectiveness and superiority of the hybrid decision-making framework in complex autonomous driving scenarios.
Links and sources
Need this topic turned into a technical roadmap?
Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.
Request B2B research
Comments