Autonomous driving paper index

Understanding the Markov decision process for reinforcement learning in robotics perception

2026-07-31 · Scientific Reports

autonomous drivingpath planningreinforcement learningperceptionplanningcontrol

One-line summary

This paper presents a formal framework and comprehensive introduction of the Markov Decision Process (MDP), together with the description of value functions and policies.

Engineering notes

The application of robots has increased significantly in recent years in various sectors, including household use, surveillance, medical services, manufacturing, and logistics automation.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为端到端自动驾驶、BEV感知、3D目标检测、轨迹预测、路径规划、LiDAR感知等高价值论文补充中文说明。

Original abstract

The application of robots has increased significantly in recent years in various sectors, including household use, surveillance, medical services, manufacturing, and logistics automation. To guarantee the efficient and secure functioning of such robots, it is essential to develop reliable control mechanisms that can adjust to evolving settings. Reinforcement learning (RL), placed between supervised and unsupervised learning, addresses learning in successive decision-making issues characterized by inadequate feedback. This paper presents a formal framework and comprehensive introduction of the Markov Decision Process (MDP), together with the description of value functions and policies. The principles and perceptions underlying MDPs are presented in this article, along with RL algorithms for calculating optimal behaviors using dynamic programming for a Q-learning-based approach. RL environments for robotics path planning are generally designed as MDPs, with the goal of learning a control strategy to maximize the total reward. Based on multiple interpretations of optimality with regard to the objective of learning sequencing decisions, the primary focus of the present paper is to provide the introduction of fundamental information of MDP-based algorithms to acquire optimal behaviors. Lastly, theoretical aspects are validated in this study with a simulation experiment using the Q-learning approach in a custom robot environment that has various obstacles at different locations. The initial and target positions are defined for the robot in the environment in the grid map. The simulation shows a rate of accuracy of 88.80% for the robot to reach the target point. Furthermore, a comparative performance analysis has been performed between Q-learning and DQN algorithms.

5.0Engineering value
7.0Research novelty
5.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

Full Self Driving can prepare a custom autonomous driving literature review, code map, dataset map, and B2B technology assessment.

Request B2B research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment