博弈论与多智能体强化学习

最新推荐文章于 2023-12-05 15:53:27 发布

Adam婷

最新推荐文章于 2023-12-05 15:53:27 发布

阅读量1.2w

点赞数 3

分类专栏： AI程序员算法机器学习强化学习博弈论论文研读

本文链接：https://blog.csdn.net/weixin_41697507/article/details/93305670

版权

AI程序员同时被 3 个专栏收录

166 篇文章 8 订阅 ¥19.90 ¥99.00

订阅专栏

超级会员免费看

机器学习

161 篇文章 8 订阅 ¥19.90 ¥99.00

订阅专栏

超级会员免费看

强化学习

26 篇文章 1 订阅 ¥19.90 ¥99.00

订阅专栏

超级会员免费看

Ann Nowe´, Peter Vrancx, and Yann-Michae¨l De Hauwere

Abstract. Reinforcement Learning was originally developed for Markov Decision Processes (MDPs). It allows a single agent to learn a policy that maximizes a pos- sibly delayed reward signal in a stochastic stationary environment. It guarantees convergence to the optimal policy, provided that the agent can sufficiently experi- ment and the environment in which it is operating is Markovian. However, when multiple agents apply reinforcement learning in a shared environment, this might be beyond the MDP model. In such systems, the optimal policy of an agent depends not only on the environment, but on the policies of the other agents as well. These situa- tions arise naturally in a variety of domains, such as: robotics, telecommunications, economics, distributed control, auctions, traffic light control, etc. In these domains multi-agent learning is used, either because of the complexity of the domain or be- cau

了解本专栏

超级会员免费看

Adam婷

关注

3
点赞
踩
74

收藏

觉得还不错? 一键收藏
打赏
2
评论
博弈论与多智能体强化学习

Ann Nowe´, Peter Vrancx, and Yann-Michae¨l De HauwereAbstract. Reinforcement Learning was originally developed for Markov Decision Processes (MDPs). It allows a single agent to learn a policy that ma...
复制链接

扫一扫