DDPG Project

最新推荐文章于 2024-07-05 23:19:29 发布

Bourne_Boom

最新推荐文章于 2024-07-05 23:19:29 发布

阅读量343

点赞数 2

分类专栏：强化学习文章标签： DDPG

本文链接：https://blog.csdn.net/linyijiong/article/details/88942078

版权

强化学习专栏收录该内容

21 篇文章 6 订阅

订阅专栏

1. Remember the difference between the DQN and DDPG in the Q function learning is that the Target's next MAX Q value is estimated by the actor, not the critic itself. (In continuous action space, the critic cannot estimate the MAX Q value without optimization. So the best choice is to use actor directly gives the BEST action.)

The code of 1st pic is wrong:

71: the critic_target network is to output the maximum Q value based on the estimation of actor_target network, so there is no need once more max operation (But in DQN we do need that max operation because in DQN the next Max Q value is directly estimated by critic_target itself (Q value function).)

72. the critic (Q function) in DDPG can directly output the relative input action Q value, so there is not need to gather the action index relative Q value.

74. Because optimizer will accumulate the gradient values. so use optimizer.zero_grad() to clear it.(instead of network.zero_grad)

75. Optimizer should call the step() function for backward the error.

. Do not forget to add the determination of final state: 1- dones.

79. In the actor learning part, the input actions of the critic_local is not the sample action, is the action estimated by actor. (Be careful with that). Also, it should calculate the mean of it. Finally, we want to maximize the performance but the optimizer is used to minimize object, so we have to set the negative sign.