题意:OpenAI Gym 的 Lunar Lander 模型未收敛
问题背景:
I am trying to use deep reinforcement learning with keras to train an agent to learn how to play the Lunar Lander OpenAI gym environment. The problem is that my model is not converging. Here is my code:
我正在尝试使用 Keras 深度强化学习训练一个智能体,学习如何在 OpenAI Gym 环境中玩 Lunar Lander。问题是我的模型没有收敛。以下是我的代码:
import numpy as np
import gym
from keras.models import Sequential
from keras.layers import Dense
from keras import optimizers
def get_random_action(epsilon):
return np.random.rand(1) < epsilon
def get_reward_prediction(q, a):
qs_a &#