I have a custom environment with a MultiDiscrete action space. The MultiDiscrete action space allows controlling an agent with n-dimensional discrete action spaces.
In my environment, I have 4 dimensions where each dimension has 11 actions. I'm trying to use A2C with a Softmax policy. Below is the implementation of the policy and value networks. The output of the policy gives me [N, 4, 11] tensor where N is the batch size. The softmax is applied to the last dimension of this tensor so basically, I have 4 action distributions. I thought this would work but I'm getting the following error:
Do I need to make changes to the A2C or am I doing something wrong?
File "train_rl.py", line 90, in <module>
train()
File "train_rl.py", line 80, in train
experiments.train_agent_batch(
File "/home/tarik/venvs/tacto/lib/python3.8/site-packages/pfrl/experiments/train_agent_batch.py", line 86, in train_agent_batch
agent.batch_observe(obss, rs, dones, resets)
File "/home/tarik/venvs/tacto/lib/python3.8/site-packages/pfrl/agents/a2c.py", line 224, in batch_observe
self._batch_observe_train(batch_obs, batch_reward, batch_done, batch_reset)
File "/home/tarik/venvs/tacto/lib/python3.8/site-packages/pfrl/agents/a2c.py", line 288, in _batch_observe_train
self.update()
File "/home/tarik/venvs/tacto/lib/python3.8/site-packages/pfrl/agents/a2c.py", line 183, in update
action_log_probs = action_log_probs.reshape(
RuntimeError: shape '[5, 2]' is invalid for input of size 40
policy = torch.nn.Sequential(
torch.nn.Linear(44, 128),
torch.nn.Tanh(),
torch.nn.Linear(128, 128),
torch.nn.Tanh(),
torch.nn.Linear(128, 44),
torch.nn.Unflatten(1, (4, 11)),
SoftmaxCategoricalHead()
)
value = torch.nn.Sequential(
torch.nn.Linear(44, 128),
torch.nn.Tanh(),
torch.nn.Linear(128, 128),
torch.nn.Tanh(),
torch.nn.Linear(128, 1),
)
model = pfrl.nn.Branched(policy, value)
I have a custom environment with a MultiDiscrete action space. The MultiDiscrete action space allows controlling an agent with n-dimensional discrete action spaces.
In my environment, I have 4 dimensions where each dimension has 11 actions. I'm trying to use A2C with a Softmax policy. Below is the implementation of the policy and value networks. The output of the policy gives me [N, 4, 11] tensor where N is the batch size. The softmax is applied to the last dimension of this tensor so basically, I have 4 action distributions. I thought this would work but I'm getting the following error:
Do I need to make changes to the A2C or am I doing something wrong?