Effects of hyperparameters in multiple sequence alignment for Align-RUDDER using Clustal
Loading...
Date
Authors
Samwald, Christian
Journal Title
Journal ISSN
Volume Title
Publisher
Jihočeská univerzita
Abstract
Delayed rewards are detrimental to the learning of reinforcement learning agents.One approach to this problem is the usage of return decomposition and rewardredistribution. It was realised in the Align-RUDDER algorithm of Patilet al.[14].Their solution employed the multiple sequence alignment algorithm Clustal W. Iintegrated the sequence alignment Tool Clustal, Clustal W's successor, intoAlign RUDDER to increase efficiency. During the testing process, the usage ofClustal's EPA function and the effects of different sample sizes played a centralrole. The data set that was used came from the MineRL data set [6].
