Hyperparameters fine tuning for MARL comparative study [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
hello everyone. I'm training PPO variants on different multi-agent tasks from the VMAS library (Independent PPO / Graph PPO and such, see HetGPPO by Bettini et al.).
I noticed that for every architecture/scenario couple, the optimal hyperparameters sometimes tend to vary (learning rate, entropy coefficient, KL coefficient, SGD batch size, etc).
do I need - methodologically speaking - to unify the hyperparameters of all models in order to make a fair and correct comparison of architectures later on?
note: sometimes unifying these HP leads to some non converging models.
note 2 : my objective is to test these models' robustness under adversarial attack in test-time (frozen models).
thank you in advance.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.