Join GitHub today
GitHub is home to over 50 million developers working together to host and review code, manage projects, and build software together.
Sign upGitHub is where the world builds software
Millions of developers and companies build, ship, and maintain their software on GitHub — the largest and most advanced development platform in the world.
Hi.
I'm asking A common practice introduced in the paper is that we can set the learning rate to 2e-5, batch size to 32, 64. Does this mean there is no need to do hyperparameter tuning by ourselves?
Through experiments, I see that on a 6k dataset for binary classification, all hyperparameters tend to make the model overfit on the dataset. But produce slightly different test performance. In this case, is there a need to do a random search?
The paper says during pretraining, the learning rate was set to 1e-4 and weight decay was set to 0.01. Could you give some explanation of why setting a large weight decay?
Thanks a lot!
Please feel free to leave a comment!