Join GitHub today
GitHub is home to over 50 million developers working together to host and review code, manage projects, and build software together.
Sign upGitHub is home to over 50 million developers working together to host and review code, manage projects, and build software together.
Sign up
I have tried inference using quantized BERT, on MRPC dev set it only reduces time from 1:58 to 1:45, which is minor speed improvement.
I am not sure whether the reason lies in that I do not have INT8 GEMM specialized hardware.