Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

question: [Why it is unable to improve inference speed after quantization] #162

Closed
Njuapp opened this issue May 21, 2020 · 1 comment
Closed
Labels

Comments

@Njuapp
Copy link

@Njuapp Njuapp commented May 21, 2020

I have tried inference using quantized BERT, on MRPC dev set it only reduces time from 1:58 to 1:45, which is minor speed improvement.

I am not sure whether the reason lies in that I do not have INT8 GEMM specialized hardware.

@Njuapp Njuapp added the question label May 21, 2020
@ofirzaf
Copy link
Collaborator

@ofirzaf ofirzaf commented Jun 9, 2020

Hi, similar issue to #90 and #86. In short, quantization implemented in Q8BERT is only a simulation and ops still happen in fp32 precision even though the values are int8. To get performance gain optimized kernels making use of specialized hardware for quantization.

@ofirzaf ofirzaf closed this Jun 9, 2020
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Projects
None yet
Linked pull requests

Successfully merging a pull request may close this issue.

None yet
2 participants
You can’t perform that action at this time.