i use cudnnconvolutionforward function with int8 to speed inference. when w x ,y all int8 type and convdescripter is int32 type . y often overflow。
i find paper to scale it
https://arxiv.org/pdf/1712.05877
where to set value mo and exponents in cudnnconvolutionforward?
or other solutions to avoid int8 y overflow?
