EDBT 2026 Demo / reviewers in the wild / expert
Wenxun Wang
dblp:314/8279
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-1999-6441ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Transformer Inference Through Optimized Nonlinear Operations With Quantization-Approximation-Based StrategyabstractTransformers have recently shown significant performance across various tasks, such as natural language processing (NLP) and computer vision (CV). However, the performance comes at the cost of large memory and computation overhead. Existing researches primarily focus on accelerating matrix multiplication (MatMul) through techniques like quantization and pruning, notably increasing the proportion of nonlinear operations in inference runtime. Meanwhile, previous approaches designed for nonlinear operations struggle with inefficient implementation as they are incapable of achieving both computation and memory efficiency. Additionally, these methods often require retraining or fine-tuning leading to substantial costs and inconveniences. To overcome these problems, we propose efficient implementation of nonlinear operations with quantization-approximation-based strategy. Through an in-depth analysis of the dataflow and data distribution of nonlinear operations, we design distinct quantization and approximation strategies tailored for different operations. Specifically, log2 quantization and power-of-two factor quantization have been employed in Softmax and LayerNorm, complemented by logarithmic function and low-precision statistic calculation as approximation strategies. Furthermore, the proposed efficient GeLU implementation integrates a nonuniform lookup procedure alongside low-bit-width quantization. Experimental results demonstrate negligible accuracy drops without the need for retraining or fine-tuning. By implementing the hardware design, it achieves$3.14\times - 6.34\times $energy-efficiency and$3.01\times - 10.1\times $area-efficiency improvements compared to state-of-the-art application-specific-integrated-circuit (ASIC) designs. In system-level evaluation, substantial speedup and reductions in energy consumption of 15% to 35% are achieved for end-to-end inference across both GPU and ASIC accelerator platforms. Wenxun Wang, Wenyu Sun, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Block-Wise Dynamic-Precision Neural Network Training Acceleration via Online Quantization Sensitivity AnalyticsabstractData quantization is an effective method to accelerate neural network training and reduce power consumption. However, it is challenging to perform low-bit quantized training: the conventional equal-precision quantization will lead to either high accuracy loss or limited bit-width reduction, while existing mixed-precision methods offer high compression potential but failed to perform accurate and efficient bit-width assignment. In this work, we propose DYNASTY, a block-wise dynamic-precision neural network training framework. DYNASTY provides accurate data sensitivity information through fast online analytics, and maintains stable training convergence with an adaptive bit-width map generator. Network training experiments on CIFAR-100 and ImageNet dataset are carried out, and compared to 8-bit quantization baseline, DYNASTY brings up to 5.1× speedup and 4.7× energy consumption reduction with no accuracy drop and negligible hardware overhead. Ruoyang Liu, Chenhan Wei, Yixiong Yang, Wenxun Wang, Huazhong Yang, Yongpan Liu |
ASP-DAC | 4 |
| 2023 | SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer InferenceabstractTransformers have shown remarkable performance in both natural language processing (NLP) and computer vision (CV) tasks. However, their real-time inference speed and efficiency are limited due to the inefficiency in Softmax and Layer Normalization (LayerNorm). Previous works based on function approximation suffer from inefficient implementation as they place emphasis on computation while disregarding memory overhead concerns. Moreover, such methods rely on retraining to compensate for approximation error which can be costly and inconvenient. In this paper, we present SOLE, a hardware-software co-design for Softmax and LayerNorm which is composed of E2Softmax and AILayerNorm. E2Softmax utilizes log2 quantization of exponent function and log-based division to approximate Softmax while AILayerNorm adopts low-precision statistic calculation. Compared with state-of-the-art designs, we achieve both low-precision calculation and low bit-width storage on Softmax and LayerNorm. Experiments show that SOLE maintains inference accuracy without retraining while offering orders of magnitude speedup and energy savings over GPU, achieving 3.04×, 3.86× energy-efficiency improvements and 2.82×, 3.32× area-efficiency improvements over prior state-of-the-art custom hardware for Softmax and LayerNorm, respectively. Wenxun Wang, Shuchang Zhou 0001, Wenyu Sun, Peiqin Sun, Yongpan Liu |
ICCAD | 1 |
| 2022 | Dynamic CNN Accelerator Supporting Efficient Filter Generator with Kernel Enhancement and Online Channel PruningabstractDeep neural network achieves exciting performance in several tasks with heavy storing and computing costs. Previous works adopt pruning-based methods to slim deep network. For traditional pruning, either the convolution kernel or the network inference is static, which cannot fully compress the model parameter and restrains their performance. In this paper, we propose an online pruning algorithm to support dynamic kernel generation and dynamic network inference at the same time. Two novel techniques including the filter generator and the importance-level based channel pruning are proposed. Moreover, we validate the success of the proposed method by the implementation on Ultra96-v2 FPGA. Compared with state-of-art static or dynamic pruning methods, our method can reduce the top-5 accuracy drop by nearly 50% for ResNet model on ImageNet at similar compressing level. It can also achieve better accuracy while up to 50% fewer weights are reduced to be saved on chip. Wenyu Sun, Wenxun Wang, Yongpan Liu |
ASP-DAC | 3 |
| 2022 | Efficient Neural Networks with Spatial Wise Sparsity Using Unified Importance MapabstractExploiting neural network sparsity is one of the most important directions to accelerate CNN executions. Plenty of techniques are proposed to exploit neural network sparsity, where spatial-wise pruning is quite effective for input image. However, previous spatial-wise pruning methods need nontrivial hardware overhead for dynamic execution, due to layer-by-layer binary sampling and online scheduling. This paper proposes a structured configured, spatial-wise pruning technique. Numerous computation will be saved by skipping unimportant region. By using a unified importance map, the computing graph could be compiled in advance to make it more hardware friendly. Additionally, due to multi-level measurement of importance for each region, our method can have a better performance on various tasks. On image classification task, the method can have around 50% fewer top-1 accuracy drop than previous spatialwise pruning methods at similar sparse level. On super resolution and image deraining task, the method can bring $5 \times$ to $19 \times$ acceleration while causing neglectable effect on reconstruction quality. Hardware implementation is also included. Wenyu Sun, Wenxun Wang, Zhuqing Yuan, Yongpan Liu |
ISCAS | 3 |