Wenhao Dai

dblp:228/3779 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating Sparse Transformer Inference on GPU
abstract
Large language models (LLMs) are popular around the world due to their powerful understanding capabilities. As the core component of LLMs, accelerating Transformer through parallelization has gradually become a hot research topic. Mask layers introduce sparsity into Transformer to reduce calculations. However, previous works rarely focus on the performance optimization of sparse Transformer. In addition, current static operator fusion schemes fail to adapt to diverse application scenarios. To address the above problems, we propose STOF, a framework that incorporates optimizations for Sparse Transformer that enables flexible masking and Operator Fusion on GPU. For multi-head attention (MHA) structure, STOF maps the computation to row-wise or block-wise kernels with unique storage formats according to analytical modeling. For downstream operators, STOF maps the fusion scheme to compilation templates and determines the optimal running configuration through two-stage searching. The experimental results show that compared to the state-of-the-art work, STOF achieves maximum speedups of 1.6× in MHA computation and 1.4× in end-to-end inference.
Wenhao Dai, Haodong Deng, Mengfei Rong, Fangxin Liu, Hailong Yang 0002, Qianwen Cao, Qingxiao Sun
PPoPP1
2026 A fuzzy-autoencoder-based evolutionary multiobjective algorithm for high-dimensional feature selection
Weiping Ding, Qichong Hua, Jiaru Yang, Yuepeng Chen, Wenhao Dai
Neurocomputing5
2026 Deep learning algorithms for license plate recognition: A review
Laixiang Xu, Zihan Shang, Xiangjun Chen, Tiwei Zeng, Wenhao Dai, Junmin Zhao
Neural Networks5
2025 LaRED: Efficient IR Drop Predictor with Layout-Preserving Rebuilder-Encoder-Decoder Architecture
abstract
In the realm of integrated circuit verification, IR drop analysis plays a crucial role. Recent advancements in machine learning (ML) significantly enhance its efficiency, yet many current approaches fail to fully leverage the input structure of feature maps and the transmission mechanism of Power Delivery Network (PDN) layouts. To bridge these gaps, we introduce Layout-Preserving Rebuilder-Encoder-Decoder Architecture Predictor (LaRED), which employs a novel Rebuilder-Encoder-Decoder (RED) architecture and utilizes an innovative downsampling approach and upsampling framework to optimize its perception of instances and the transmission of features. LaRED captures information from various regions with asymmetric topological structure while preserving and transferring layout characteristics through deformable convolution, hybrid downsampling, cascaded upsampling, and attentional feature fusion. The rebuilder rebuilds raw input, whereas the encoder ensures comprehensive feature transmission across all instances. The decoder then facilitates seamless transfer of feature information across layers. This approach enables LaRED to integrate chip features of varying topologies and scales, enhancing its representational power. Compared to the current State-Of-The-Art (SOTA), MAUnet, LaRED achieves accuracy improvements of 34.6% to 42.6% in benchmark tests, establishing it as the new standard in static IR drop analysis for integrated circuit design with ML techniques. The code is available at https://github.com/Todi85/LaRED.
Chengxuan Yu, Yanshuang Teng, Wenhao Dai, Yongjiang Li, Wei W. Xing, Dan Niu, Zhou Jin 0001
DATE3
2025 Convergence-aware operator-wise mixed-precision training
abstract
Abstract With the support of more precision formats in emerging hardware architectures, mixed-precision has become a popular approach to accelerate deep learning (DL) training. Applying low-precision formats such as FP16 and BF16 to neural operators can save GPU memory while improving bandwidth. However, DL frameworks use black and white lists as default mixed-precision selections and cannot flexibly adapt to a variety of neural networks. In addition, existing work on automatic precision adjustment does not consider model convergence, and the decision cost of precision selection is high. To address the above problems, this paper proposes CoMP, a non-intrusive framework for Convergence-aware operator-wise Mixed-precision training. CoMP uses two-stage precision adjustment based on epochs and batches to ensure convergence and performance respectively. After that, CoMP performs subsequent training according to the searched optimal operator-wise mixed-precision plan. The experimental results on A100 GPU show that CoMP achieves a maximum performance speedup of 1.15 $$\times$$ × compared with PyTorch AMP implementation, while also saving up to 29.81% of GPU memory.
Wenhao Dai, Yuesi Bai, Qingxiao Sun
CCF Trans. High Perform. Comput.1
2025 Optimization strategy for batch-stochastic configuration network models and their application in component content prediction
Rongxiu Lu, Xingrong Hu, Cong Pei, Hui Yang 0005, Wenhao Dai, Jianyong Zhu
Eng. Appl. Artif. Intell.5
2020 A Supervised Anonymous Issuance Scheme of Central Bank Digital Currency Based on Blockchain
Wenhao Dai, Xiaozhuo Gu, Yajun Teng
ICA3PP (3)1
2019 Secure Multi-receiver Communications: Models, Proofs, and Implementation
Maomao Fu, Xiaozhuo Gu, Wenhao Dai, Jingqiang Lin 0001
ICA3PP (1)3