EDBT 2026 Demo / reviewers in the wild / expert
Qingwen Wei
dblp:361/7558
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GRAG-ProSafe QAS: Graph retrieval-augmented generation for process production safety intelligent question and answer system
Jianrong Zhang, Qingwen Wei, Huayu Zhong |
Expert Syst. Appl. | 3 |
| 2025 | SUArch: Accelerating Layer-wise N: M Sparse Pattern with a Unified Architecture for Deep-learning Edge DeviceabstractDeep neural networks are of the essence for user applications on edge devices. However, the computation and memory-intensive nature of deep neural networks conflicts with the resource-constrained devices. Moreover, the heterogeneity across different models imposes new challenges on deployment on edge devices. To boost the capabilities of edge devices, we propose SUArch, which innovates on three fronts: 1) a layer-wise N:M sparsity aware training approach to strike a balance between accuracy and training cost; 2) a sparsity alignment unit based on the butterfly network to maximize hardware utilization and eliminate extra overhead; 3) a mode-heterogenous processing element array to effectively accomplish the unified support for Convolution Neural Network and Transformer. The experimental results demonstrate that when running convolution-based and attention-based models under an industrial 28-nm process, the proposed SUArch realizes an energy efficiency of 52.1 TOPS/W. Compared to state-of-the-art architecture, SUArch achieves an energy efficiency improvement of 2.07× while accuracy loss is within 0.7%. Xilong Kang, Qingwen Wei, Ningyuan Li 0004, Xingyu Xu 0008, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 2 |
| 2025 | TWDP: A Vision Transformer Accelerator with Token-Weight Dual-Pruning Strategy for Edge Device DeploymentabstractVision Transformers (ViTs) have attracted significant attention due to their superior accuracy compared to convolutional neural networks (CNNs) in various computer vision tasks. However, their substantial computational load and significant memory footprint lead to excessive delay and considerable data storage overhead, posing challenges for resource-limited edge device deployment. To address these issues, we present TWDP, a vision transformer accelerator employing a Token-Weight Dual-Pruning strategy to enhance the efficiency of the inference process. Firstly, we propose a parameter-free self-adaptive token pruning method to skip redundant computations in an image-dependent manner. Secondly, we apply a Hessian-aware layer-wise N:M weight pruning approach to minimize storage overhead, memory access, and computational power consumption. Additionally, to manage the complex computing patterns in ViTs, an overlapping dataflow is utilized to further reduce temporal storage and inference latency. Implemented and evaluated under an industrial 28nm technology, the proposed TWDP framework reduces 66.1% weight storage requirements and achieves an energy efficiency of 2070.9 FPS/W. Compared to state-of-the-art architectures, TWDP obtains a 1.6× energy efficiency improvement with negligible accuracy loss, demonstrating the superiority of TWDP in edge device deployment scenarios. Guang Yang 0036, Xinming Yan, Hui Kou, Zihan Zou, Qingwen Wei, Hao Cai 0001, Bo Liu 0019 |
ASP-DAC | 5 |
| 2025 | A Layer-wise N: M Sparsity Aware Transformer Accelerator leveraging Temporal Locality with Butterfly NetworkabstractIn recent years, Transformer-based neural networks have driven significant advancements across various fields. However, the attention mechanism results in high memory access and computational demands, making deployment on edge devices challenging. To address these issues, we propose a layer-wise N:M sparsity-aware Transformer accelerator with two key innovations: 1) a layer-wise N:M sparsity-based parallel dimension compression strategy that effectively leverages the temporal locality of feature values, and 2) a butterfly network-based scattering unit that ensures correct accumulation of partial products after multiplication. Experimental results demonstrate that the proposed accelerator achieves up to a 3.86× improvement in runtime cycles. When running Deit-s on Imagenet-1k dataset, the accelerator delivers an energy efficiency of up to 34.4 TOPS/W at 68.9% sparsity under an industrial 28-nm technology, with only a 0.39% accuracy loss, representing a 1.46~1.48× improvement over prior work. Qinfan Wang, Xilong Kang, Qingwen Wei, Cai Hao, Bo Liu 0019 |
ISCAS | 4 |
| 2025 | Performance Analysis of Direct Acyclic Graph-Based Ledgers in Low-to-High Load RegimeabstractDirect acyclic graph (DAG)-based ledgers and distributed consensus algorithms have been proposed for use in the Internet of Things (IoT). The DAG-based ledgers have many advantages over single-chain blockchains, such as low resource consumption, low transaction fee, high transaction throughput, and short confirmation delay. However, the scalability of the DAG consensus has not been comprehensively verified on a large scale. This paper explores the scalability of DAG consensus within the low-to-high load regime (L2HR) using the tangle model, where L2HR characterizes the transition from a phase of low network load to another phase of high network load. In particular, we determine the average number of tips in the tangle in L2HR when adopting the uniform random tip selection (URTS) and rigorously prove that using the tangle model, the average number of tips at the end of L2HR converges to a constant. We also analyze the probability that a transaction in L2HR becomes an abandoned tip, the approximate average time required for the network load to transition from low load regime (LR) to high load regime (HR), and the average time required for a tip being approved for the first time in L2HR. All analytics are verified by numerical simulations. Qingwen Wei, Shuping Dang, Zhihui Ge, Xiangcheng Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | FDCA: Fine-grained Digital-CIM based CNN Accelerator with Hybrid Quantization and Weight-Stationary DataflowabstractDigital-Compute-in-memory (DCIM) has demonstrated significant energy and area efficiency in convolutional neural network (CNN) accelerators, particularly for high precision applications. However, to mitigate parasitic effects on word and bit lines, most DCIMs employ fine-grained multiply-accumulate operations, which introduces new challenges and opportunities but has not been widely explored. This paper proposes FDCA: a fine-grained digital-CIM based CNN accelerator with hybrid quantization and weight-stationary dataflow, in which the key contributions are :1) a hybrid quantization approach for CNNs leveraging hessian trace and approximation is utilized. This method incorporates the ratio of computation time and storage time into quantization, achieving high efficiency while maintaining accuracy; 2) a Cartesian Genetic Programming based approximate shift and accumulate with error compensation is proposed, where an approximate adder tree is generated to compensate for errors introduced by DCIM; 3) an optimized weight-stationary dataflow is used to improve the utilization of CIM and eliminate dataflow stalls. The experimental results demonstrate that under 28-nm process, when running VGG16 and ResNet50 on CIFAR100, the proposed FDCA achieves 17.1TOPS/W and 18.79TOPS/W with only a slight decrease in accuracy by 0.71% and 0.98%, respectively. Compared to previous works, this work achieves 1.76× and 1.28× better in energy efficiency with less accuracy loss. Bo Liu 0019, Qingwen Wei, Yang Zhang 0132, Xingyu Xu 0008, Zihan Zou, Xinxiang Huang, Xin Si, Hao Cai 0001 |
DAC | 2 |
| 2024 | Timing Error Tolerant CNN Accelerator With Layerwise Approximate MultiplicationabstractExploiting the error tolerance in computation, approximate circuits become an emerging computing paradigm to increase the energy efficiency in digital systems, which is crucial in high-performance and low-power systems for the edge Internet-of-Things (EIoT) devices. Inspired by the state-of-the-art high-efficiency NN accelerators, three techniques are proposed for effectively integrating the approximate computing unit into CNN accelerator to achieve a dynamic energy-accuracy trade-off: (1) An approximate multiplier that can be configured to three precision modes is proposed. A weight pre-encoding method is used to save hardware overhead. (2) For hybrid-accuracy layer-wise mapping, the hessian-aware layer-wise accuracy scaling is proposed, which concerns inference accuracy and hardware overhead simultaneously. A progressive re-training approach is proposed to enable an aggressive approximation configuration and higher power reduction. (3) A tensor multiplication unit (TMU) with timing error detection and correction (TEDC) approach is proposed, enabling an aggressive voltage scaling and a 41.5% power reduction is obtained. An energy-efficient CNN accelerator is proposed and shows how deep learning can be brought to EIoT devices by running each layer at its appropriate computational accuracy. Implemented under 28-nm CMOS technology, the CNN accelerator achieves the energy efficiency of 14.4 TOPS/W. The proposed accelerator and method are conducted on the applications of keyword spotting of GSCD, CIFAR10 and CIFAR100, 44.5%~46.7% multiplication energy is saved while reducing the accuracy by less than 0.6%. Bo Liu 0019, Na Xie, Qingwen Wei, Guang Yang 0036, Chonghang Xie, Weiqiang Liu 0001, Hao Cai 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Work-in-Process: Error-Compensation-Based Energy-Efficient MAC Unit for CNNsabstractApproximate circuits sacrifice accuracy in exchange for energy efficiency and have been widely used in hardware deployment of neural networks (NNs). Since convolution accounts for most of the power consumption in NNs, it is necessary to design an approximate multiplication and accumulation (MAC) unit which improve the energy efficiency of hardware with ignorable accuarcy loss. In this work, an error-compensation-based energy-efficient MAC unit is proposed in which approximate multipliers are designed by Boolean matrix factorization and approximate adders are generated by Cartesian genetic programming. The proposed MAC unit is conducted on CIFAR10 using ResNet-18, where PDP is reduced by 58.8% with an accuracy loss of 0.81%. Xingyu Xu 0008, Qingwen Wei, Yang Zhang 0132, Hao Cai 0001, Bo Liu 0019 |
CASES | 2 |