Wendi Sun

dblp:264/3067 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
0009-0002-3198-2345ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AO-BFP: An Adaptive Mixed-Precision and Outlier-Aware Block Floating-Point Accelerator for Large Language Model Inference
abstract
Large Language Models (LLMs) have achieved remarkable success in Natural Language Processing (NLP) tasks, but their deployment is severely constrained by intensive computation and memory costs. Block Floating-Point (BFP) extends the dynamic range beyond INT with shared exponents, while reducing memory and alignment overhead compared to floating-point formats. However, when bit-widths are further reduced, BFP becomes sensitive to outliers; existing mixed-precision BFP methods largely rely on heuristic settings of mantissa width and block size, and the induced bit-level sparsity has yet to be systematically leveraged in hardware. In this paper, we propose AO-BFP, an adaptive BFP framework for LLM inference. At the algorithm level, we propose an adaptive outlier exponent mapping mechanism combined with mixed-precision exploration driven by layer-wise sensitivity analysis. At the hardware level, we design a reconfigurable bit-serial accelerator with a unified datapath that efficiently leverages BFP-induced bit sparsity. Compared with prior LLM accelerators such as ANT, OliVe, and BitMoD, AO-BFP achieves superior performance while preserving model accuracy, delivering speedups of 1.61×, 1.39×, and 1.11×, respectively.
Zetao Guo, Wendi Sun, Qiyan Fang, Song Chen 0001, Yi Kang
DATE3
2026 Model-Hardware Co-Design of Depthwise Separable Convolution With Interlayer Pipelining for Stereo Matching
Zetao Guo, Wendi Sun, Jiaheng Ruan, Yukang Han, Song Chen 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2025 IOPS: A Unified SpMM Accelerator Based on Inner-Outer-Hybrid Product
abstract
Sparse matrix multiplication (SpMM) is widely applied to numerous domains, such as graph processing and machine learning. However, inner product (IP) induces redundant zero-element computing for mismatched nonzero operands, while outer product (OP) lacks input reuse across Process Elements (PEs). Besides, current accelerators only focus on sparse-sparse matrix multiplication (SSMM) or sparse-dense matrix multiplication (SDMM), rarely performing efficiently for both. To compensate for the shortcomings of IP and OP, we propose an inner-outer-hybrid product (IOHP) method, which reuses the input matrix among PEs with IP and removes zero-element calculations with OP in each PE. Based on IOHP, we co-design a accelerator with a unified computing flow, called IOPS, to efficiently process both SSMM and SDMM. It divides the SpMM into three stages: encoding, partial sum (psum) calculation, and address mapping, where the input matrices can be reused among PEs after encoding (IP) and the zero element can be skipped in the latter two stages (OP). Furthermore, an adaptive partition strategy is proposed to tile the input matrices based on their sparsity ratios, effectively utilizing the on-chip storage and reducing DRAM access. Compared with SpArch, we achieve 1.2×~4.3× performance and 1.3×~4.8× energy efficiency, with 1.4×~2.1× DRAM access saving.
Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Computers2
2025 Comma: A Communication-Minimized Model-Architecture Framework for Efficient Convolution Acceleration
Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Very Large Scale Integr. Syst.1
2024 Parallel Multi-Objective Bayesian Optimization Framework for CGRA Microarchitecture
abstract
Recently, due to the flexibility and reconfigurability of Coarse-Grained Reconfigurable Architecture (CGRA), CGRA microarchitecture has become an inevitable trend to accelerate the convolution calculation in diverse deep neural networks. However, since the vast microarchitecture design space and the complicated VLSI verification flow, it is a huge challenge to explore a perfect microarchitecture to compromise between multiple performance metrics. In this paper, we formulate the CGRA microarchitecture design as a design space exploration problem, and propose a parallel multi-objective Bayesian optimization framework (PAMBOF) to automatically explore the CGRA microarchitecture design space. Meanwhile, high-precision performance and area models are built to enable fast design space exploration. To approximate the black-box objective function in the design space, the PAMBOF framework first builds multiple Gaussian processes (GP) with deep regularization kernel learning functions (DRKL-GP). Then a parallel Bayesian optimization algorithm is developed to sample a batch of candidate design points, which are simulated in parallel by the performance and area models. Experimental results demonstrate that compared to the prior arts, the proposed PAMBOF framework can search for a CGRA microarchitecture design with the better area and performance in a shorter runtime.
Wendi Sun, Xiaobing Ni, Kaixuan He, Qi Xu 0004, Song Chen 0001, Yi Kang
DATE2
2024 Communication Minimized Model-Architecture Co-design for Efficient Convolution Acceleration
abstract
CNN is indispensable for today’s Artificial Intelligence (AI) applications, but brings dominantly large overhead of data communication. Current works mainly focus on prior off-chip or intuitive/heuristic on-chip access optimization, but with the development of Near Memory Processing (NMP), DRAM access cost has greatly dropped and on&off-chip access optimization needs rethinking as a whole. Thus, this paper proposes a holistic on&off-chip communication-minimized model-architecture acceleration scheme for CNN. First, we derive the layer-wise off-chip communication Lower Bound (LB) based on different data reuse strategies. Second, on-chip LB is derived and overall on&off-chip communication analysis model is presented to provide a solid guidance for on-chip storage allocation, dataflow and architecture design. Finally, we design Window-Primitive (WP) dataflow and a Systolic-Cross-Line (SCL) CNN accelerator based on proposed theoretical model. SCL achieves 3.8 × pJ/MAC energy reduction at 1.4 × less on-chip storage area compared with Eyeriss and 1.3~1.8 × reduction at 3~4 × less area compared with CLB. For NMP, we reduce around 2 × access energy compared with previous systolic NMP architecture.
Wendi Sun, Yi Kang, Song Chen 0001
ACM Great Lakes Symposium on VLSI1
2024 Spectrum utilization improvement for multi-channel EH-CRN with spectrum sensing
abstract
Abstract Due to the ever‐growing applications and services of the Internet of Things (IoT), designing energy‐efficient and spectral‐efficient transmission schemes to support IoT devices for the 6G space–air–ground integrated networks becomes much more challenging. Fortunately, energy harvesting (EH) and cognitive radio (CR) technologies have been proposed to alleviate these challenges. Inspired by this fact, this paper studies the issue of spectrum reuse in terms of spectrum utilization efficiency (SUE) in the energy harvesting cognitive radio network (EH‐CRN), where multiple primary transceiver pairs, one multi‐antenna secondary transmitter (ST), and one secondary base station (SBS) coexist. To characterize the impact of small‐scale fading and improve the SUE of the EH‐CRN with perfect spectrum sensing (SS), an adaptive scheme concerning SS, channel selection, EH, and data transmission (SCED) scheme are proposed, where the ST selects the channels for SS based on the residual energy, and adjusts the duration of EH and data transmission with respect to the sensing results. Then the Markov decision process problem of SUE is formulated, which is challenging due to the infinite system space and action space. To tackle the Markov decision process problem, the system space and action space are discreted, and divide the ST into the energy‐limited case and energy‐sufficient case according to specific energy condition. Moreover, theoretical results are extended to the EH‐CRN with imperfect SS. Numerical results show that the SUE under the SCED scheme in perfect SS and imperfect SS scenarios is better than that under other schemes.
Kechen Zheng, Jiahong Wang, Anping Chen, Wendi Sun, Xiaoying Liu 0001, Jia Liu 0009
IET Commun.4
2024 Bit-Balance: Model-Hardware Codesign for Accelerating NNs by Exploiting Bit-Level Sparsity
abstract
Bit-serial architectures can handle Neural Networks (NNs) with different weight precision, achieving higher resource efficiency compared with bit-parallel architectures. Besides, the weights contain abundant zero bits owing to the fault tolerance of NNs, indicating that bit sparsity of NNs can be further exploited for performance improvement. However, the irregular proportion of zero bits in each weight causes imbalanced workloads in the Processing Element (PE) array, which degrades performance or induces overhead for sparse processing. Thus, this article proposed a channel-wise bit-sparsity quantization method that keeps the non-zero bit number of each weight in each channel from exceeding a certain threshold and clusters the channels with the same threshold to balance the workloads in PE array with little accuracy loss. Then, we co-designed a sparse bit-serial architecture, called Bit-balance, to improve overall performance, supporting weight-bit sparsity and adaptive bitwidth computation. The whole design was implemented with 65 nm technology at 1 GHz and performs at 447-, 37-, 59-, 240-, and 19-frame/s for AlexNet, VGG-16, ResNet-50, GoogleNet, and Yolo-v3 respectively. Compared with sparse bit-serial accelerator, Bitlet, Bit-balance achieves 1.6$\boldsymbol{\times}$2.1$\boldsymbol{\times}$energy efficiency (frame/J) and 2.3$\boldsymbol{\times}$3.6$\boldsymbol{\times}$resource efficiency (frame/mm${}^{\mathbf{2}}$).
Zhiwei Zou, Deng Liu, Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Computers4
2023 Sense: Model-Hardware Codesign for Accelerating Sparse CNNs on Systolic Arrays
abstract
Sparsity is an intrinsic property of convolutional neural networks (CNNs), worth exploiting for CNN accelerators. However, the extra processing involved comes with hardware overhead, resulting in only marginal profits for most architectures. Meanwhile, systolic arrays have become increasingly competitive on CNN acceleration for its high spatiotemporal locality and low hardware overhead. However, the irregularity of sparsity induces imbalanced workloads under the rigid systolic dataflow, causing performance degradation. Thus, this article proposed a systolic-array-based architecture, called Sense, for sparse CNN acceleration by model-hardware codesign, enabling large performance gains. To balance input feature map (IFM) and weight loads across the processing element (PE) array, we applied channel clustering to gather IFMs with approximate sparsity for array computation and codesigned a load-balancing weight pruning method to keep the sparsity ratio of each kernel at a certain value with little accuracy loss, improving PE utilization and overall performance. In addition, adaptive dataflow configuration was applied to determine the computing strategy based on the storage ratio of IFMs and weights, lowering$1.17\times $–$1.8\times $dynamic random access memory (DRAM) access compared with Swallow and further reducing system energy consumption. The whole design was implemented on ZynqZCU102 with 200 MHz and performs at 471, 34, 53, and 191 image/s for AlexNet, VGG-16, ResNet-50, and GoogleNet, respectively. Compared with sparse systolic-array-based accelerators, Swallow, fusion-enabled systolic architecture (FESA), and SPOTS, Sense achieves$0.97\times $–$2.18\times $,$1.3\times $–$1.67\times $, and$0.94\times $–$1.82\times $energy efficiency (image/J) on these CNNs, respectively.
Deng Liu, Zhiwei Zou, Wendi Sun, Song Chen 0001, Yi Kang
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Throughput maximisation for multi-channel energy harvesting cognitive radio networks with hybrid overlay/underlay transmission
abstract
Abstract This paper focuses on the issue of joint time and power allocation in multi‐channel energy harvesting CR networks (EH‐CRNs), where the multi‐antenna secondary transmitter (ST) opportunistically accesses the licensed subchannels by a hybrid overlay/underlay transmission approach. To improve spectrum efficiency and energy efficiency of the EH‐CRNs, the ST scavenges energy from the radio‐frequency signal radiated by the primary transmitter, and exploits the harvested energy for data transmission through subchannels of different states in overlay/underlay mode simultaneously. Moreover, under the interference power constraint, energy constraint, and maximum power constraint, the secondary throughput is improved by optimising the allocation of subchannels, the time scheduling between energy harvesting and data transmission, and the power allocation of the ST among different subchannels. A subchannel allocation scheme with low time complexity is proposed, and the secondary throughput optimisation problem is formulated with respect to the time scheduling and power allocation of the ST. Then it is proved the problem is convex, and the problem is solved by a proposed joint time and power allocation algorithm. Numerical results show that the proposed scheme has an advantage of secondary throughput over the other schemes. Finally, the impacts of key relevant factors on the secondary throughput are explored.
Kechen Zheng, Wendi Sun, Xiaoying Liu 0001, Yang Xu 0012, Jia Liu 0009
IET Commun.2