VLDB 2026 Research / reviewers in the wild / expert
Shiwei Liu 0002
dblp:234/8697-2
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-0564-4900ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLARE: Finetuning ReLU With FIRE for Efficient Long-Context InferenceabstractDeploying large language models (LLMs) on resource-constrained edge devices, such as mobile phones or IoT devices, is highly desirable for enabling secure, personalized on-device AI. However, there are significant challenges due to these models’ high computational and memory demands. A key bottleneck lies in the Transformer’s attention block, especially when handling long contexts. Techniques like model architectures with Rectified Linear Unit (ReLU) activations for Softmax and FIRE positional encoding (a resource-efficient, automatic context-length-scaling alternative to Rotary Positional Embedding (RoPE)) have each independently shown promise in reducing the computational complexity of the attention block, but the proper alchemy for combining their benefits remains underexplored. In this paper, we show a method for combining FIRE and ReLU that maintains low-validation loss at long contexts. We also introduce FLARE, a new algorithm that further improves efficiency by removing operations from the learned relative position encoding in FIRE. Our approach leads to faster inference on long sequences, robust generalization to varying context lengths, and lower validation loss compared to baseline models. FLARE achieves a significant reduction in power and area consumption. On custom hardware, it achieves a 6× higher operating frequency than Softmax, while occupying 57× less silicon area (measured under different throughput settings) and consuming 600× less energy. Our results indicate that FLARE represents a significant step towards deploying powerful LLMs efficiently on resource-limited devices. Code and hardware designs are publicly available at: https://github.com/ReaLLMASIC/nanoGPT. Michael Moffatt, Junyi Luo, Xinting Jiang, Guanchen Tao, Shiwei Liu 0002, Kauna Lei, Gregory Kielian, Mehdi Saligane |
DATE | 7 |
| 2026 | A 1024-Ch 583-nW/Ch Spike-Sorting SoC With Sparsity-Aware Spike Detection Scratchpad and Ultra-Low-Leakage Dual-Voltage 5T-SRAM for 16K-Template ClusteringabstractThis paper presents an energy-efficient spike-sorting system-on-chip (SoC) designed for closed-loop brain-computer interfaces of massive probing channels. The design first incorporates a sparsity/similarity-aware spike detection scratchpad, leveraging a bit-wise differential encoder and zero-friendly read-out circuits, reducing the dynamic power consumption of spike detection by 77.7%. To mitigate static power dissipation, it also introduces an ultra-low-leakage dual-voltage 5T-SRAM array with level-shifter embedded sense amplifiers, achieving an 82.2% leakage power reduction of neural signal buffering by applying half$V_{DD}$on SRAM cells. Additionally, a memory hierarchy architecture combining on-chip SRAM and off-chip FeRAM, along with a firing-rate-based Osort for cluster template management, minimizes off-chip memory access to only 9.7% with a latency of$11.7\mu $s for 1024-channel spike sorting. A silicon prototype is fabricated in 28-nm CMOS technology, which achieves a power consumption of 583nW/channel and an area consumption of 0.0012mm2/channel. The chip supports real-time spike sorting with up to 16K templates,$21.3\times $greater than the state-of-the-art spike-sorting processor. Hao Jiang 0024, Zexing Chen, Jiajun Lu, Siqi He, Liangjian Lyu, Jiamin Xu, Shiwei Liu 0002, Yingping Chen, Chixiao Chen, Qi Liu 0010, Ming Liu 0022 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | DRAFT: Decoupling Backpropagation from Pre-trained Backbone for Efficient Transformer Fine-Tuning on EdgeabstractTransformers have demonstrated outstanding performance across diverse applications recently, necessitating finetuning to optimize their performance for downstream tasks. However, fine-tuning remains challenging due to the substantial computational costs and storage overhead of backpropagation (BP). The existing fine-tuning techniques require the BP through the massive pre-trained backbone weights for computing the input gradient, resulting in significant computing overhead and memory footprint for resource-constrained edge devices. To address the challenge, this work proposes an algorithm-hardware co-design framework, DRAFT, for efficient Transformer finetuning by decoupling the BP from the backbone weights, thereby efficiently reducing the BP overhead. The framework employs Feedback Decoupling Approximation (FDA), an efficient finetuning algorithm that decouples BP into two low-complexity pathways: trainable adapter pathway and sparse ternary Bypass Network (BPN) pathway. The two pathways work collaboratively to approximate the conventional BP process. Further, a DRAFT accelerator is proposed, featuring a reconfigurable design with lightweight sparse gather networks and dynamic workflows to fully harness the sparsity and data parallelism inherent to the FDA. Experimental results demonstrate that DRAFT achieves a speedup of $4.9 \times$ and an energy efficiency improvement of $4.2 \times$ on average compared to baseline fine-tuning methods across multiple fine-tuning tasks with negligible accuracy loss. Zhirui Huang, Shiwei Liu 0002, Haozhe Zhu, Qi Liu 0010, Chixiao Chen |
DAC | 2 |
| 2025 | McPAL: Scaling Unstructured Sparse Inference with Multi-Chiplet HBM-PIM Architecture for LLMsabstractLarge language models (LLMs) have gained significant attention recently. However, executing LLM is memory-bound due to the extensive memory accesses. Process-in-memory (PIM) emerges as an energy-efficient solution for LLMs, delivering high memory bandwidth and compute parallelism. Nevertheless, the trend towards larger LLMs introduces escalating memory footprint challenges for monolithic PIM chips. This paper proposes McPAL, which tackles this challenge by emphasizing unstructured sparse compute within PIM and hierarchical multi-chiplet scaling. McPAL decomposes arbitrary sparse weight matrix into multiple irregular sparse vectors. The non-skipped computations in each vector are then routed via an in-memory butterfly network to the standard PIM array, enhancing the PIM utilization. In addition, we scale McPAL vertically by strategically organizing the 3D-HBM hierarchy to minimize the internal long-distance data travel. Meanwhile, a 2.5D IO chiplet scales McPAL horizontally, reducing die-to-die (D2D) data transfer and ensuring sparse workload balance. We conducted extensive experiments from Llama-7B to Llama-70B. The results show that McPAL achieves $1.57 \times$ to $3.12 \times$ speedup and $10.43 \times$ to $35.66 \times$ energy efficiency over Nvidia A100 GPU. Compared to SOTAs, McPAL also achieves $1.08 \times$ to $2.15 \times$ speedup and $1.65 \times$ to $5.14 \times$ energy efficiency. Shiwei Liu 0002, Zhirui Huang, Jiangnan Yu, Qi Liu 0010, Chixiao Chen |
DAC | 1 |
| 2024 | ConSmax: Hardware-Friendly Alternative Softmax with Learnable ParametersabstractThe self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvement, achieving real-time LLM inference on silicon remains challenging due to the extensive use of Softmax in self-attention. In addition to the non-linearity, the low arithmetic intensity significantly limits processing parallelism, especially when working with longer contexts. To address this challenge, we propose Constant Softmax (ConSmax), a software-hardware co-design that serves as an efficient alternative to Softmax. ConSmax utilizes differentiable normalization parameters to eliminate the need for maximum searching and denominator summation in Softmax. This approach enables extensive parallelization while still executing the essential functions of Softmax. Moreover, a scalable ConSmax hardware design with a bitwidth-split look-up table (LUT) can achieve lossless non-linear operations and support mixed-precision computing. Experimental results show that ConSmax achieves a minuscule power consumption of 0.2mW and an area of 0.0008mm2 at 1250MHz working frequency in 16nm FinFET technology. For open-source contribution, we further implement our design with the OpenROAD toolchain under SkyWater's 130nm CMOS technology. The corresponding power is 2.69mW and the area is 0.007mm2. ConSmax achieves 3.35× power savings and 2.75× area savings in 16nm technology, and 3.15× power savings and 4.14× area savings with the open-source EDA toolchain. In the meantime, it also maintains comparable accuracy on the GPT-2 model and the WikiText103 dataset. The project is available at https://github.com/ReaLLMASIC/ConSmax. Shiwei Liu 0002, Guanchen Tao, Yifei Zou, Derek Chow, Zichen Fan, Kauna Lei, Bangfei Pan, Dennis Sylvester, Gregory Kielian, Mehdi Saligane |
ICCAD | 1 |
| 2024 | HARDSEA: Hybrid Analog-ReRAM Clustering and Digital-SRAM In-Memory Computing Accelerator for Dynamic Sparse Self-Attention in TransformerabstractSelf-attention-based transformers have outperformed recurrent and convolutional neural networks (RNN/ CNNs) in many applications. Despite the effectiveness, calculating self-attention is prohibitively costly due to quadratic computation and memory requirements. To solve this challenge, this article proposes a hybrid analog-ReRAM and digital-SRAM in-memory computing accelerator (HARDSEA), a computing-in-memory (CIM) accelerator supporting self-attention in transformer applications. To trade off between energy efficiency and algorithm accuracy, HARDSEA features an algorithm-architecture-circuit codesign. A product-quantization-based scheme dynamically facilitates self-attention sparsity by predicting lightweight token relevance. A hybrid in-memory computing architecture employs both high-efficiency analog ReRAM-CIM and high-precision digital SRAM-CIM to implement the proposed new scheme. The ReRAM-CIM, whose precision is sensitive to circuit nonidealities, takes charge of token relevance prediction where only computing monotonicity is demanded. The SRAM-CIM, utilized for exact sparse attention computing, is reorganized as an on-memory-boundary computing scheme, thus adapting to irregular sparsity patterns. In addition, we propose a time-domain winner-take-all (WTA) circuit to replace the expensive ADCs in ReRAM-CIM macros. Experimental results show that HARDSEA prunes BERT and GPT-2 models to 12%–33% sparsity without accuracy loss, achieving$13.5\times $–$28.5\times $speedup and$291.6\times $–$1894.3\times $energy efficiency over GPU. Compared to state-of-the-art transformer accelerators, HARDSEA has$1.2\times $–$14.9\times $better energy efficiency at the same level of throughput. Shiwei Liu 0002, Chen Mu, Hao Jiang 0024, Yunzhengmao Wang, Jinshan Zhang 0006, Keji Zhou, Qi Liu 0010, Chixiao Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D IntegrationabstractChiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost. Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen |
ISCAS | 4 |
| 2021 | Systolic-Array Deep-Learning Acceleration Exploring Pattern-Indexed Coordinate-Assisted Sparsity for Real-Time On-Device Speech ProcessingabstractThis paper presents a hardware-software co-design for efficient sparse deep neural networks (DNNs) implementation in a regular systolic array for real-time on-device speech processing. The weight pruning format, exploring pattern-based coordinate-assisted (PICA) sparsity, expands the pattern-based pruning into both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). It reduces the index storage overhead as well as avoids accuracy degradation. The proposed systolic accelerator leverages the intrinsic data reuse and locality to accommodate the PICA-based sparsity without using complex data distribution networks. It also supports DNNs with different topologies. By reducing the model size by 16x, PICA sparsification reduces 6.02x index storage overhead while still achieving 20.7% WER in TIMIT dataset. For the pruned WaveNet and LSTM, the accelerator achieves 0.62 and 2.69 TOPS/W energy efficiency, 1.7x to 10x higher than the state-of-the-art. Shiwei Liu 0002, Qiaosha Zou, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 1 |