EDBT 2026 Demo / reviewers in the wild / expert
Shengbai Luo
dblp:370/3070
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0007-5551-2897ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HIVE+: An Enhanced High-Priority Victim Cache to Accelerate GPU Memory Accesses
Yuhan Tang, Sheng Ma, Hanqing Li, Shengbai Luo, Jixuan Tang, Siqing Fu, Lizhou Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | HIVE: A High-Priority Victim Cache for Accelerating GPU Memory AccessesabstractThe victim cache was originally designed as a secondary cache to handle misses in the L1 data (L1D) cache in CPUs. However, this design is often sub-optimal for GPUs. Accessing the high-latency L1D cache and its victim cache can lead to significant latency overhead, severely degrading the performance of certain applications. We introduce HIVE, a high-priority victim cache designed to accelerate GPU memory accesses. HIVE handles memory requests first, before they reach the L1D cache. Our experimental results show that HIVE achieves an average performance improvement of $\mathbf{7 7. 1 \%}$ and $\mathbf{2 1. 7 \%}$ compared to the baseline and the state-of-the-art architecture, respectively. Yuhan Tang, Sheng Ma, Hanqing Li, Shengbai Luo, Jixuan Tang, Lizhou Wu |
DAC | 6 |
| 2025 | HeadTile: A Scalable and Efficient Accelerator for Large Language Model Inference with 3D Memory IntegrationabstractLarge Language Models (LLMs) inference have gained popularity over the past two years, driven by their high performance achieved through rapid increases in the number of parameters. The inference process of LLMs consists of two distinct stages: prefill and decode, each with unique computational characteristics. While existing neural network inference platforms, such as Google's TPU, perform well during the prefill stage, they often suffer from poor resource utilization during the decode stage. To address this challenge, we propose the scalable Headtile architecture, specifically designed to improve hardware resource utilization. By analyzing the inference behavior of LLMs, we examine how each layer executes on TPUv3 and introduce the Maarg paradigm in Headtile for inter-layer scheduling and mapping. Experimental results show that Headtile can achieve up to 24 × higher throughput in the decode stage compared to TPUv3. In addition, the Maarg paradigm reduces memory accesses by up to 60 % during the prefill stage. Qingshan Xue, Yihao Shi, Shengbai Luo, Xueyi Zhang 0001, Sheng Ma |
HPCC | 4 |
| 2025 | SpMARD: A Sparse-Sparse Matrix Multiplication Accelerator with Reconfigurable Dataflow for DNN WorkloadsabstractDeep learning becomes increasingly popular, and its main workload is Sparse-Sparse Matrix Multiplication (SpMSpM). Most SpMSpM accelerators usually only support a single dataflow. Different dataflows have different performance in different computing environments. Therefore, the single-dataflow accelerator cannot maintain the highest performance in all environments. Compared with single-dataflow accelerators, multi-dataflow accelerators provide flexible options for different workloads and improve the overall performance. Flexagon, Sparm, and SPADA are state-of-the-art multi-dataflow accelerators. However, the computation process of Flexagon and Sparm is not fully pipelined, and SPADA cannot support inner product dataflow. Additionally, Flexagon, Sparm, and SPADA cannot switch dataflows quickly and accurately. Inspired by these observations, we present SpMARD, a SpMSpM accelerator with reconfigurable dataflow. The computation process of SpMARD is fully pipelined, and SpMARD can support six dataflow variants simultaneously. Through the design of a Two-stage Pipeline Adder Network (TPAN) and a Position-based Psum Array (PPA), SpMARD can execute element-level merging, which can hide the merging overhead. Through the quantitative analysis of dataflows, we implement a Dataflow Switcher (DSwitcher), which can switch dataflows more efficiently. For the SpMSpM workload, the performance (GOPS) of the SpMARD we proposed is 1.27 times that of Flexagon, 1.18 times that of Sparm, and 1.22 times that of SPADA. Bo Wang 0159, Sheng Ma, Yunping Zhao, Shengbai Luo, Lizhou Wu, Dongsheng Li 0001, Zhuojun Chen |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Sparm: A Sparse Matrix Multiplication Accelerator Supporting Multiple DataflowsabstractAs the main workload of many scientific and machine learning applications, sparse matrix-matrix multiplication (spGEMM) has become a hot research field. The current spG EMM workloads exhibit sparsity and irregularity, leading to computational inefficiencies on traditional hardware platforms and motivating numerous customized accelerators. These spe-cialized hardware designs typically accelerate only one type of spGEMM dataflow (such as Inner-Product, Outer-Product, or Gustavson), yet the computational efficiency of the same spG EMM kernel can vary significantly under different dataflows. Flexagon is the first spGEMM accelerator to support multiple dataflows, but its MRN (Merger-Reduction Network) design causes a lot of data blocking and load imbalance, which limits its performance. In this work, we propose Sparm, which achieves efficient merging of psums (partial sums) for different dataflows through a specialized indexing unit. Sparm addresses the performance bottlenecks encountered by Flexagon when facing highly sparse matrices. Furthermore, Sparm employs a row/column prefetcher to load the streaming matrices proactively and thus significantly reduces the amount of DRAM access. We conduct simulations using a cycle-accurate simulator on workloads from various application domains, and the results demonstrate that Sparm achieves average performance gains of 2.62× and 1.35× compared to state-of-the-art spGEMM accelerators SIGMA and Flexagon. Meanwhile, Sparm brings only a small amount of additional hardware overhead over Flexagon. Shengbai Luo, Yihao Shi, Xueyi Zhang 0001, Qingshan Xue, Sheng Ma |
ASAP | 1 |
| 2024 | Understanding and Mitigating the Soft Error of Contrastive Language-Image Pre-training ModelsabstractIn recent years, MultiModal Large Language Models (MM-LLMs), based on the Contrastive Language-Image Pretraining models (CLIP), have achieved the best results in many fields. CLIP breaks through the gaps between language models and image models, realizes zero-shot image classification, and achieves excellent performance in tasks such as text-to-image generation, image style transformation, and long video generation. However, there are few studies on the fault tolerance of CLIP with soft errors, which hinders the application of multimodal large models in the field of security. Based on the analysis of the fault tolerance of common multimodal large models, we proposes a soft error mitigation framework. According to the experiments in this paper, the framework can effectively detect soft errors and mitigate the errors. Yihao Shi, Shengbai Luo, Qingshan Xue, Xueyi Zhang 0001, Sheng Ma |
ITC-Asia | 3 |
| 2024 | SparGD: A Sparse GEMM Accelerator with Dynamic DataflowabstractDeep learning has become a highly popular research field, and previously deep learning algorithms ran primarily on CPUs and GPUs. However, with the rapid development of deep learning, it was discovered that existing processors could not meet the specific large-scale computing requirements of deep learning, and custom deep learning accelerators have become popular. The majority of the primary workloads in deep learning are general matrix-matrix multiplications (GEMMs), and emerging GEMMs are highly sparse and irregular. The TPU and SIGMA are typical GEMM accelerators in recent years, but the TPU does not support sparsity, and both the TPU and SIGMA have insufficient utilization rates of the Processing Element (PE). We design and implement SparGD, a sparse GEMM accelerator with dynamic dataflow. SparGD has specific PE structures, flexible distribution networks and reduction networks, and a simple dataflow switching module. When running sparse and irregular GEMMs, SparGD can maintain high PE utilization while utilizing sparsity, and can switch to the optimal dataflow according to the computing environment. For sparse, irregular GEMMs, our experimental results show that SparGD outperforms systolic arrays by 30 times and SIGMA by 3.6 times. Bo Wang 0159, Sheng Ma, Shengbai Luo, Lizhou Wu, Chunyuan Zhang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | A Hybrid Kernel Pruning Approach for Efficient and Accurate CNNs
Xiao Yi, Shengbai Luo, Lizhou Wu, Kenli Li 0001, Sheng Ma |
ICA3PP (7) | 3 |