Huize Li

dblp:278/8981 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-8710-4472ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 8 first-author · 18 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DANMP: Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing Architecture
abstract
Multi-Scale Deformable Attention (MSDAttn) has become a fundamental component in various vision tasks due to its effective multi-scale grid sampling (MSGS). However, its reliance on random sampling results in highly irregular memory access patterns, making it a memory-intensive operation inefficient for GPUs. Near-memory processing (NMP) offers a promising solution for accelerating memory-bound kernels, yet existing NMP-based attention accelerators remain suboptimal for MSDAttn due to incompatible load balancing and data reuse strategies. Specifically, current NMP solutions uniformly distribute processing elements (PEs) across all banks, leading to significant PE underutilization and excessive cross-bank data transfers. Moreover, most rely on locality-based reuse, which fails under MSDAttn’s unpredictable sampling patterns.
Huize Li, Qinggang Wang, Bin Gao 0013, Dan Chen 0006, Yu Huang 0013, Xin Xin 0008
ICS1
2026 HBM-CASO: A Coordinated Approach to HBM System-Level and On-Die ECC
Ruizhi Zhu, Yanan Guo 0002, Huize Li, Weidong Cao 0001, Qian Lou, Xin Xin 0008
ISCA3
2026 HighP: In-Memory Acceleration of SpGEMM With High Bank-Level Parallelism
abstract
Generalized sparse matrix-matrix multiplication (SpGEMM) is a critical computational primitive that is highly memory-bound due to its inherent irregular data-dependent access pattern. Near-bank processing-in-memory (PIM) is a promising technique to overcome the memory bottleneck of SpGEMM by performing computations near the bank where the data is stored. However, earlier PIM studies fail to fully utilize the high memory bandwidth when performing SpGEMM due to low bank-level parallelism. As a result, 80% memory bandwidth is wasted as observed in our in-depth experimental analysis.Our key insight in this paper is that non-conflicting matrix columns in SpGEMM, where each row of these columns has no more than one non-zero element, can be processed simultaneously in different banks. We hence propose HighP, a near-bank PIM accelerator for SpGEMM with high bank-level parallelism. We first propose a set-based search mechanism, which finds non-conflicting columns through set operations automatically. We then develop a DIMM-based PIM architecture with detailed hardware and workflow designs for SpGEMM. Set operation logic and unified scratchpad memory management are designed to perform set operations with high computational parallelism and to enhance data reuse, respectively. HighP provides up to 17.88× performance improvement compared to the state-of-th-eart SpGEMM accelerator and achieves up to 8.19× performance improvement over the state-of-the-art PIM solution.
Dan Chen 0006, Huize Li, Huiying Lan, Zhaoying Li 0004, Pengcheng Yao, Tulika Mitra
IEEE Trans. Computers2
2025 SCREME: A Scalable Framework for Resilient Memory Design
abstract
The continuing advancement of memory technology has not only fueled a surge in performance, but also substantially exacerbated reliability challenges. Traditional solutions have primarily focused on improving the efficiency of protection schemes, i.e., Error Correction Codes (ECC), under the assumption that allocating additional memory space for ECC parity is always costly and therefore unsustainable as parity scales. We break the stereotype by proposing an orthogonal approach that provides additional, cost-effective memory space for resilient memory design. In particular, we recognize that ECC chips (used for parity storage) do not necessarily require the same performance level as regular data chips. This offers two-fold benefits: First, the bandwidth originally provisioned for a regularperformance ECC chip can instead be used to accommodate multiple low-performance chips. Second, the cost of ECC chips can be effectively reduced, as lower performance often correlates with lower expense. In addition, we observe that server-class memory chips are often provisioned with ample, yet underutilized I/O resources. This suggests an opportunity to repurpose these resources for flexible on-DIMM interconnections. Based on the above two insights, we finally propose SCREME, a scalable memory framework that leverages cost-effective yet slower chips - byproducts of rapid technology evolution - to meet the growing reliability demands driven by this evolution.
Mimi Xie, Yanan Guo 0002, Huize Li, Xin Xin 0008
PACT4
2025 ATLAS: Efficient Dynamic GNN System Through Abstraction-Driven Incremental Execution
Yu Huang 0013, Long Zheng 0003, Yang Wu 0010, Huize Li, Amelie Chi Zhou, Xiaofei Liao, Hai Jin 0001, Jingling Xue
APPT5
2025 Rewire: Advancing CGRA Mapping Through a Consolidated Routing Paradigm
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) balance the performance and power efficiency in computing systems. Effective compilers play a crucial role in fully realizing its potential. The compiler maps Data Flow Graphs (DFGs), which represent compute-intensive loop kernels, onto CGRAs. However, existing compilers often tackle DFG nodes individually, neglecting their intricate inter-dependencies. We introduce a novel mapping paradigm called Rewire that can place and route multiple nodes in one shot. Rewire first generates routing information that is shareable among multiple nodes via propagation. Then, Rewire intersects the routing information to generate individual placement candidates for each node. Finally, Rewire innovatively utilizes data dependencies as constraints to quickly find suitable placement for multiple nodes together. Our evaluation demonstrates that Rewire can generate more near-optimal mappings than prior works. Rewire achieves 2.1x and 1.3x performance improvement and 13.5x and 4.7x compilation time reduction, respectively, compared to two popular mappers.
Zhaoying Li 0004, Dhananjaya Wijerathne, Dan Chen 0006, Huize Li, Cheng Tan 0002, Tulika Mitra
DAC5
2025 SeIM: In-Memory Acceleration for Approximate Nearest Neighbor Search
abstract
Approximate nearest neighbor search (ANNS) is crucial in many applications to find semantically similar matches for user queries. Especially with the development of large language models (LLMs), ANNS is becoming increasingly important in retrieval-augmented generation (RAG). An in-depth analysis of ANNS reveals that its diverse operations, from extensive memory access to intensive sorting, are key performance bottlenecks, imposing significant strain on both the memory system and computing resources. Based on these observations, we present SeIM, a hierarchical in-memory architecture to accelerate ANNS. SeIM is designed to accommodate the diverse operational characteristics of ANNS. Specifically, SeIM offloads highly parallel memorybound operations to the memory bank level and introduces a unified execution model to reuse hardware units, requiring only lightweight modifications to standard DRAM architecture. Additionally, SeIM places compute-bound sorting operations, which require cross-unit data access, at the memory controller level and employs an adaptive transmission filtering technique to reduce unnecessary data transfers and processing during sorting. Our evaluation shows that SeIM achieves $268 \times 22 \times$, and $5 \times$ higher throughput, $306 \times 59 \times$, and $4 \times$ lower latency, and $3081 \times$, $287 \times$, and $2 \times$ higher power efficiency than state-of-the-art CPU-, GPU-, and ASIC-based ANNS solutions.
Chaoqiang Liu, Dan Chen 0006, Yu Huang 0013, Wenjing Xiao, Haifeng Liu 0003, Yi Zhang 0191, Huize Li, Xiaofei Liao, Hai Jin 0001
DAC7
2025 HyAtten: Hybrid Photonic-Digital Architecture for Accelerating Attention Mechanism
abstract
The wide adoption and substantial computational resource requirements of attention-based Transformers have spurred the demand for efficient hardware accelerators. Unlike digital-based accelerators, there is growing interest in exploring photonics due to its high energy efficiency and ultra-fast processing speeds. However, the significant signal conversion overhead limits the performance of photonic-based accelerators. In this work, we propose HyAtten, a photonic-based attention accelerator with minimize signal conversion overhead. HyAtten incorporates a signal comparator to classify signals into two categories based on whether they can be processed by low-resolution converters. HyAtten integrates low-resolution converters to process all low-resolution signals, thereby boosting the parallelism of photonic computing. For signals requiring high-resolution conversion, Hy-Atten uses digital circuits instead of signal converters to reduce area and latency overhead. Compared to state-of-the-art photonic-based Transformer accelerator, HyAtten achieves 9.8x performance/area and 2.2 x energy-efficiency/area improvement.
Huize Li, Dan Chen 0006, Tulika Mitra
DATE1
2025 PIM-SUM: Fast and Reliable In-Memory Summation for Recommendation Systems
abstract
Embedding aggregation in large-scale recommendation systems creates severe memory bandwidth bottlenecks, as each query sums many high-dimensional vectors. Bitwise-operation-based PIM can exploit subarray bandwidth, but traditional designs struggle with summation because long carry propagation limits parallelism. We propose PIM-SUM, an in-DRAM summation primitive for Sparse Length Sum (SLS) in recommendation workloads. PIM-SUM reformulates vector summation as column-wise accumulation via popcount, truncating carry propagation and avoiding redundant in-DRAM computation. This enables highthroughput summation using native DRAM bitwise primitives. PIM-SUM integrates a Reed-Solomon-inspired error correction at the DRAM row level. Its linearity supports parity propagation, enabling integrity checking and multi-bit error correction with low overhead. On DLRM workloads, PIM-SUM achieves up to$5.14 \times$speedup, logarithmic I/O reduction, and large energy savings. It also reduces silent data corruption by$1778 \times$and improves detection by over$170 \times$. These results show PIM-SUM is a scalable, fault-tolerant, and energy-efficient solution for memory-bound inference at data center scale.
Ruizhi Zhu, Huize Li, Di Wu 0016, Xin Xin 0008
ICCD3
2025 SADIMM: Accelerating $\underline{\text{S}}$S - parse $\underline{\text{A}}$A - ttention Using $\underline{\text{DIMM}}$DIMM - -Based Near-Memory Processing
abstract
Self-attention mechanism is the performance bottleneck of Transformer based language models. In response, researchers have proposed sparse attention to expedite Transformer execution. However, sparse attention involves massive random access, rendering it as a memory-intensive kernel. Memory-based architectures, such asnear-memory processing(NMP), demonstrate notable performance enhancements in memory-intensive applications. Nonetheless, existing NMP-based sparse attention accelerators face suboptimal performance due to hardware and software challenges. On the hardware front, current solutions employ homogeneous logic integration, struggling to support the diverse operations in sparse attention. On the software side, token-based dataflow is commonly adopted, leading to load imbalance after the pruning of weakly connected tokens. To address these challenges, this paper introduces SADIMM, a hardware-software co-designed NMP-based sparse attention accelerator. In hardware, we propose a heterogeneous integration approach to efficiently support various operations within the attention mechanism. This involves employing different logic units for different operations, thereby improving hardware efficiency. In software, we implement a dimension-based dataflow, dividing input sequences by model dimensions. This approach achieves load balancing after the pruning of weakly connected tokens. Compared to NVIDIA RTX A6000 GPU, the experimental results on BERT, BART, and GPT-2 models demonstrate that SADIMM achieves 48$\boldsymbol{\times}$, 35$\boldsymbol{\times}$, 37$\boldsymbol{\times}$speedups and 194$\boldsymbol{\times}$, 202$\boldsymbol{\times}$, 191$\boldsymbol{\times}$energy efficiency improvement, respectively.
Huize Li, Dan Chen 0006, Tulika Mitra
IEEE Trans. Computers1
2025 SPLIM: Bridging the Gap Between Unstructured SpGEMM and Structured In-Situ Computing
abstract
Sparse matrix-matrix multiplication (SpGEMM) is a critical kernel widely employed in machine learning and graph algorithms. However, high sparsity of real-world matrices makes SpGEMM memory-intensive. In-situ computing offers the potential to accelerate memory-intensive applications through high bandwidth and parallelism. Nevertheless, the irregular distribution of nonzeros renders software SpGEMM computation unstructured. In contrast, in-situ hardware platforms follow a fixed computation pattern, making them structured. The mismatch between unstructured software and structured hardware leads to suboptimal performance of current solutions. In this article, we propose SPLIM, a novel in-situ computing SpGEMM accelerator. SPLIM involves two innovations. First, we present a novel computation paradigm that converts SpGEMM into structured in-situ multiplication and unstructured accumulation. Second, we develop a unique coordinates alignment method utilizing in-situ search operations, effectively transforming unstructured accumulation into highly parallel search operations. Our experimental results demonstrate that SPLIM achieves$276\times $performance improvement and$687\times $energy saving compared to NVIDIA RTX A6000 GPU.
Huize Li, Dan Chen 0006, Tulika Mitra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 A ReRAM-Based Processing-In-Memory Architecture for Hyperdimensional Computing
abstract
Hyperdimensional computing (HDC) is a human brain-inspired computing paradigm that processes neural activity patterns with high dimensional vectors. Existing HDC accelerators usually utilize different hardware architectures to process encoding phases and comparison phases of HDC applications separately. They are unable to adapt to dynamic workloads for various datasets, resulting in resource underutilization. In this article, we propose a resistive random access memory (ReRAM)-based HDC accelerator called ReHDC for general HDC. We abstract the computing paradigms in encoding and comparison phases, and provide uniform primitive operators to efficiently process these two phases with the same hardware architecture. In the unified processing engine, ReHDC utilizes analog crossbar arrays to accelerate accumulation operations, and digital crossbar arrays to speed up high-dimensional element-wise operations (xor). Experimental results show that ReHDC can accelerate the HDC training by$69.4\times $and$1.93\times $, and can also improve the energy efficiency by$51.5\times $and$2.2\times $, compared with NVIDIA Tesla P100 GPU and the ReRAM-based HDC accelerator DUAL, respectively. Moreover, the performance speedup and energy efficiency for HDC inference are similar to that of HDC training.
Cong Liu 0028, Kaibo Wu, Haikun Liu, Hai Jin 0001, Xiaofei Liao, Zhuohui Duan, Huize Li, Yu Zhang 0027, Jing Yang 0051
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs
abstract
Efficiently supporting long context length is crucial for Transformer models. The quadratic complexity of the self-attention computation plagues traditional Transformers. Sliding window-based static sparse attention mitigates the problem by limiting the attention scope of the input tokens, reducing the theoretical complexity from quadratic to linear. Although the sparsity induced by window attention is highly structured, it does not align perfectly with the microarchitecture of the conventional accelerators, leading to sub-optimal implementation. In response, we propose a dataflow-aware FPGA-based accelerator design, SWAT, that efficiently leverages the sparsity to achieve scalable performance for long input. The proposed microarchitecture is based on a design that maximizes data reuse by using a combination of row-wise dataflow, kernel fusion optimization, and an input-stationary design considering the distributed memory and computation resources of FPGA. Consequently, it achieves up to 22× and 5.7× improvement in latency and energy efficiency compared to the baseline FPGA-based accelerator and 15× energy efficiency compared to GPU-based solution.
Zhenyu Bai, Pranav Dangi, Huize Li, Tulika Mitra
DAC3
2024 ASADI: Accelerating Sparse Attention Using Diagonal-based In-Situ Computing
abstract
The self-attention mechanism is the performance bottleneck of Transformer-based language models, particularly for long sequences. Researchers have proposed using sparse attention to speed up the Transformer. However, sparse attention introduces significant random access overhead, limiting computational efficiency. To mitigate this issue, researchers attempt to improve data reuse by utilizing row/column locality. Unfortunately, we find that sparse attention does not naturally exhibit strong row/column locality, but instead has excellent diagonal locality. Thus, it is worthwhile to use diagonal compression (DIA) format. However, existing sparse matrix computation paradigms struggle to efficiently support DIA format in attention computation. To address this problem, we propose ASADI, a novel software-hardware co-designed sparse attention accelerator. In the soft-ware side, we propose a new sparse matrix computation paradigm that directly supports the DIA format in self-attention computation. In the hardware side, we present a novel sparse attention accelerator that efficiently implements our computation paradigm using highly parallel in-situ computing. We thoroughly evaluate ASADI across various models and datasets. Our experimental results demonstrate an average performance improvement of 18.6 × and energy savings of 2.9× compared to a PIM-based baseline.
Huize Li, Zhaoying Li 0004, Zhenyu Bai, Tulika Mitra
HPCA1
2024 ReHarvest: An ADC Resource-Harvesting Crossbar Architecture for ReRAM-Based DNN Accelerators
abstract
ReRAM-based Processing-In-Memory (PIM) architectures have been increasingly explored to accelerate various Deep Neural Network (DNN) applications because they can achieve extremely high performance and energy-efficiency for in-situ analog Matrix-Vector Multiplication (MVM) operations. However, since ReRAM crossbar arrays’ peripheral circuits– analog-to-digital converters (ADCs) often feature high latency and low area efficiency, AD conversion has become a performance bottleneck of in-situ analog MVMs. Moreover, since each crossbar array is tightly coupled with very limited ADCs in current ReRAM-based PIM architectures, the scarce ADC resource is often underutilized. In this article, we propose ReHarvest, an ADC-crossbar decoupled architecture to improve the utilization of ADC resource. Particularly, we design a many-to-many mapping structure between crossbars and ADCs to share all ADCs in a tile as a resource pool, and thus one crossbar array can harvest much more ADCs to parallelize the AD conversion for each MVM operation. Moreover, we propose a multi-tile matrix mapping (MTMM) scheme to further improve the ADC utilization across multiple tiles by enhancing data parallelism. To support fine-grained data dispatching for the MTMM, we also design a bus-based interconnection network to multicast input vectors among multiple tiles, and thus eliminate data redundancy and potential network congestion during multicasting. Extensive experimental results show that ReHarvest can improve the ADC utilization by 3.2×, and achieve 3.5× performance speedup while reducing the ReRAM resource consumption by 3.1× on average compared with the state-of-the-art PIM architecture–FORMS.
Haikun Liu, Zhuohui Duan, Xiaofei Liao, Hai Jin 0001, Xiaokang Yang 0004, Huize Li, Cong Liu 0028, Fubing Mao, Yu Zhang 0027
ACM Trans. Archit. Code Optim.7
2024 CPSAA: Accelerating Sparse Attention Using Crossbar-Based Processing-In-Memory Architecture
abstract
The attention-based neural network attracts great interest due to its excellent accuracy enhancement. However, the attention mechanism requires huge computational efforts to process unnecessary calculations, significantly limiting the system’s performance. To reduce the unnecessary calculations, researchers propose sparse attention to convert some dense-dense matrices multiplication (DDMM) operations to sampled dense-dense matrix multiplication (SDDMM) and sparse matrix multiplication (SpMM) operations. However, current sparse attention solutions introduce massive off-chip random memory access since the sparse attention matrix is generally unstructured. We propose CPSAA, a novel crossbar-based processing-in-memory (PIM)-featured sparse attention accelerator to eliminate off-chip data transmissions. First, we present a novel attention calculation mode to balance the crossbar writing and crossbar processing latency. Second, we design a novel PIM-based sparsity pruning architecture to eliminate the pruning phase’s off-chip data transfers. Finally, we present novel crossbar-based SDDMM and SpMM methods to process unstructured sparse attention matrices by coupling two types of crossbar arrays. Experimental results show that CPSAA has an average of 89.6×, 32.2×, 17.8×, 3.39×, and 3.84× performance improvement and 755.6×, 55.3×, 21.3×, 5.7×, and 4.9× energy-saving when compare with GPU, FPGA, SANGER, ReBERT, and ReTransformer.
Huize Li, Hai Jin 0001, Long Zheng 0003, Xiaofei Liao, Yu Huang 0013, Cong Liu 0028, Zhuohui Duan, Dan Chen 0006, Chuangyi Gui
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 ReSMA: accelerating approximate string matching using ReRAM-based content addressable memory
abstract
Approximate string matching (ASM) functions as the basic operation kernel for a large number of string processing applications. Existing Von-Neumann-based ASM accelerators suffer from huge intermediate data with the ever-increasing string data, leading to massive off-chip data transmissions. This paper presents a novel ASM processing-in-memory (PIM) accelerator, namely ReSMA, based on ReCAM- and ReRAM-arrays to eliminate the off-chip data transmissions in ASM. We develop a novel ReCAM-friendly filter-and-filtering algorithm to process the q-grams filtering in ReCAM memory. We also design a new data mapping strategy and a new verification algorithm, which enables computing the edit distances totally in ReRAM crossbars for energy saving. Experimental results show that ReSMA outperforms the CPU-, GPU-, FPGA-, ASIC-, and PIM-based solutions by 268.7×, 38.6×, 20.9×, 707.8×, and 14.7× in terms of performance, and 153.8×, 42.2×, 31.6×, 18.3×, and 5.3× in terms of energy-saving, respectively.
Huize Li, Hai Jin 0001, Long Zheng 0003, Yu Huang 0013, Xiaofei Liao, Zhuohui Duan, Dan Chen 0006, Chuangyi Gui
DAC1
2022 ReGNN: a ReRAM-based heterogeneous architecture for general graph neural networks
abstract
Graph Neural Networks (GNNs) have both graph processing and neural network computational features. Traditional graph accelerators and NN accelerators cannot meet these dual characteristics of GNN applications simultaneously. In this work, we propose a ReRAM-based processing-in-memory (PIM) architecture called ReGNN for GNN acceleration. ReGNN is composed of analog PIM (APIM) modules for accelerating matrix vector multiplication (MVM) operations, and digital PIM (DPIM) modules for accelerating non-MVM aggregation operations. To improve data parallelism, ReGNN maps data to aggregation sub-engines based on the degree of vertices and the dimension of feature vectors. Experimental results show that ReGNN speeds up GNN inference by 228x and 8.4x, and reduces energy consumption by 305.2x and 10.5x, compared with GPU and the ReRAM-based GNN accelerator ReGraphX, respectively.
Cong Liu 0028, Haikun Liu, Hai Jin 0001, Xiaofei Liao, Yu Zhang 0027, Zhuohui Duan, Huize Li
DAC8
2022 ReCSA: a dedicated sort accelerator using ReRAM-based content addressable memory
Huize Li, Hai Jin 0001, Long Zheng 0003, Yu Huang 0013, Xiaofei Liao
Frontiers Comput. Sci.1
2020 ReSQM: Accelerating Database Operations Using ReRAM-Based Content Addressable Memory
abstract
The huge amount of data enforces great pressure on the processing efficiency of database systems. By leveraging the in-situ computing ability of emerging nonvolatile memory, processing-in-memory (PIM) technology shows great potential in accelerating database operations against traditional architectures without data movement overheads. In this article, we introduce ReSQM, a novel ReCAM-based accelerator, which can dramatically reduce the response time of database systems. The key novelty of ReSQM is that some commonly used database queries that would be otherwise processed inefficiently in previous studies can be in-situ accomplished with massively high parallelism by exploiting the PIM-enabled ReCAM array. ReSQM supports some typical database queries (such as SELECTION, SORT, and JOIN) effectively based on the limited computational mode of the ReCAM array. ReSQM is also equipped with a series of hardware-algorithm co-designs to maximize efficiency. We present a new data mapping mechanism that allows enjoying in-situ in-memory computations for SELECTION operating upon intermediate results. We also develop a count-based ReCAM-specific algorithm to enable the in-memory sorting without any row swapping. The relational comparisons are integrated for accelerating inequality join by making a few modifications to the ReCAM cells with negligible hardware overhead. The experimental results show that ReSQM can improve the (energy) efficiency by 611x (193x ), 19x (17x ), 59x (43x ), and 307x (181x) in comparison to a 10-core Intel Xeon E5-2630v4 processor for SELECTION, SORT, equi-join, and inequality join, respectively. In contrast to state-of-the-art CMOS-based CAM, GPU, FPGA, NDP, and PIM solutions, ReSQM can also offer 2.2x 39x speedups.
Huize Li, Hai Jin 0001, Long Zheng 0003, Xiaofei Liao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1