EDBT 2026 Demo / reviewers in the wild / expert
Xing Li 0031
dblp:26/379-31
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-4509-3859ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GIFTS: Efficient GCN Inference Framework on PyTorch-CPU via Exploring the SparsityabstractGraph Convolutional Networks (GCNs) are gaining attraction in AI research due to their ability to learn from graph data effectively. However, deploying GCNs on CPUs presents substantial challenges, as large-scale graph structures can lead to intensive computational and memory demands. While acceleration methods dedicatedly designed for GCNs on CPU platforms have emerged, they may not be the optimal solution as they tend to overlook the opportunity gained from the sparsity in the feature and adjacency matrices, leading to unsatisfactory memory and computation savings. In this paper, we propose a progressive two-step algorithm called GIFTS to accelerate GCN inference on CPUs by making use of the dynamic sparsity in the feature matrix and the static sparsity in the adjacency matrix. The first step introduces an online adaptive compression approach that performs a selective value-level compression on the part of feature vectors that will be frequently accessed during GCN inference. To further reduce the redundant bit width in feature vectors, we employ a bit-level pruning approach to narrow down bandwidth. The second step designs an offline degree-aware scheduling approach, which aims at balancing workloads caused by the irregular sparsity in the adjacency matrix during the GCN training process, given the observation that the adjacency matrix remains unchanged during inference. This step distributes workloads based on the nodes' degree in a coarse-grained fashion, considering both workload distribution and data locality. Through comprehensive evaluation, we demonstrate that GIFTS consistently outperforms the PyTorch Geometric (PyG) and DistGNN frameworks in terms of execution time, while incurring negligible accuracy loss. The source code is available at https://github.com/ACA-Lab/GIFTS. Xing Li 0031, Xiaoyao Liang, Zhuoran Song |
IPDPS | 2 |
| 2025 | An Efficient Bit-Sparse DNN Accelerator Exploiting Adaptive Bit-Serial ComputationsabstractBit sparsity, an intrinsic attribute of binary representation, has been widely utilized in DNN inference acceleration. Despite the advantages in performance and energy efficiency demonstrated by existing bit-serial-based bit-sparse accelerators, they still face two notable limitations: 1) At the low-level bit-serial multiplier level, existing methods either statically select weight or activation as the serialized object during the design phase, or simply serialize both without considering the distribution of non-zero bits in different operands, thereby failing to achieve optimal performance; 2) At the high-level dataflow level, existing approaches do not eliminate zero values in data movement and computation, leading to considerable energy and latency overhead, as well as suboptimal PE utilization. In this work, we propose AdaS-Pro accelerator for fast and energy-efficient DNN inference. At the multiplier level, AdaSPro employs an adaptive bit-serial computation scheme, which dynamically serializes the input operand with fewer non-zero bits at runtime, thereby minimizing compute cycles. To further enhance performance, AdaS-Pro introduces an improved Booth encoding method to reduce the number of non-zero bits in each operand. At the dataflow level, AdaS-Pro employs a compressed format to eliminate zero values and proposes a bi-directional inner-join unit coupled with a ring-shaped scheduler to achieve efficient non-zero workload extraction and balancing. Experimental results show that AdaS-Pro outperforms existing state-of-theart bit-sparse accelerators, such as BitLet, BitX, and Laconic, with performance improvements of 4.03×, 6.78×, and 1.43×, respectively. Jiayao Ling, Gang Li 0015, Xiaolong Lin, Xing Li 0031, Jian Cheng 0001, Xiaoyao Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Vision Transformer Acceleration via a Versatile Attention Optimization FrameworkabstractVision Transformers (ViTs) have achieved remarkable success across various tasks. However, their deployment is hindered by challenges, such as high memory requirements, long inference latency, and significant power consumption. To address these challenges, existing works optimize one of the two key stages on ViT’s critical path: linear projection or self-attention. Regrettably, we have noticed that both linear projection and self-attention can potentially become bottlenecks as the input image resolution varies, which makes the existing approaches lack generality. Accordingly, in this article, we propose a versatile attention optimization framework. On the algorithm side, we present a SpQuant algorithm that sparsifies weight matrices offline and input matrices online during linear projection as well as tunes the bit-width of the probabilities matrix according to their importance. On the hardware side, we design SQArch architectures to improve the performance of the SpQuant algorithm. The proposed SQArch architecture offers a low-cost preprocess module that predicts and prunes nonkey elements of the input matrix on the fly. Moreover, we design a compute module that supports sparse-sparse matrix multiplications (SpMSpM) and multiple precision computations on a single systolic array for generality. Furthermore, we can address the underutilization and workload imbalance problems by 1) decoupling the rows in the systolic array for enough flexibility and 2) proposing a workload balance scheme for SpMSpM that allows the array to accept data of similar sparsity, thereby reducing synchronization between computing units. Extensive experiment results demonstrate that SQArch can achieve satisfactory performance speedups and energy saving compared to state-of-the-art designs. Xuhang Wang, Qiyue Huang, Xing Li 0031, Haozhe Jiang, Qiang Xu 0001, Xiaoyao Liang, Zhuoran Song |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | MoC: A Morton-Code-Based Fine-Grained Quantization for Accelerating Point Cloud Neural NetworksabstractPoint Cloud Neural Network (PCNN) plays an essential role in various 3D applications, with some of them even being time-sensitive and safety-critical. However, the large scale of unordered points with lengthy features results in heavy computational workloads, making them far from real-time processing. To address this challenge, we propose MoC, a Morton-code-based fine-grained quantization for accelerating PCNNs. Specifically, we utilize Morton code to capture the spatial locality among points. Then, we gather nearby points with similar features into a region. Considering the similarity in features of nearby points, we propose to decompose features into base and offsets, where the offsets fall within a narrow range. Building upon this, we introduce a two-level mixed-precision quantization. In the first level, we quantize offsets with low precision, while keeping the base in high precision to ensure accuracy. For the second level, noticing the different data distribution of offsets across various regions, we employ two types of low precision at the region level, which provides opportunities to further accelerate feature computations. To support our algorithm, we design a hardware architecture that parallelizes the Morton code path with the critical path. In our extensive experiments on various datasets, our algorithm-architecture co-designed method demonstrates 12X, 6.3X, 4.7X, 3.8X, 3.4X and 2.8X speedup and 19.3X, 9.7X, 6.0X, 5.2X, 4.6X and 4.1X energy savings over CPU, Server and Edge GPUs, state-of-the-art ASICs (incl. PointAcc, MARS, PRADA) with negligible accuracy loss. Xueyuan Liu 0001, Zhuoran Song, Hao Chen 0126, Xing Li 0031, Xiaoyao Liang |
DAC | 4 |
| 2024 | Sava: A Spatial- and Value-Aware Accelerator for Point Cloud TransformerabstractPoint Cloud Transformer is undergoing a rising trend in both industry and academia. It aligns traditional point cloud feature extraction methods with the latest transformer architecture and achieves remarkable performance. However, accelerators for traditional point cloud neural networks (PCNNs) and those solely for transformers fail to capture the characteristics of point cloud transformers, thus exhibiting poor performance. To address this challenge, we propose Sava, a co-designed accelerator that adopts a spatial- and value-aware hybrid pruning strategy for point cloud transformers. In terms of the spatial domain, we observe that points in regions of various densities exhibit different levels of importance. In the value space, a minor input contributes less to features, indicating lower importance. Considering both perspectives, we hybridize the information inherited from the spatial and value spaces to prune less significant values in attention, which converts data to sparse patterns and makes it readily accelerated. Furthermore, we adopt low-bit quantization to boost computations and apply varying quantization precisions across different network layers based on their sensitivity. In support of our algorithm, we propose an architecture that employs a configurable mixed-precision systolic array for various computing loads under diverse precisions. To address the workload imbalance of the unstructured sparse computations, we introduce a data rearrangement mechanism, which improves resource utilization while hiding latency. We evaluate our Sava on four point cloud transformer models and achieve notable accuracy and performance gains. In comparison with CPU, GPUs, and ASICs, our Sava offers 10.3×, 3.6×, 3.3×, 2.6×, 2.2× speedup, along with 20×, 8.8×, 6.9×, 3.2×, 2.4× energy savings on average. Xueyuan Liu 0001, Zhuoran Song, Xing Li 0031, Tao Yang 0031, Fangxin Liu, Xiaoyao Liang |
DATE | 4 |
| 2024 | Watt: A Write-Optimized RRAM-Based Accelerator for Attention
Xuan Zhang 0001, Zhuoran Song, Xing Li 0031, Zhezhi He, Naifeng Jing, Li Jiang 0002, Xiaoyao Liang |
Euro-Par (2) | 3 |
| 2024 | Early: An Importance-Aware Early Firing and Exit for SNN AccelerationabstractSpiking neural networks (SNNs) have been promising applications in the image recognition domain, and their key component is the spiking neuron. SNN s mainly contain integration and firing processes, which are essentially weight accumulation and threshold comparison, respectively. However, spike trains of the neurons exhibit high sparsity and irregularity in both temporal and spatial domains, leading to inefficient memory access and computation. Therefore, designing an efficient accelerator for SNNs is urgent. This paper presents an elaborate accelerator Early in a software-hardware co-design way. At the software level: (i) Noticing the importance of weights, where larger weights disproportionately affect the membrane potential, we devise a weight importance-aware early firing solution for the firing neurons. It prioritizes the accumulation of these large weights, thereby accelerating the membrane potential's rise to surpass the threshold sooner. (ii) Meanwhile, given the observation that a large proportion of neurons do not eventually be fired even after experiencing a long delay of weight accumulation, we propose a weight importance-aware early exit mechanism. It preferentially accumulates large weights and compares the membrane potential with the predetermined threshold, which early halts the accumulation of neurons that are unlikely to be fired, enhancing efficiency. At the hardware level, we design a specialized processing element (PE) featuring the reorder engine for spikes and weights, tailored to realize the aforementioned strategies. Experimental results show that Early averagely achieves 20.3 x, 6.5 x, and 2.4 x speedup compared to the state-of-the-art accelerators Spinalflow, PTB, and SATO. Meanwhile, it averagely achieves 25.2x, 7.4x, and 3.2x energy savings with respect to the three accelerators. Xuan Zhang 0001, Zhuoran Song, Peng Zhou 0030, Xing Li 0031, Xueyuan Liu 0001, Xiaolong Lin, Zhezhi He, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang |
ICCD | 4 |
| 2024 | Janus: A Flexible Processing-in-Memory Graph Accelerator Toward SparsityabstractGraph application is ever-growing in relational data analysis. However, the memory access patterns become the performance bottleneck in graph analytics and graph neural network (GNN) suffering from single-side and dual-side sparsity, separately. Existing resistive random access memory (RRAM)-based processing-in-memory accelerators reduce data movements but fail to handle both types of sparsity in graph data. To address these issues, our work introduces Janus, a flexible highly compact architecture that is capable of being configured to enable single-sparse mode and dual-sparse mode, to accelerate graph analytics and GNN workloads in compressed mapping, respectively. Upon performing graph analytics with single-side sparsity, Janus employs a tandem-isomorphic-crossbar design both to remove zero-stored footprint, and to eliminate redundant search and sequential indexing. To address the challenge of dual-side sparsity in GNN, Janus still takes a random index access mechanism to gather data rapidly and uses a semi-SPM2 compute paradigm to boost the RRAM-based analog multiplication-and-accumulation in the compressed format. Compared with the state-of-the-art works, Janus outperforms them in both performance and energy efficiency for graph analytics and GNN, respectively. Xing Li 0031, Zhuoran Song, Rachata Ausavarungnirun, Xiao Liu 0033, Xueyuan Liu 0001, Xuan Zhang 0001, Xuhang Wang, Jiayao Ling, Gang Li 0015, Naifeng Jing, Xiaoyao Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | HyAcc: A Hybrid CAM-MAC RRAM-based Accelerator for Recommendation ModelabstractThe deep learning recommendation model (DLRM) plays a crucial role in online services, whose key component is the embedding layer. The embedding layer is to gather and reduce several rows of embedding vectors from the large embedding tables given the input item IDs, which poses challenges due to its memory-intensive nature and becomes a focus of current DLRM accelerators. One potential solution for accelerating DLRM is the use of resistive random access memory (RRAM), which exploits process-in-memory (PIM) capability. However, current RRAM-based DLRM accelerators encounter issues with expensive serial embedding vector searches.Accordingly, this paper proposes a Hybrid CAM-MAC RRAM-based Accelerator (HyAcc) to address the challenges of the embedding layer. Firstly, we recognize that content-addressable-memory (CAM) crossbar can broadcast the input item IDs across all rows to gather the stored item IDs at one cycle. Hence, we design RRAM-based CAM crossbars to gather item IDs efficiently. In the meantime, we utilize the multiplication-and-accumulation (MAC) crossbars to implement the reduction operation in the embedding layer. Whereas, during the gather operation, the RRAM-based CAM crossbar inevitably encounters the access inefficiency problem because only one item ID can be gathered per cycle. To overcome this, we propose the hot/cold item engines containing fine-grained/coarse-grained CAM crossbars for the input item IDs with high-frequency/low-frequency (termed as hot/cold item IDs). Additionally, since the input cold item IDs are unevenly distributed in the coarse-grained CAM crossbars, they may cause the workload imbalance problem. To alleviate it, we present the access-aware dynamic pruning solution to dynamically prune the redundant input cold item IDs and average the workload of the coarse-grained CAM crossbars. Extensive experiments validate the effectiveness of the proposed HyAcc architecture. Xuan Zhang 0001, Zhuoran Song, Xing Li 0031, Zhezhi He, Li Jiang 0002, Naifeng Jing, Xiaoyao Liang |
ICCD | 3 |
| 2022 | Gzippo: Highly-Compact Processing-in-Memory Graph Accelerator Alleviating Sparsity and RedundancyabstractGraph application plays a significant role in real-world data computation. However, the memory access patterns become the performance bottleneck of the graph applications, which include low compute-to-communication ratio, poor temporal locality, and poor spatial locality. Existing RRAM-based processing-in-memory accelerators reduce the data movements but fail to address both sparsity and redundancy of graph data. In this work, we present Gzippo, a highly-compact design that supports graph computation in the compressed sparse format. Gzippo employs a tandem-isomorphic-crossbar architecture both to eliminate redundant searches and sequential indexing during iterations, and to remove sparsity leading to non-effective computation on zero values. Gzippo achieves a 3.0× (up to 17.4×) performance speedup, 23.9× (up to 163.2×) energy efficiency over state-of-the-art RRAM-based PIM accelerator, respectively. Xing Li 0031, Rachata Ausavarungnirun, Xiao Liu 0033, Xueyuan Liu 0001, Xuan Zhang 0001, Zhuoran Song, Naifeng Jing, Xiaoyao Liang |
ICCAD | 1 |
| 2022 | GCNTrain: A Unified and Efficient Accelerator for Graph Convolutional Neural Network TrainingabstractGraph convolutional neural networks (GCNs) have been emerging as a promising category of neural network models for extending deep learning to graph data analytics. Serving as a type of semi-supervised models, GCNs need training before being used to extract any input graph’s features. The challenge is that the existing GCN accelerators often target the sparse-dense matrix multiplications (SpDM) in GCN inference while ignoring the compute-intensive GCN training. Obviously, this poses momentous performance demands and design challenges.In this paper, we categorize the computations of GCN training into sparse-sparse matrix multiplications (SpSpM) and sparse-dense matrix multiplications (SpDM); and then introduce the GCNTrain-v1 architecture that uniformly performs both SpSpM and SpDM by the column-wise-product-based method. To ad-dress the bank conflict problem in the GCNTrain-v1 architecture, we further propose the GCNTrain-v2 architecture with the conflict-free bank access strategy. This strategy is able to coalesce all requests to one bank by broadcasting elements. Moreover, to alleviate the workload imbalance problem in the GCNTrain-v2 architecture, we offer the GCNTrain-v3 architecture with the offline reshuffle technique that offline reshuffles and balances the non-zero elements in the matrix before GCN training. Overall, the GCNTrain-v3 architecture implements both SpSpM and SpDM for accelerating GCN training without bank conflict and work-load imbalance problems. On five graph datasets, experiment results demonstrate considerable performance speedups over CPU (80.47×), GPU (10.88×), and GCNAX (1.65×). Zhuoran Song, Xing Li 0031, Naifeng Jing, Xiaoyao Liang |
ICCD | 3 |