Chaoqiang Liu

dblp:50/474 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Gopher: Efficient Dynamic Graph Pattern Mining via DAG-Driven Execution
abstract
Graph pattern mining is essential for analyzing dynamic networks, where graphs evolve over time. To accommodate these changes, existing solutions update match sets incrementally, avoiding the need to re-mine the entire graph and achieving significant performance improvements. However, these methods suffer from inefficiencies due to redundant set intersection operations across subgraph instances, causing performance degradation.
Yi Zhang 0191, Yu Huang 0013, Chaoqiang Liu, Haifeng Liu 0003, Jingrui Yuan, Jianhui Yue, Xiaofei Liao, Hai Jin 0001, Jingling Xue
EuroSys3
2026 Meridian: In-Memory Acceleration for RAG with Document Attention Decomposition
Chaoqiang Liu, Yu Huang 0013, Haifeng Liu 0003, Yi Zhang 0191, Qihang Qiu, Xueqi Li 0001, Xiaofei Liao, Hai Jin, Jingling Xue
ISCA1
2025 SeIM: In-Memory Acceleration for Approximate Nearest Neighbor Search
abstract
Approximate nearest neighbor search (ANNS) is crucial in many applications to find semantically similar matches for user queries. Especially with the development of large language models (LLMs), ANNS is becoming increasingly important in retrieval-augmented generation (RAG). An in-depth analysis of ANNS reveals that its diverse operations, from extensive memory access to intensive sorting, are key performance bottlenecks, imposing significant strain on both the memory system and computing resources. Based on these observations, we present SeIM, a hierarchical in-memory architecture to accelerate ANNS. SeIM is designed to accommodate the diverse operational characteristics of ANNS. Specifically, SeIM offloads highly parallel memorybound operations to the memory bank level and introduces a unified execution model to reuse hardware units, requiring only lightweight modifications to standard DRAM architecture. Additionally, SeIM places compute-bound sorting operations, which require cross-unit data access, at the memory controller level and employs an adaptive transmission filtering technique to reduce unnecessary data transfers and processing during sorting. Our evaluation shows that SeIM achieves $268 \times 22 \times$, and $5 \times$ higher throughput, $306 \times 59 \times$, and $4 \times$ lower latency, and $3081 \times$, $287 \times$, and $2 \times$ higher power efficiency than state-of-the-art CPU-, GPU-, and ASIC-based ANNS solutions.
Chaoqiang Liu, Dan Chen 0006, Yu Huang 0013, Wenjing Xiao, Haifeng Liu 0003, Yi Zhang 0191, Huize Li, Xiaofei Liao, Hai Jin 0001
DAC1
2025 HeterRAG: Heterogeneous Processing-in-Memory Acceleration for Retrieval-augmented Generation
abstract
By integrating external knowledge bases, Retrieval-augmented Generation (RAG) enhances natural language generation for knowledgeintensive scenarios and specialized domains, producing content that is both more informative and personalized.RAG systems typically consist of two fundamental stages: retrieval and generation.The retrieval stage experiences low bandwidth utilization due to its random and irregular memory access patterns.Meanwhile, the generation stage is also constrained by memory bandwidth limitations, which arise from involving a significant number of General Matrix-Vector Multiplications (GEMV) operations.These two stages collectively lead to memory bottlenecks within RAG systems.Recent efforts leverage HBM-based Processing-in-Memory (PIM) to accelerate conventional Large Language Models (LLMs).However, the retrieval stage incurs substantial storage overhead due to the need to maintain large-scale knowledge bases, resulting in a capacity bottleneck.Solely relying on HBM-based PIM in RAG is both costly and insufficient to meet the capacity demands.Fortunately, DIMM-based PIM provides a low-cost, high-capacity alternative that complements HBM.In this work, we propose HeterRAG, a novel heterogeneous PIM acceleration system for RAG.It combines
Chaoqiang Liu, Haifeng Liu 0003, Dan Chen 0006, Yu Huang 0013, Yi Zhang 0191, Wenjing Xiao, Xiaofei Liao, Hai Jin 0001
ISCA1
2025 Cycle-conditional diffusion model for noise correction of diffusion-weighted images using unpaired data
abstract
Diffusion-weighted imaging (DWI) is a key modality for studying brain microstructure, but its signals are highly susceptible to noise due to the thermal motion of water molecules and interactions with tissue microarchitecture, leading to significant signal attenuation and a low signal-to-noise ratio (SNR). In this paper, we propose a novel approach, a Cycle-Conditional Diffusion Model (Cycle-CDM) using unpaired data learning, aimed at improving DWI quality and reliability through noise correction. Cycle-CDM leverages a cycle-consistent translation architecture to bridge the domain gap between noise-contaminated and noise-free DWIs, enabling the restoration of high-quality images without requiring paired datasets. By utilizing two conditional diffusion models, Cycle-CDM establishes data interrelationships between the two types of DWIs, while incorporating synthesized anatomical priors from the cycle translation process to guide noise removal. In addition, we introduce specific constraints to preserve anatomical fidelity, allowing Cycle-CDM to effectively learn the underlying noise distribution and achieve accurate denoising. Our experiments conducted on simulated datasets, as well as children and adolescents' datasets with strong clinical relevance. Our results demonstrate that Cycle-CDM outperforms comparative methods, such as U-Net, CycleGAN, Pix2Pix, MUNIT and MPPCA, in terms of noise correction performance. We demonstrated that Cycle-CDM can be generalized to DWIs with head motion when they were acquired using different MRI scannsers. Importantly, the denoised DWI data produced by Cycle-CDM exhibit accurate preservation of underlying tissue microstructure, thus substantially improving their medical applicability.
Pengli Zhu, Chaoqiang Liu, Yingji Fu, Nanguang Chen, Anqi Qiu
Medical Image Anal.2
2024 Enabling Efficient Large Recommendation Model Training with Near CXL Memory Processing
abstract
Personalized recommendation systems have become one of the most important Internet services nowadays. A critical challenge of training and deploying the recommendation models is their high memory capacity and bandwidth demands, with the embedding layers occupying hundreds of GBs to TBs of storage. The advent of memory disaggregation technology and Compute Express Link (CXL) provides a promising solution for memory capacity scaling. However, relocating memory-intensive embedding layers to CXL memory incurs noticeable performance degradation due to its limited transmission bandwidth, which is significantly lower than the host memory bandwidth. To address this, we introduce ReCXL, a CXL memory disaggregation system that utilizes near-memory processing for scalable, efficient recommendation model training. ReCXL features a unified, hardwareefficient NMP architecture that processes the entire embedding training within CXL memory, minimizing data transfers over the bandwidth-limited CXL and enhancing internal bandwidth. To further improve the performance, ReCXL incorporates softwarehardware co-optimizations, including sophisticated dependencyfree prefetching and fine-grained update scheduling, to maximize hardware utilization. Evaluation results show that ReCXL outperforms the CPU-GPU baseline and the naïve CXL memory by $7.1 \times \sim 10.6 \times(9.4 \times$ on average) and $12.7 \times \sim 31.3 \times(22.6 \times$ on average), respectively.
Haifeng Liu 0003, Long Zheng 0003, Yu Huang 0013, Chaoqiang Liu, Xiaofei Liao, Hai Jin 0001, Jingling Xue
ISCA5
2024 L-FNNG: Accelerating Large-Scale KNN Graph Construction on CPU-FPGA Heterogeneous Platform
abstract
Due to the high complexity of constructing exact k -nearest neighbor graphs, approximate construction has become a popular research topic. The NN-Descent algorithm is one of the representative in-memory algorithms. To effectively handle large datasets, existing state-of-the-art solutions combine the divide-and-conquer approach and the NN-Descent algorithm, where large datasets are divided into multiple partitions, and a subgraph is constructed for each partition before all the subgraphs are merged, reducing the memory pressure significantly. However, such solutions fail to address inefficiencies in large-scale k -nearest neighbor graph construction. In this paper, we propose L-FNNG, a novel solution for accelerating large-scale k -nearest neighbor graph construction on CPU-FPGA heterogeneous platform. The CPU is responsible for dividing data and determining the order of partition processing, while the FPGA executes all construction tasks to utilize the acceleration capability fully. To accelerate the execution of construction tasks, we design an efficient FPGA accelerator, which includes the Block-based Scheduling (BS) and Useless Computation Aborting (UCA) techniques to address the problems of memory access and computation in the NN-Descent algorithm. We also propose an efficient scheduling strategy that includes a KD-tree-based data partitioning method and a hierarchical processing method to address scheduling inefficiency. We evaluate L-FNNG on a Xilinx Alveo U280 board hosted by a 64-core Xeon server. On multiple large-scale datasets, L-FNNG achieves, on average, 2.3× construction speedup over the state-of-the-art GPU-based solution.
Chaoqiang Liu, Xiaofei Liao, Long Zheng 0003, Yu Huang 0013, Haifeng Liu 0003, Yi Zhang 0191, Haiheng He, Haoyan Huang, Hai Jin 0001
ACM Trans. Reconfigurable Technol. Syst.1
2023 FNNG: A High-Performance FPGA-based Accelerator for K-Nearest Neighbor Graph Construction
abstract
The k-nearest neighbor graph has emerged as the key data structure for many critical applications. However, it can be notoriously challenging to construct k-nearest neighbor graphs over large graph datasets, especially with a high-dimensional vector feature. Many solutions have been recently proposed to support the construction of k-nearest neighbor graphs. However, these solutions involve substantial memory access and computational overheads and an architecture-level solution is still absent. To address these issues, we architect FNNG, the first FPGA-based accelerator to support k-nearest neighbor graph construction. Specifically, FNNG is equipped with the block-based scheduling technique to exploit the inherent data locality between vertices. It divides the vertices that are close in space into blocks and process the vertices according to the granularity of the blocks during the construction process. FNNG also adopts the useless computation aborting technique to identify superfluous useless computations. It keeps the existing maximum similarity values of all vertices inside the computing unit. In addition, we propose an improved architecture in order to fully utilize both techniques. We implement FNNG on the Xilinx Alveo U280 FPGA card. The results show that FNNG achieves 190x and 2.1x speedups over the state-of-the-art CPU and GPU solutions, running on Intel Xeon Gold 5117 CPU and NVIDIA GeForce RTX 3090 GPU, respectively.
Chaoqiang Liu, Haifeng Liu 0003, Long Zheng 0003, Yu Huang 0013, Xiangyu Ye, Xiaofei Liao, Hai Jin 0001
FPGA1
2023 GraphMetaP: Efficient MetaPath Generation for Dynamic Heterogeneous Graph Models
abstract
Metapath-based heterogeneous graph models (MHGM) show excellent performance in learning semantic and structural information in heterogeneous graphs. Metapath matching is an essential processing step in MHGM to find all metapath instances, bringing significant overhead compared to the total model execution time. Even worse, in dynamic heterogeneous graphs, metapath instances require to be rematched while graph updated. In this paper, we observe that only a small fraction of metapath instances change and propose GraphMetaP, an efficient incremental metapath maintenance method in order to eliminate the matching overhead in dynamic heterogeneous graphs. GraphMetaP introduces a novel format for metapath instances to capture the dependencies among the metapath instances. The format incrementally maintains metapath instances based on the graph updates to avoide the rematching metapath overhead for the updated graph. Furthermore, GraphMetaP uses the fold way to simplify the format in order to recover all metapath instances faster. Experiments show that GraphMetaP enables efficient maintenance of metapath instances on dynamic heterogeneous graphs and outperforms 172.4X on average compared to the matching metapath method.
Haiheng He, Dan Chen 0006, Long Zheng 0003, Yu Huang 0013, Haifeng Liu 0003, Chaoqiang Liu, Xiaofei Liao, Hai Jin 0001
IPDPS6
2023 Accelerating Personalized Recommendation with Cross-level Near-Memory Processing
abstract
The memory-intensive embedding layers of the personalized recommendation systems are the performance bottleneck as they demand large memory bandwidth and exhibit irregular and sparse memory access patterns. Recent studies propose near memory processing (NMP) to accelerate memory-bound embedding operations. However, due to the load imbalance caused by the skewed access frequency of the embedding data, existing NMP solutions that exploit fine-grained memory parallelism fail to translate the increasingly massive internal bandwidth to performance improvements, leading to resource underutilization and hardware overhead.
Haifeng Liu 0003, Long Zheng 0003, Yu Huang 0013, Chaoqiang Liu, Xiangyu Ye, Jingrui Yuan, Xiaofei Liao, Hai Jin 0001, Jingling Xue
ISCA4
2023 Monte Carlo Ensemble Neural Network for the diagnosis of Alzheimer's disease
Chaoqiang Liu, Anqi Qiu
Neural Networks1
2022 A deep feature alignment adaptation network for rolling bearing intelligent fault diagnosis
Hongkai Jiang, Chaoqiang Liu
Adv. Eng. Informatics5
2022 A deep feature enhanced reinforcement learning method for rolling bearing fault diagnosis
Hongkai Jiang, Chaoqiang Liu
Adv. Eng. Informatics5
2022 Data-augmented wavelet capsule generative adversarial network for rolling bearing fault diagnosis
Hongkai Jiang, Chaoqiang Liu, Wangfeng Yang
Knowl. Based Syst.3
2022 A new data generation approach with modified Wasserstein auto-encoder for rotating machinery fault diagnosis with limited fault data
Hongkai Jiang, Chaoqiang Liu
Knowl. Based Syst.3
2022 A Flexible Yet Efficient DNN Pruning Approach for Crossbar-Based Processing-in-Memory Architectures
abstract
Pruning deep neural networks (DNNs) can reduce the model size and thus save hardware resources of a resistive-random-access-memory (ReRAM)-based DNN accelerator. For the tightly coupled crossbar structure, existing ReRAM-based pruning techniques prune the weights of a DNN in a structured manner, thereby attaining low pruning ratios. This article presents a novel pruning technique, SegPrune, for pruning the weights of a DNN flexibly on crossbar architectures in order to maximize the pruning ratio achieved while preserving crossbar efficiency. We observe that different filters of a weight matrix share a large number of matrix subcolumns (in the same rows), called segments, that can be pruned by using the same segment shape in the sense that the weights at the same column position of these segments are either simultaneously accuracy-sensitive (and should thus be reserved) or simultaneously accuracy-insensitive (and can thus be pruned). Due to the bit-line exchangeability in the crossbar, segments with the same pruning shape can be assembled together into the same crossbar to ensure crossbar execution efficiency. We propose a projection-based shape voting algorithm to select suitable segment shapes to drive the weight pruning process. Accordingly, we also introduce a low-overhead data path that can be easily integrated into any existing ReRAM-based DNN accelerator, achieving a high pruning ratio and a high execution efficiency. Our evaluation shows that SegPrune outperforms the state-of-the-art, Hybrid-P, and FORMAS, by up to$14.6\times $and$3.6\times $in pruning ratio,$13.9\times $and$3.4\times $in inference speedup, and$12.5\times $and$3.1\times $in energy reduction, respectively, while achieving an even higher accuracy at the cost of less than 0.27% extra hardware area overhead.
Long Zheng 0003, Haifeng Liu 0003, Yu Huang 0013, Dan Chen 0006, Chaoqiang Liu, Haiheng He, Xiaofei Liao, Hai Jin 0001, Jingling Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2021 Fast vertex-based graph convolutional neural network and its application to brain images
Chaoqiang Liu, Hui Ji 0002, Anqi Qiu
Neurocomputing1
2012 Framelet-Based Blind Motion Deblurring From a Single Image
abstract
How to recover a clear image from a single motion-blurred image has long been a challenging open problem in digital imaging. In this paper, we focus on how to recover a motion-blurred image due to camera shake. A regularization-based approach is proposed to remove motion blurring from the image by regularizing the sparsity of both the original image and the motion-blur kernel under tight wavelet frame systems. Furthermore, an adapted version of the split Bregman method is proposed to efficiently solve the resulting minimization problem. The experiments on both synthesized images and real images show that our algorithm can effectively remove complex motion blurring from natural images without requiring any prior information of the motion-blur kernel.
Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen
IEEE Trans. Image Process.3
2011 A Variational Model for Segmentation of Overlapping Objects With Additive Intensity Value
abstract
We propose a variant of the Mumford-Shah model for the segmentation of a pair of overlapping objects with additive intensity value. Unlike standard segmentation models, it does not only determine distinct objects in the image, but also recover the possibly multiple membership of the pixels. To accomplish this, some a priori knowledge about the smoothness of the object boundary is integrated into the model. Additivity is imposed through a soft constraint which allows the user to control the degree of additivity and is more robust than the hard constraint. We also show analytically that the additivity parameter can be chosen to achieve some stability conditions. To solve the optimization problem involving geometric quantities efficiently, we apply a multiphase level set method. Segmentation results on synthetic and real images validate the good performance of our model, and demonstrate the model's applicability to images with multiple channels and multiple objects.
Yan Nei Law, Hwee Kuan Lee, Chaoqiang Liu, Andy M. Yip
IEEE Trans. Image Process.3
2010 Robust video denoising using low rank matrix completion
abstract
Most existing video denoising algorithms assume a single statistical model of image noise, e.g. additive Gaussian white noise, which often is violated in practice. In this paper, we present a new patch-based video denoising algorithm capable of removing serious mixed noise from the video data. By grouping similar patches in both spatial and temporal domain, we formulate the problem of removing mixed noise as a low-rank matrix completion problem, which leads to a denoising scheme without strong assumptions on the statistical properties of noise. The resulting nuclear norm related minimization problem can be efficiently solved by many recently developed methods. The robustness and effectiveness of our proposed denoising algorithm on removing mixed noise, e.g. heavy Gaussian noise mixed with impulsive noise, is validated in the experiments and our proposed approach compares favorably against some existing video denoising algorithms.
Hui Ji 0002, Chaoqiang Liu, Zuowei Shen, Yuhong Xu
CVPR2
2009 Blind motion deblurring from a single image using sparse approximation
abstract
Restoring a clear image from a single motion-blurred image due to camera shake has long been a challenging problem in digital imaging. Existing blind deblurring techniques either only remove simple motion blurring, or need user interactions to work on more complex cases. In this paper, we present an approach to remove motion blurring from a single image by formulating the blind blurring as a new joint optimization problem, which simultaneously maximizes the sparsity of the blur kernel and the sparsity of the clear image under certain suitable redundant tight frame systems (curvelet system for kernels and framelet system for images). Without requiring any prior information of the blur kernel as the input, our proposed approach is able to recover high-quality images from given blurred images. Furthermore, the new sparsity constraints under tight frame systems enable the application of a fast algorithm called linearized Bregman iteration to efficiently solve the proposed minimization problem. The experiments on both simulated images and real images showed that our algorithm can effectively removing complex motion blurring from nature images.
Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen
CVPR3
2009 High-quality curvelet-based motion deblurring from an image pair
abstract
One promising approach to remove motion deblurring is to recover one clear image using an image pair. Existing dual-image methods require an accurate image alignment between the image pair, which could be very challenging even with the help of user interactions. Based on the observation that typical motion-blur kernels will have an extremely sparse representation in the redundant curvelet system, we propose a new minimization model to recover a clear image from the blurred image pair by enhancing the sparsity of blur kernels in the curvelet system. The sparsity prior on the motion-blur kernels improves the robustness of our algorithm to image alignment errors and image formation noise. Also, a numerical method is presented to efficiently solve the resulted minimization problem. The experiments showed that our proposed algorithm is capable of accurately estimating the blur kernels of complex camera motions with low requirement on the accuracy of image alignment, which in turn led to a high-quality recovered image from the blurred image pair.
Jian-Feng Cai 0001, Hui Ji 0002, Chaoqiang Liu, Zuowei Shen
CVPR3
2008 Motion blur identification from image gradients
abstract
Restoration of a degraded image from motion blurring is highly dependent on the estimation of the blurring kernel. Most of the existing motion deblurring techniques model the blurring kernel with a shift-invariant box filter, which holds true only if the motion among images is of uniform velocity. In this paper, we present a spectral analysis of image gradients, which leads to a better configuration for identifying the blurring kernel of more general motion types (uniform velocity motion, accelerated motion and vibration). Furthermore, we introduce a hybrid Fourier-Radon transform to estimate the parameters of the blurring kernel with improved robustness to noise over available techniques. The experiments on both simulated images and real images show that our algorithm is capable of accurately identifying the blurring kernel for a wider range of motion types.
Hui Ji 0002, Chaoqiang Liu
CVPR2