Xueqi Li 0001

dblp:163/2079-1 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0001-6396-0058ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CoCoTree: A Computation-Capable Architecture for Collective Communication in Scalable PIM
abstract
The growing demand for high-bandwidth and largecapacity memory access in data-intensive workloads has driven the development and deployment of Processing-in-Memory (PIM) architectures. However, existing DIMM-based PIM systems suffer from the severe communication bottleneck between the processing elements (PEs) near the PIM banks due to their requirement on host CPU forwarding. This bottleneck limits the efficiency of collective operations and degrades scalability and performance for workloads that require inter-PE communication. To address the communication limitation, we propose CoCoTree, a computation-capable architecture for collective communication in scalable DIMM-based PIM. CoCoTree supports direct and high-throughput inter-PE communication without host intervention. CoCoTree accelerates key collective communication using novel hierarchical binary tree topology and lightweight in-network computation support. We design and implement microarchitectures for the main building blocks: Co-Leaf and Co-Node, to efficiently handle the data packing, routing, and processing in CoCoTree. Furthermore, we also introduce a packet-based communication protocol tailored to the CoCoTree architecture, which decouples control and data through a twophase configuration-computation communication mechanism to efficiently support a wide range of collective communication operations. CoCoTree effectively mitigates inter-PE communication bottlenecks, enabling scalable PIM systems capable of meeting the demands of growing data size. Experimental results show that CoCoTree achieves up to$95.6 \times$improvement for collective operations and improves end-to-end application performance by up to$10.5 \times$across various workloads over the baseline PIM, while outperforming state-of-the-art PIM communication architectures in both performance and scalability.
Shunchen Shi, Qijia Yang, Fan Yang 0096, Yu Huang 0013, Youwei Zhuo, Zhichun Li, Ninghui Sun, Xueqi Li 0001
HPCA8
2026 TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification
abstract
Verification is a critical process for ensuring the correctness of modern processors. The increasing complexity of processor designs and the emergence of new instruction set architectures (ISAs) like RISC-V have created demands for more agile and efficient verification methodologies, particularly regarding verification efficiency and faster coverage convergence. While simulation-based approaches now attempt to incorporate advanced software testing techniques such as fuzzing to improve coverage, they face significant limitations when applied to processor verification, notably poor performance and inadequate test case quality. Hardware-accelerated solutions using FPGA or ASIC platforms have tried to address these issues, yet they struggle with challenges including host-FPGA communication overhead, inefficient test pattern generation, and suboptimal implementation of the entire multi-step verification process. In this paper, we present TurboFuzz, an end-to-end hardwareaccelerated verification framework that implements the entire Test Generation-Simulation-Coverage Feedback loop on a single FPGA for modern processor verification. TurboFuzz enhances test quality through optimized test case (seed) control flow, efficient inter-seed scheduling, and hybrid fuzzer integration, thereby improving coverage and execution efficiency. Additionally, it employs a feedback-driven generation mechanism to accelerate coverage convergence. Experimental results show that TurboFuzz achieves up to 2.23× more coverage collection than software-based fuzzers within the same time budget, and up to 571× performance speedup when detecting real-world issues, while maintaining full visibility and debugging capabilities with moderate area overhead.
Xueqi Li 0001, Sa Wang, David Boland, Yungang Bao, Kan Shi
HPCA3
2026 Meridian: In-Memory Acceleration for RAG with Document Attention Decomposition
Chaoqiang Liu, Yu Huang 0013, Haifeng Liu 0003, Yi Zhang 0191, Qihang Qiu, Xueqi Li 0001, Xiaofei Liao, Hai Jin, Jingling Xue
ISCA6
2026 HOPESim: A Lightweight and Modern C++ based Accelerator Simulation Approach
Xueqi Li 0001, Ruihao Gao, Xiaoyu Zhang 0009, Xiaoming Chen 0003, Shunchen Shi, Fan Yang 0096, Ninghui Sun
ISCAS1
2026 GPA: A General-Purpose In-Memory Computing Accelerator
Xiaoyu Zhang 0009, Zerun Li, Rui Liu 0045, Libo Shen, Boyu Long, Xueqi Li 0001, Yinhe Han 0001, Xiaoming Chen 0003
ISCAS6
2025 BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM Inference
abstract
Processing-In-Memory (PIM) architectures alleviate the memory bottleneck in the decode phase of large language model (LLM) inference by performing operations like GEMV and Softmax in memory. However, the fragmented data layout in current PIM architectures limits end-to-end acceleration for long-context LLMs. In this paper, we propose BlockPIM, a cross-channel block memory layout strategy that maximizes memory utilization and eliminates the context length constraint. Additionally, we introduce a cross-channel attention computation scheme that is compatible with the current architecture to support distributed attention operations on BlockPIM. Experimental results demonstrate that our approach achieves a 62% average throughput increase compared to existing state-of-the-art PIM solutions, enabling efficient and scalable deployment of large language models on PIM architectures.
Zhichun Li, Xueqi Li 0001, Ninghui Sun
DAC3
2025 upTSA: A DIMM-Based Near Data Processing Accelerator for Time Series Analysis
Shunchen Shi, Fan Yang 0096, Qijia Yang, Xiaohui Peng 0002, Xueqi Li 0001, Ninghui Sun
NPC (1)5
2025 EdgeInferFlow: A Distributed Inference Acceleration Method for Deep Learning Chained Structure Models for Edge Devices
Hanfeng Zhai, Yifan Wang 0005, Xiaohui Peng 0002, Xueqi Li 0001
NPC (1)5
2025 FuHsi: Shifting Base-Calling Closer to Sequencer via In-Cache Acceleration
Yewen Li, Guangming Tan, Xueqi Li 0001
J. Comput. Sci. Technol.3
2024 BeeZip: Towards An Organized and Scalable Architecture for Data Compression
abstract
Data compression plays a critical role in operating systems and large-scale computing workloads. Its primary objective is to reduce network bandwidth consumption and memory/storage capacity utilization. Given the need to manipulate hash tables, and execute matching operations on extensive data volumes, data compression software has transformed into a resource-intensive CPU task. To tackle this challenge, numerous prior studies have introduced hardware acceleration methods. For example, they have utilized Content-Addressable Memory (CAM) for string matches, incorporated redundant historical copies for each matching component, and so on. While these methods amplify the compression throughput, they often compromise an essential aspect of compression performance: the compression ratio (C.R.). Moreover, hardware accelerators face significant resource costs, especially in memory, when dealing with new large sliding window algorithms.
Ruihao Gao, Zhichun Li, Guangming Tan, Xueqi Li 0001
ASPLOS (3)4
2024 ProXplore: A GPU-Enhanced Protein Discovery Engine
abstract
Language models have achieved unprecedented success in natural language processing tasks and have recently been adapted for biological sequences. However, GPUs still encounter significant performance bottlenecks when running BERT-style natural language models. Moreover, due to the vastly greater number of tokens in protein sequences compared to human languages (as illustrated in Figure 1), directly transferring language models from human languages to proteins results in significant changes in runtime behavior. Consequently, optimizations designed for short-input BERT models are less effective for protein language models.In this paper, we propose a novel GPU-based cross-layer optimization strategy. From an architectural perspective, our approach leverages intra- and inter-operator parallelism, multilevel data computation and warp parallelism to fully utilize GPU features. From an algorithmic perspective, we address the issue of varying lengths in protein input sequences and reduce memory overhead through padding removal. Experimental results demonstrate that ProXplore achieves a 1.8× increase in inference speed on the TAPE benchmark without a significant loss in prediction accuracy. This improvement effectively overcomes performance bottlenecks and provides substantial benefits for protein sequence analysis and other bioinformatics applications.
Zewen Sun, Peiheng Zhang, Xingzhuo Fu, Xueqi Li 0001
BIBM4
2024 LazyCAT: Efficient Fine-Grained Cache Partitioning with Two Boundaries
abstract
Intel CAT is a widely available cache partitioning technique in commercial hardware but falls short in partitioning granularity. We propose LazyCAT, a fine-grained, on-demand, and easy-to-use cache partitioning technique, which not only can ensure the QoS of High-Priority (HP) applications but also yield the under-utilized cache sets to other Best-Effort (BE) applications for better resource efficiency. LazyCAT retains the easy-to-use philosophy of CAT and introduces a new soft LLC partitioning boundary, which is lower than the original CAT partitioning boundary (hard boundary). LazyCAT detects and selects under-utilized cache sets in HP applications during profiling, and specifies them to the soft boundary at runtime dynamically, yielding the cache blocks between these two boundaries for other applications. Meanwhile, LazyCAT provides users with a simple software interface to guide the set-level space allocation according to their needs. Experimental results show that LazyCAT exhibits substantial performance improvements (up to 12.2%) for BE applications with less than 3% performance degradation of HP applications.
Chuanqi Zhang, Xueqi Li 0001, Ninghui Sun, Yungang Bao, Sa Wang
HPCC2
2024 CoPIM: A Collaborative Scheduling Framework for Commodity Processing-in-memory Systems
abstract
Processing in memory (PIM) is a promising paradigm to effectively alleviate the bottleneck of memory access for data-intensive applications. UPMEM is the first publicly-available real-world processing-in-memory (PIM) platform. However, current approaches mostly treat the CPU as the controller rather than considering the CPU-PIM system as a whole, thus failing to fully utilize the computational resources in modern CPUs. Additionally, the use of commodity PIM hardware as accelerators with independent memory addresses leads to significant data communication between the CPU and PIM, diminishing the performance benefits of reducing data movement. To address these challenges, we propose CoPIM, a scheduling framework that automates efficient workload scheduling. CoPIM introduces a fine-grained load modeling approach based on computational graphs and establishes a performance model to estimate workload performance and energy consumption. Our framework includes a scheduler that generates optimal scheduling parameters, enabling CPU-PIM collaborative processing. CoPIM enhances system parallelism and reduces data transfer between the CPU and PIM. We implement CoPIM on the UPMEM PIM system and demonstrate through experiments that our design achieves 4.0 x performance improvement on average when compared to general-purpose processors while reducing energy consumption by 56.6%.
Shunchen Shi, Xueqi Li 0001, Zhaowu Pan, Peiheng Zhang, Ninghui Sun
ICCD2
2023 TransCaller: An End-to-end Accelerated Transformer-based Nanopore Basecaller on GPUs
abstract
Basecalling is a crucial step in nanopore sequencing as it transforms the raw electrical signal obtained from the nanopore into a readable sequence. The accuracy of basecalling directly affects the quality and reliability of the sequenced data. Various algorithms and models are employed to perform accurate basecalling, with advancements in deep learning techniques. However, due to the difficulty of combining feature extraction capability, high parallelism and long sequence modeling capability of existing models, the accuracy and speed of basecalling model are still bottlenecks in data analysis.In this paper, we present TransCaller, an end-to-end accelerated transformer-based nanopore basecaller model. TransCaller comprises a low-parameter electrical signal sampler, a fully transformer-based encoder integrating multiple algorithm-specific modules, and a CTC decoder enhanced with a filter. To further refine and enhance the TransCaller model, we introduce a three-stage pyramid model structure and leverage knowledge distillation techniques to balance the accuracy and speed. Experimental results on 9 datasets demonstrate that TransCaller achieves an average accuracy of 92.95%, surpassing the state-of-the-art Bonito sup model by a margin of 0.65% to 2.34%. Moreover, the distilled TransCallermodel showcases a remarkable 2.97× acceleration in inference speed compared to the original model, effectively catering to the variegated requirements of distinct inference speeds and precision levels across a wide spectrum of scenarios.
Peiheng Zhang, Shunchen Shi, Xueqi Li 0001
BIBM5
2023 NvWa: Enhancing Sequence Alignment Accelerator Throughput via Hardware Scheduling
abstract
Sequence alignment is the most time-consuming step in the genome analysis pipeline. Since sequence alignment generally follows the seed-and-extension paradigm, prior proposed hardware accelerators either opt to accelerate the seeding phase or the seed-extension phase. However, the diversity of each sequence in the alignment workflow leads to the pipeline stall or bubbles, which finally results in a decreased throughput for the end-to-end sequence alignment.In this paper, we propose NvWa, which is a hardware scheduling accelerator for sequence alignment. To solve the diversity problem, we propose three novel scheduling mechanisms and corresponding architecture, which target the seeding phase, the seed-extension phase, and the interaction between the two phases, respectively. For the seeding phase, we propose a Seeding Scheduler to schedule all idle seeding units in only one cycle. For the seed-extension phase, the Extension Scheduler can achieve both lower latency and higher parallelism when facing seed-extension tasks with different scales. Between the two phases, an efficient Coordinator caches and dispatches seeding hits to optimal and sub-optimal seed-extension units. Furthermore, to avoid algorithmic obsolescence for the new sequence technologies, we propose a loosely coupled design, which decouples the data path and the control scheduling path. Experimental results show that NvWa can achieve 493×, 200×, 12.11×, 2.30× speedup and 14.21×, 5.60×, 4.34×, 5.85× energy reduction when compared with a 16-thread CPU baseline, an NVIDIA A100 GPU, and two state-of-the-art accelerators, respectively.
Yewen Li, Xueqi Li 0001, Ruihao Gao, Wanqi Liu, Guangming Tan
HPCA2
2023 Accelerating k-Shape Time Series Clustering Algorithm Using GPU
abstract
In the data space, time-series analysis has emerged in many fields, including biology, healthcare, and numerous large-scale scientific facilities like astronomy, climate science, particle physics, and genomics. Clustering is one of the most critical methods in time-series analysis. So far, the state-of-art time series clustering algorithm k-Shape has been widely used not only because of its high accuracy, but also because of its relatively low computation cost. However, due to the high heterogeneity of time series data, it can not be simply regarded as a high-dimensional vector. Two time series often need some alignment method in similarity comparison. The alignment between sequences is often a time-consuming process. For example, when using dynamic time warping as a sequence alignment algorithm and if the length of time series is greater than 1,000, a single iteration in the clustering process may take hundreds to tens of thousands of seconds, while the entire clustering cycle often requires dozens of iterations. In this article, we propose a set of novel parallel strategies suitable for GPU's computation model, called Times-C, which is an abbreviation for Time Series Clustering. We define three stages in the analysis process: aggregation, centroid, and class assignment. Times-C includes efficient parallel algorithms and corresponding implementations for these three stages. Overall, the experimental results show that the Times-C algorithm exhibits a performance improvement of one to two orders of magnitude compared to the multi-core CPU version of k-Shape. Furthermore, compared to the GPU version of the k-Shape algorithm, the Times-C algorithm achieves a maximum acceleration of up to 345 times.
Xun Wang 0010, Ruibao Song, Junmin Xiao, Xueqi Li 0001
IEEE Trans. Parallel Distributed Syst.5
2022 Crescent: A GPU-based Targeted Nanopore Sequence Selector
abstract
With the development of three-generation genome sequencing technology, targeted sequencing has become a fundamental need in virus detection, metagenomics expansion, and human polymorphisms detection. However, previous studies on CPU/GPU platforms have much lower throughput than Nanopore sequencers, and hardware designs have high design cost and flexibility issues. In this paper, we propose Crescent, a GPU-based targeted sequence selector. Specifically, on the algorithm level, we modify the popular-used algorithm without significantly decreasing the accuracy and then propose four key observations. On the architecture level, based on four key observations, we propose two-stage parallelism mechanisms (i.e., inter-task and intra-task) and other optimization strategies (i.e., data reuse and redundant computation elimination) to accelerate our modified algorithm. Our experimental results show that the throughput of Crescent on NVIDIA A100 is $10.84 \times$ and $2.30 \times$ higher than a state-of-the-art 16-thread and 72-thread CPU baseline, respectively.
Xueqi Li 0001, Yewen Li, Ruibao Song, Xun Wang 0010
BIBM2
2022 MetaZip: a high-throughput and efficient accelerator for DEFLATE
abstract
Booming data volume has become an important challenge for data center storage and bandwidth resources. Consequently, fast and efficient compression architecture is becoming the most fundamental design in data centers. However, the compression ratio (CR) and compression throughput are often difficult to achieve at the same time on existing computing platforms. DEFLATE is a widely used compression format in data centers, which is an ideal case for hardware acceleration. Unfortunately, Deflate has an inherent connection among its special memory access pattern, which limits a higher throughput.
Ruihao Gao, Xueqi Li 0001, Yewen Li, Xun Wang 0010, Guangming Tan
DAC2
2022 AlphaSparse: Generating High Performance SpMV Codes Directly from Sparse Matrices
abstract
Sparse Matrix-Vector multiplication (SpMV) is an essential computational kernel in many application scenarios. Tens of sparse matrix formats and implementations have been proposed to compress the memory storage and speed up SpMV performance. We develop AlphaSparse, a superset of all existing works that goes beyond the scope of human-designed format(s) and implementation(s). AlphaSparse automatically creates novel machine-designed formats and SpMV kernel implementations en-tirely from the knowledge of input sparsity patterns and hard-ware architectures. Based on our proposed Operator Graph that expresses the path of SpMV format and kernel design, AlphaS-parse consists of three main components: Designer, Format & Kernel Generator, and Search Engine. It takes an arbitrary sparse matrix as input while outputs the performance machine-designed format and SpMV implementation. By extensively evaluating 843 matrices from SuiteSparse Matrix Collection, AlphaSparse achieves significant performance improvement by 3.2 × on average compared to five state-of-the-art artificial formats and 1.5 × on average (up to 2.7×) over the up-to-date implementation of traditional auto-tuning philosophy.
Zhen Du, Jiajia Li 0001, Yinshan Wang, Xueqi Li 0001, Guangming Tan, Ninghui Sun
SC4
2021 PIM-Align: A Processing-in-Memory Architecture for FM-Index Search Algorithm
Xueqi Li 0001, Guangming Tan, Ninghui Sun
J. Comput. Sci. Technol.1
2019 Efficient System Architecture in the Era of Monolithic 3D: Dynamic Inter-tier Interconnect and Processing-in-Memory
abstract
Emerging Monolithic Three-Dimensional (M3D) integration technology will not only provide improved circuit density through the high-bandwidth coupling of multiple vertically-stacked layers, but it can also provide new architectural opportunities for on-chip computation, memory, and communication that are beyond the capabilities of existing process and packaging technologies. For example, with massive parallel communication between heterogeneous memory and compute layers, existing processing-in-memory architectures can be optimized and expanded, developing into efficient and flexible near-data processors. Additionally, multiple tiers of interconnect can be dynamically leveraged to provide an efficient, scalable interconnect fabric that spans the three-dimensional system. This work explores some of the challenges and opportunities presented by M3D technology for emerging computer architectures, with focus on improving efficiency and increasing system flexibility.
Dylan C. Stow, Itir Akgun, Wenqin Huangfu, Yuan Xie 0001, Xueqi Li 0001, Gabriel H. Loh
DAC5
2019 MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm
abstract
Computational genomics has proven its great potential to support precise and customized health care. However, with the wide adoption of the Next Generation Sequencing (NGS) technology, 'DNA Alignment', as the crucial step in computational genomics, is becoming more and more challenging due to the booming bio-data. Consequently, various hardware approaches have been explored to accelerate DNA seeding - the core and most time consuming step in DNA alignment.
Wenqin Huangfu, Xueqi Li 0001, Shuangchen Li, Xing Hu 0001, Peng Gu 0008, Yuan Xie 0001
MICRO2
2018 Accelerating FM-index Search for Genomic Data Processing
abstract
The deluge of genomics data is incurring prohibitively high computational costs. As an important building block for genomic data processing algorithms, FM-index search occupies most of execution time in sequence alignment. Due to massive random streaming memory references relative to only small amount of computations, FM-index search algorithm exhibits extremely low efficiency on conventional architectures. This paper proposes Niubility, an accelerator for FM-index search in genomic sequence alignment. Based on our algorithm-architecture co-design analysis, we found that conventional architectures exploit low memory-level parallelism so that the available memory bandwidth cannot be fully utilized. Niubility accelerator customizes bit-wise operations and exploit data-level parallelism, that produces maximal concurrent memory accesses to saturate memory bandwidth. We implement an accelerator ASIC in a ST 28nm process that achieves up to 990x speedup over the state-of-the-art software.
Yuanrong Wang, Xueqi Li 0001, Dawei Zang, Guangming Tan, Ninghui Sun
ICPP2
2018 High-performance genomic analysis framework with in-memory computing
abstract
In this paper, we propose an in-memory computing framework (called GPF) that provides a set of genomic formats, APIs and a fast genomic engine for large-scale genomic data processing. Our GPF comprises two main components: (1) scalable genomic data formats and API. (2) an advanced execution engine that supports efficient compression of genomic data and eliminates redundancies in the execution engine of our GPF. We further present both system and algorithm-specific implementations for users to build genomic analysis pipeline without any acquaintance of Spark parallel programming. To test the performance of GPF, we built a WGS pipeline on top of our GPF as a test case. Our experimental data indicate that GPF completes Whole-Genome-Sequencing (WGS) analysis of 146.9G bases Human Platinum Genome in running time of 24 minutes, with over 50% parallel efficiency when used on 2048 CPU cores. Together, our GPF framework provides a fast and general engine for large-scale genomic data processing which supports in-memory computing.
Xueqi Li 0001, Guangming Tan, Bingchen Wang, Ninghui Sun
PPoPP1
2016 Accelerating large-scale genomic analysis with Spark
abstract
High-throughput next-generation sequencing technologies are producing a flood of cheap genomic information, providing precision medicine with the opportunity to better understand the primary cause of complicated diseases like cancer. However, even current state-of-the-art approaches still have large gaps with data generation due to limited scalability, accuracy and computational efficiency. To explore how to efficiently and effectively synthesize genomic data into knowledge, we propose GATK-Spark, a balanced parallelization approach that implements an in-memory version of GATK using Apache Spark. First, we performed a rigorous analysis of current GATK optimization strategies. We identify that compute resource utilization, text-based data format and long time single-thread file cutting and mergence operations are three major scalable bottlenecks. Second, we share our experiences designing a new approach optimized for GATK with big-data computing frameworks Apache Spark - GATK-Spark, which reduces the original execution of 20 hours to 30 minutes with a speedup in excess of 37 at 256 CPU cores. This work will facilitate the understanding of genomics analytics pipeline and design of strategies for accelerating large scale genomic analysis applications.
Xueqi Li 0001, Guangming Tan, Zhonghai Zhang, Ninghui Sun
BIBM1