Xilong Xie

dblp:340/7826 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0005-9988-2940ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Accelerating LLM Inference via Low-Bit Fine-Grained Quantization Algorithm and Bit-Level Accelerator Co-Design
abstract
Large language models (LLMs) have emerged as one of the most impactful and transformative paradigms in natural language processing. Despite their remarkable success, the intensive computational demands and substantial memory footprint impose a significant barrier to efficient LLM inference.In this paper, we present a comprehensive solution to improve LLM inference performance under ultra-low weight precision, meticulously optimized through algorithm and architecture co-design. To achieve this, we first propose a fine-grained intra-cluster bit allocation method that partitions the weights into small clusters and explicitly considers the distribution of outliers and salient points within each cluster. Then, an intra-cluster protection mechanism is proposed to selectively preserve important weights during quantization, where an extended integer format and group-wise scale factor search are further introduced to mitigate accuracy degradation caused by aggressive bit-width reduction. Furthermore, we develop a memory-aligned encoding scheme to facilitate efficient memory access while enabling flexible identification of mixed-precision representations. Finally, we design a lightweight bit-level accelerator for low-bit LLM inference, offering simplified hardware design and enhanced adaptability through parallel bit-level computation. Compared to existing state-of-the-art quantization algorithms, our algorithm achieves higher model accuracy under ultra-low weight precision. Meanwhile, the proposed bit-level accelerator delivers speedups of 1.59×, 1.38×, and 1.61×, along with energy efficiency improvements of 1.52×, 1.42×, and 1.22× over ANT, OliVe, and FineQ, respectively.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Tairan Zhang, Jinquan Wang, Yongyue Wang, Xiaojian Liao
IEEE Trans. Computers1
2025 FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs
abstract
Large language models (LLMs) have significantly advanced the natural language processing paradigm but impose substantial demands on memory and computational resources. Quantization is one of the most effective ways to reduce memory consumption of LLMs. However, advanced single-precision quantization methods experience significant accuracy degradation when quantizing to ultra-low bits. Existing mixed-precision quantization methods are quantized by groups with coarse granularity. Employing high precision for group data leads to substantial memory overhead, whereas low precision severely impacts model accuracy. To address this issue, we propose FineQ, software-hardware co-design for low-bit fine-grained mixed-precision quantization of LLMs. First, FineQ partitions the weights into finer-grained clusters and considers the distribution of outliers within these clusters, thus achieving a balance between model accuracy and memory overhead. Then, we propose an outlier protection mechanism within clusters that uses 3 bits to represent outliers and introduce an encoding scheme for index and data concatenation to enable aligned memory access. Finally, we introduce an accelerator utilizing temporal coding that effectively supports the quantization algorithm while simplifying the multipliers in the systolic array. FineQ achieves higher model accuracy compared to the SOTA mixed-precision quantization algorithm at a close average bit-width. Meanwhile, the accelerator achieves up to 1.79x energy efficiency and reduces the area of the systolic array by 61.2%.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Xiangrong Xu 0002
DATE1
2025 PointISA: ISA-Extensions for Efficient Point Cloud Analytics via Architecture and Algorithm Co-Design
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu
MICRO6
2025 Amove: Accelerating LLMs through Mitigating Outliers and Salient Points via Fine-Grained Grouped Vectorized Data Type
abstract
The quantization of Large Language Models (LLMs) poses significant challenges due to the heterogeneous nature of feature point distributions in low-bit quantization scenarios, including salient points, normal outliers, and massive outliers.These challenges are particularly pronounced in supporting both weight-only and weight-activation quantization modes, as existing methods often focus on a single mode and fail to address the diverse feature characteristics holistically, resulting in suboptimal model accuracy and hardware efficiency trade-offs.To tackle these limitations, we introduce Amove, a novel codesign framework that synergistically integrates data type and hardware architecture design for efficient LLM quantization.Our approach is threefold: First, we conduct a comprehensive analysis of quantization granularity and propose a residual approximation mechanism that balances model accuracy and memory overhead under fine-grained quantization.Second, we design a flexible finegrained grouped vectorized data type, enabling seamless support for both weight-activation and low-bit weight-only quantization modes within a unified framework.Third, we implement the hardware architecture of Amove on both GPU tensor core and systolic arraybased architectures.The Amove-enhanced tensor core achieves an average speedup of 2.13× and a 1.70× reduction in energy consumption over the state-of-the-art OliVe design.Furthermore, an Amove-based accelerator achieves up to 2.67× speedup and 1.68× energy reduction over the state-of-the-art accelerator.
Xilong Xie, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Xiangrong Xu 0002, Jinquan Wang, Xiaojian Liao
MICRO1
2025 Exploiting intra-chip locality for multi-chip GPUs via two-level shared L1 cache
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Xiaojian Liao
J. Syst. Archit.8
2024 FuseFPS: Accelerating Farthest Point Sampling with Fusing KD-tree Construction for Point Clouds
abstract
Point cloud analytics has become a critical workload for embedded and mobile platforms across various applications. Farthest point sampling (FPS) is a fundamental and widely used kernel in point cloud processing. However, the heavy external memory access makes FPS a performance bottleneck for real-time point cloud processing. Although bucket-based farthest point sampling can significantly reduce unnecessary memory accesses during the point sampling stage, the KD-tree construction stage becomes the predominant contributor to execution time. In this paper, we present FuseFPS, an architecture and algorithm co-design for bucket-based farthest point sampling. We first propose a hardware-friendly sampling-driven KD-tree construction algorithm. The algorithm fuses the KD-tree construction stage into the point sampling stage, further reducing memory accesses. Then, we design an efficient accelerator for bucket-based point sampling. The accelerator can offload the entire bucket-based FPS kernel at a low hardware cost. Finally, we evaluate our approach on various point cloud datasets. The detailed experiments show that compared to the state-of-the-art accelerator QuickFPS, FuseFPS achieves about $4.3\times $ and about $6.1\times $ improvements on speed and power efficiency, respectively.
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie
ASPDAC6
2024 ATA-Cache: Contention Mitigation for GPU Shared L1 Cache With Aggregated Tag Array
abstract
To fully exploit the locality of GPU applications, the GPU shared L1 cache architecture, which shares L1 cache among multiple GPU cores, is a promising architecture while still suffering from high resource contentions. We present a GPU shared L1 cache architecture with an aggregated tag array that minimizes the L1 cache contentions and takes full advantage of inter-core locality. The key idea is to decouple and aggregate the tag arrays of multiple L1 caches so that the cache requests can be compared with all tag arrays in parallel to probe the replicated data in other caches. The GPU caches are only accessed by other GPU cores when replicated data exists, filtering out unnecessary cache accesses that cause high resource contentions. We also develop a two-level thread-block scheduling policy adapted for the shared L1 cache architecture to maximize the available locality. The experimental results show that GPU performance can be improved by 14.5% on average for applications with a high inter-core locality.
Xiangrong Xu 0002, Liang Wang 0020, Limin Xiao 0001, Lei Liu 0037, Yuanqiu Lv, Xilong Xie, Hao Liu 0107
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6