Chenyang Guan

dblp:323/2825 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0008-6422-7729ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SpecQuant: Spectral Decomposition and Adaptive Truncation for Ultra-Low-Bit LLMs Quantization
abstract
The emergence of accurate open large language models (LLMs) has sparked a push for advanced quantization techniques to enable efficient deployment on end-user devices. In this paper, we revisit the challenge of extreme LLM compression---targeting ultra-low-bit quantization for both activations and weights---from a Fourier frequency domain perspective. We propose SpecQuant, a two-stage framework that tackles activation outliers and cross-channel variance. In the first stage, activation outliers are smoothed and transferred into the weight matrix to simplify downstream quantization. In the second stage, we apply channel-wise low-frequency Fourier truncation to suppress high-frequency components while preserving essential signal energy, improving quantization robustness. Our method builds on the principle that most of the weight energy is concentrated in low-frequency components, which can be retained with minimal impact on model accuracy. To enable runtime adaptability, we introduce a lightweight truncation module during inference that adjusts truncation thresholds based on channel characteristics. On LLaMA-3 8B, SpecQuant achieves 4-bit quantization for both weights and activations, narrowing the zero-shot accuracy gap to only 1.5% compared to full precision, while delivering 2× faster inference and 3× lower memory usage.
Zhixiong Zhao, Fangxin Liu, Chenyang Guan, Zongwu Wang, Li Jiang 0002, Haibing Guan
AAAI4
2026 When Low-Rank Meets Mixed-Precision: Training-Free Joint Compression for Efficient LLM Inference
abstract
The rapid growth of Large Language Models (LLMs) raises significant challenges for deployment in resource-constrained environments. Existing compression approaches, such as low-rank decomposition and quantization, are typically applied independently, which limits their effectiveness and fails to exploit their complementarity. To address this issue, we present a training-free framework for joint compression that integrates low-rank decomposition with mixed-precision quantization. We formulate the allocation of layer-wise rank and bit-width as a combinatorial optimization problem, guided by an input-aware sensitivity metric to allocate resources where they yield the highest accuracy retention. We further develop a sample-aware low-rank decomposition scheme with theoretical guarantees, and introduce a unified difference matrix to mitigate the coupled errors from structural approximation and quantization. Extensive experiments on diverse LLM architectures and datasets demonstrate that our method achieves state-of-the-art compression, reducing model size to 20% of the original while preserving inference accuracy. The code is available at https://github.com/zzzzzjq0126/HALO.git
Fangxin Liu, Jinqi Zhu, Chenyang Guan, Tao Yang 0031, Li Jiang 0002, Haibing Guan
ASP-DAC4
2026 TFLOP: Towards Energy-Efficient LLM Inference An FPGA-Affinity Accelerator with Unified LUT-based OPtimization
abstract
Large Language Models (LLMs) suffer from significant performance and energy efficiency bottlenecks during the memory-bound decoding stage, where GPUs are often underutilized. We propose TFLOP, a novel CPU-FPGA heterogeneous prototype system that addresses this challenge by employing a 4-bit product quantization scheme on model weights and the KV cache. This approach decomposes GEMV operations in decoding stage into two hardware-friendly steps: centroid reconstruction and table lookup, which are efficiently mapped onto an FPGA’s heterogeneous resources. Our key innovation is a unified FPGA architecture that can handle both row- and column-wise quantization, simplifying hardware design and improving efficiency. Evaluations show that TFLOP achieves superior performance, delivering a 2.76$\times$ speedup over the NVIDIA A100 GPU on the LLaMA-2-7B model, while maintaining high accuracy and exceptional energy efficiency.
Zongwu Wang, Zhongyi Tang, Fangxin Liu, Chenyang Guan, Li Jiang 0002, Haibing Guan
ASP-DAC4
2026 EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result Reuse
abstract
Mixture-of-Experts (MoE) models significantly reduce computation in large language models by activating only a subset of experts per input token, but they introduce severe memory bottlenecks due to the large number of expert parameters. Existing offloading and prefetching strategies either incur accuracy loss, prohibitively high memory traffic, or high decoding overhead, limiting deployment on resource-constrained hardware. In this work, we present EARTH, a hardware–software co-design that addresses these challenges through three key innovations. First, we propose a dual-entropy encoding scheme that decomposes each expert into a high-information base and a delta component, enabling compact storage while preserving accuracy via adaptive precision management. Second, we introduce a delta-aware speculative prefetching and reuse mechanism that preloads base components of predicted experts and selectively fetches deltas, reusing previously computed delta patterns to reduce memory traffic and redundant computation. Third, we design a hardware accelerator that is co-designed to efficiently support this encoding and prefetching strategy, optimizing execution order, parallelism, and memory utilization. Across representative MoE workloads, EARTH reduces data movement overhead, improves prefetch efficiency, and achieves up to 2.10× speedup compared to state-of-the-art baselines, while maintaining high model accuracy.
Fangxin Liu, Ning Yang 0012, Jingkui Yang, Zongwu Wang, Chenyang Guan, Yu Feng 0007, Li Jiang 0002, Haibing Guan
ASPLOS (2)5
2026 LaMoS: Enabling Efficient Large Number Modular Multiplication through SRAM-based CiM Acceleration
abstract
Barrett’s algorithm is one of the most widely used methods for performing modular multiplication, a critical nonlinear operation in modern privacy computing techniques such as homomorphic encryption (HE) and zero-knowledge proofs (ZKP). Since modular multiplication dominates the processing time in these applications, computational complexity and memory limitations significantly impact performance. Computing-in-Memory (CiM) is a promising approach to tackle this problem. However, existing schemes currently suffer from two main problems: 1) Most works focus on low bit-width modular multiplication, which is inadequate for mainstream cryptographic algorithms such as elliptic curve cryptography (ECC) and the RSA algorithm, both of which require high bit-width operations; 2) Recent efforts targeting large number modular multiplication rely on inefficient in-memory logic operations, resulting in high scaling costs for larger bit-widths and increased latency. To address these issues, we propose LaMoS, an efficient SRAM-based CiM design for large-number modular multiplication, offering high scalability and area efficiency. First, we analyze the Barrett’s modular multiplication method and map the workload onto SRAM CiM macros for high bit-width cases. Additionally, we develop an efficient CiM architecture and dataflow to optimize large-number modular multiplication. Finally, we refine the mapping scheme for better scalability in high bit-width scenarios using workload grouping. Experimental results show that LaMoS achieves a 7.02 × speedup and reduces high bit-width scaling costs compared to existing SRAM-based CiM designs.
Haomin Li 0002, Fangxin Liu, Chenyang Guan, Zongwu Wang, Li Jiang 0002, Haibing Guan
DATE3
2026 STEP: Adaptive Spatio-Temporal Expert Prefetching for Low-Latency and Memory-Efficient MoE Inference
Fangxin Liu, Ning Yang 0012, Zongwu Wang, Chenyang Guan, Haomin Li 0002, Yu Feng 0007, Liqiang Lu, Siran Yang, Jiamang Wang, Lin Qu, Li Jiang 0002, Haibing Guan
ISCA4
2026 Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix Multiplication
Jingkui Yang, Fangxin Liu, Ning Yang 0012, Chenyang Guan, Zongwu Wang, Mei Wen, Li Jiang 0002, Haibing Guan
ISCA5
2024 FIAless: Asynchronous Programming for Large-Scale Burst Requests in Serverless Computing
abstract
As emerging microservice hosting platforms, serverless computing platforms are highly favour for their simplicity and high level of automation when deploying and running stateless functions. However, when any number of users sends many simultaneous requests to one or more functions, the frequent synchronization between stateless functions and remote storage can lead to high latency and additional system overhead during the request execution process on serverless platforms. In this paper, we design a concurrent scheduling policy called "Function in Asynchronous Serverless" (FIAless), which is based on asynchronous input/output (I/O), to handle large-scale burst requests on serverless platforms. FIAless breaks down the request execution process into subtasks with different execution characteristics and optimizes the scheduling and resource allocation schemes of each subtask to reduce the end-to-end latency and improve the overall performance of the system. In addition, FIAless monitors the entire lifecycle of each running request, ensuring that suspended requests can be promptly resumed. Finally, FIAless also periodically processes the residual data contained in memory after the operation execution procedure to promptly release resources. We evaluate FIAless by running multiple serverless computing benchmarks on Knative. Compared with running on the original Knative platform, the average request execution time is reduced by 61%, the memory consumption level is decreased by 80%, and high CPU utilization and low latency rates are maintained.
Chenyang Guan, Junjie Yin
HPCC1