Yuzhuang Xu

dblp:306/0476 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
abstract
Large language models (LLMs) utilize key-value (KV) cache to store historical information during sequence processing. The size of KV cache grows linearly as the length of the sequence extends, which seriously affects memory usage and decoding efficiency. Current methods for KV cache eviction typically utilize the last window from the pre-filling phase as queries to compute the KV importance scores for eviction. Although this scheme is simple to implement, it tends to overly focus on local information, potentially leading to the neglect or omission of crucial global information. To mitigate this issue, we propose **Judge Q**, a novel training method which incorporates a soft token list. This method only tunes the model’s embedding layer at a low training cost. By concatenating the soft token list at the end of the input sequence, we train these tokens' attention map to the original input sequence to align with that of the actual decoded tokens. In this way, the queries corresponding to the soft tokens can effectively capture global information and better evaluate the importance of the keys and values within the KV cache, thus maintaining decoding quality when KV cache is evicted. Under the same eviction budget, our method exhibits less performance degradation compared to existing eviction approaches. We validate our approach through experiments conducted on models such as Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3, using benchmarks including LongBench, RULER, and Needle-in-a-Haystack. Results indicate an improvement of approximately 1 point on the LongBench and over 3 points on RULER. This proposed methodology can be seamlessly integrated into existing open-source models with minimal training overhead, thereby enhancing performance in KV cache eviction scenarios.
Yuzhuang Xu, Shiyu Ji, Yang Xu 0049, Qingfu Zhu, Wanxiang Che
AAAI3
2026 CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
abstract
Large Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU.
Yuzhuang Xu, Xu Han 0007, Yuanchi Zhang, Shiyu Ji, Qingfu Zhu, Wanxiang Che
AAAI1
2025 ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
abstract
Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang, Peng Li, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziyue Wang 0002, Chi Chen 0005, Fuwen Luo, Yurui Dong 0001, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang 0014, Peng Li 0030, Yang Liu 0005
ACL (1)6
2025 Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
abstract
Large language models (LLMs) rely on keyvalue cache (KV cache) to accelerate decoding by reducing redundant computations.However, the KV cache memory usage grows substantially with longer text sequences, posing challenges for efficient deployment.Existing KV cache eviction methods prune tokens using prefilling-stage attention scores, causing inconsistency with actual inference queries, especially under tight memory budgets.In this paper, we propose Lookahead Q-Cache (LAQ), a novel eviction framework that generates lowcost pseudo lookahead queries to better approximate the true decoding-stage queries.By using these lookahead queries as the observation window for importance estimation, LAQ achieves more consistent and accurate KV cache eviction aligned with real inference scenarios.Experimental results on LongBench and Needlein-a-Haystack benchmarks show that LAQ outperforms existing methods across various budget levels, achieving a 1 ∼ 4 point improvement on LongBench under limited cache budget.Moreover, LAQ is complementary to existing approaches and can be flexibly combined to yield further improvements.
Shiyu Ji, Yuzhuang Xu, Yang Xu 0049, Qingfu Zhu, Wanxiang Che
EMNLP4
2025 CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
abstract
Abstract Powerful large language models (LLMs) are increasingly expected to be deployed with lower computational costs, enabling their capabilities on resource-constrained devices. Post-training quantization (PTQ) has emerged as a star approach to achieve this ambition, with best methods compressing weights to less than 2 bit on average. In this paper, we propose Channel-Relaxed Vector Quantization (CRVQ), a novel technique that significantly improves the performance of PTQ baselines at the cost of only minimal additional bits. This state-of-the-art extreme compression method achieves its results through two key innovations: (1) carefully selecting and reordering a very small subset of critical weight channels, and (2) leveraging extended codebooks to relax the constraint of critical channels. With our method, we demonstrate a 38.9% improvement over the current strongest sub-2-bit PTQ baseline, enabling nearer lossless 1-bit compression. Furthermore, our approach offers flexible customization of quantization bit-width and performance, providing a wider range of deployment options for diverse hardware platforms. Code and checkpoints are available at https://github.com/xuyuzhuang11/CRVQ.
Yuzhuang Xu, Shiyu Ji, Qingfu Zhu, Wanxiang Che
Trans. Assoc. Comput. Linguistics1
2024 UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
abstract
Haoyu Wang, Shuo Wang, Yukun Yan, Xujia Wang, Zhiyu Yang, Yuzhuang Xu, Zhenghao Liu, Liner Yang, Ning Ding, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shuo Wang 0013, Yukun Yan, Xujia Wang, Zhiyu Yang 0001, Yuzhuang Xu, Zhenghao Liu 0001, Liner Yang, Ning Ding 0002, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
ACL (1)6
2024 Pluggable Neural Machine Translation Models via Memory-augmented Adapters
abstract
Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratch for each user requirement, we propose a memory-augmented adapter to steer pretrained NMT models in a pluggable manner. Specifically, we construct a multi-granular memory based on the user-provided text samples and propose a new adapter architecture to combine the model representations and the retrieved results. We also propose a training strategy using memory dropout to reduce spurious dependencies between the NMT model and the memory. We validate our approach on both style- and domain-specific experiments and the results indicate that our method can outperform several representative pluggable baselines.
Yuzhuang Xu, Shuo Wang 0013, Peng Li 0030, Xuebo Liu 0002, Xiaolong Wang 0014, Yang Liu 0005
LREC/COLING1
2024 Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
abstract
Fine-tuning is a crucial process for adapting large language models (LLMs) to diverse applications. In certain scenarios, such as multi-tenant serving, deploying multiple LLMs becomes necessary to meet complex demands. Recent studies suggest decomposing a fine-tuned LLM into a base model and corresponding delta weights, which are then compressed using low-rank or low-bit approaches to reduce costs. In this work, we observe that existing low-rank and low-bit compression methods can significantly harm the model performance for task-specific fine-tuned LLMs (e.g., WizardMath for math problems). Motivated by the long-tail distribution of singular values in the delta weights, we propose a delta quantization approach using mixed-precision. This method employs higher-bit representation for singular vectors corresponding to larger singular values. We evaluate our approach on various fine-tuned LLMs, including math LLMs, code LLMs, chat LLMs, and even VLMs. Experimental results demonstrate that our approach performs comparably to full fine-tuned LLMs, surpassing both low-rank and low-bit baselines by a considerable margin. Additionally, we show that our method is compatible with various backbone LLMs, such as Llama-2, Llama-3, and Mistral, highlighting its generalizability.
Bowen Ping, Shuo Wang 0013, Hanqing Wang 0003, Xu Han 0007, Yuzhuang Xu, Yukun Yan, Yun Chen 0007, Baobao Chang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS5
2024 OneBit: Towards Extremely Low-bit Large Language Models
abstract
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance degradation when the bit-width is extremely reduced, and thus focus on utilizing 4-bit or 8-bit values to quantize models. This paper boldly quantizes the weight matrices of LLMs to 1-bit, paving the way for the extremely low bit-width deployment of LLMs. For this target, we introduce a 1-bit model compressing framework named OneBit, including a novel 1-bit parameter representation method to better quantize LLMs as well as an effective parameter initialization method based on matrix decomposition to improve the convergence speed of the quantization framework. Sufficient experimental results indicate that OneBit achieves good performance (at least 81% of the non-quantized performance on LLaMA models) with robust training processes when only using 1-bit weight matrices.
Yuzhuang Xu, Xu Han 0007, Zonghan Yang, Shuo Wang 0013, Qingfu Zhu, Zhiyuan Liu 0001, Wanxiang Che
NeurIPS1
2021 Cover Image
abstract
The cover image is based on the Original Article Verification Algebra for Multi-Tenant Applications in VaaS Architecture by Kan Luo et al., https://doi.org/10.1002/stvr.1763.
Kai Hu 0004, Ji Wan, Yuzhuang Xu, Zijing Cheng, Wei-Tek Tsai
Softw. Test. Verification Reliab.4
2021 Verification algebra for multi-tenant applications in VaaS architecture
abstract
Summary This paper proposes an algebraic system, verification algebra (VA), for reducing the number of component combinations to be verified in multi‐tenant architecture (MTA). MTA is a design architecture used in SaaS (Software‐as‐a‐Service) where a tenant can customize its applications by integrating services already stored in the SaaS databases or newly supplied services. Similar to SaaS, VaaS (Verification‐as‐a‐Service) is a verification service in a cloud that leverages the computing power offered by a cloud environment with automated provisioning, scalability and service composition. In VaaS architecture, however, there is a challenging problem called ‘combinatorial explosion’ that it is difficult to verify a large number of compositions constructed by both quantities of components and various combination structures even with computing resources in cloud. This paper proposes rules to emerge combinations status for future verification, on the basis of the existing results. Both composition patterns and properties are considered and analysed in VA rules.
Kai Hu 0004, Ji Wan, Yuzhuang Xu, Zijing Cheng, Wei-Tek Tsai
Softw. Test. Verification Reliab.4
2021 Erratum
abstract
The published online cover has been updated.
Kai Hu 0004, Ji Wan, Yuzhuang Xu, Zijing Cheng, Wei-Tek Tsai
Softw. Test. Verification Reliab.4