Xinrui Chen 0001

dblp:241/3777-1 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
0009-0002-8053-8494ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
abstract
Layer pruning is a viable technique for compressing large language models while achieving acceleration proportional to the pruning ratio. In this work, we identify that removing any layer induces a magnitude gap in hidden states, and demonstrate that a simple compensation operation leads to superior performance in iterative layer pruning. This key observation motivates us to propose Prune&Comp, a novel, plug-and-play iterative layer pruning scheme that leverages magnitude compensation to mitigate such gaps in a training-free manner. Specifically, we first estimate the magnitude gap of layer removal and then eliminate it by rescaling the remaining weights offline. We further demonstrate the advantages of Prune&Comp in improving the stability of iterative pruning. When integrated with an iterative prune-and-compensate loop, Prune&Comp consistently enhances existing layer pruning metrics. For instance, when 5 layers of LLaMA-3-8B are pruned with the prevalent Taylor+ metric, Prune&Comp reduces PPL from 512.78 to 16.34 and retains 90.57% of the original performance across 9 question-answering tasks, outperforming the baseline by 24.72%.
Xinrui Chen 0001, Fanyi Zeng, Yongxian Wei, Yizhi Wang 0002, Xitong Ling, Guanghao Li 0003, Chun Yuan 0003
AAAI1
2025 Zero-shot Quantization for Large-kernels via Shape-based Distribution and Diversity Self-distillation
abstract
Zero-shot quantization (ZSQ) has emerged as an effective method to reduce model complexity and memory footprint without using original training data, thereby mitigating data privacy and security concerns during model deployment. Recently, Large-Kernel Convolutional Neural Networks (LKCNNs) have achieved state-of-the-art performance on various vision tasks, which introduce challenges in terms of increased parameters and network complexity, making them difficult to deploy on resource-constrained edge devices. Despite the success of ZSQ, existing methods fail to apply to LKCNNs due to architectural differences such as Batch Normalization (BN) layers in models and thus result in significant performance declines. In this paper, we propose a novel ZSQ framework tailored specifically for LKCNNs, considering their two key characteristics: the large receptive field and the reliance on shape bias. Correspondingly, we first employ an edge detection-based loss to optimize synthetic images that closely mimic the distribution of real images, and a diversity self-distillation loss to maintain consistency in feature representation to enable the generation of synthetic images. Afterward, we use these synthetic images to fine-tune the quantization parameters with a shape-enhance data augmentation strategy. Experiment results demonstrate the superiority of the proposed framework over existing methods, with significant improvements in maintaining accuracy after quantization across various quantization configurations on the ImageNet dataset.
Zhuozhen Yu, Xinrui Chen 0001, Shunzhou Wang, Wei Gao 0003
ICASSP3
2025 Eff-DFQT: Efficient Model Inversion for Data-free Quantization of Vision Transformers
abstract
Model inversion is a promising technique for raw data reconstruction, especially in data-free quantization of Vision Transformers (ViTs). Previous inversion methods for ViTs have focused on extracting necessary foreground information while discarding irrelevant noise. However, these mode inversion methods for ViTs are inefficient in terms of data synthesis speed. In this paper, we propose a novel method to accelerate model inversion for efficient data-free quantization of ViTs(Eff-DFQT). Our method has the following features. 1) Token fusion strategy tailored for model inversion. We propose a token fusion strategy tailored for model inversion to lower the computations for image inversion. 2) Label compensation function. We propose a label compensation function to model the label uncertainty of the inverted image and accurately capture the real labels, which improves the quality of the inverted data by compensating for the negative effects of token reduction. Extensive experimental results demonstrate that Eff-DFQT, significantly accelerates the inversion process through token fusion strategy tailored for model inversion and label compensation function, while maintaining or even improving model performance in data-free quantization of ViTs.
Mengkui Li, Xinrui Chen 0001, Hai Chen, Fulan Qian
ICME2
2025 Zero-shot Quantization of Vision Transformers: Leveraging Multi-model Ensembles and Attention Mixup
abstract
Zero-shot quantization (ZSQ) shows promise in compressing and accelerating deep neural networks in scenarios where the original training data is inaccessible. Recently, ZSQ for vision transformers (ViTs) has been proposed to synthesize samples for ViT network quantization, the quality of which significantly impacts the performance of quantized models. Nonetheless, we observe that the synthetic samples produced by current ZSQ techniques exhibit severe bias and insufficient optimization, which deviate from real data and lead to substantial performance declines. On the one hand, the synthetic samples generated by a single ViT network are inaccurate and biased to the specific model. On the other hand, unlike convolutional neural networks (CNNs) with BatchNorm layers, ViTs do not store any training set statistics in the networks, hindering the generation of high-quality calibration samples. To address the above issues, we propose leveraging Multi-model Ensembles and Attention Mixup in ZSQ for ViTs (MMA-ViT). Specifically, MMA-ViT employs an ensemble of diverse pre-trained proxy CNN models to narrow the sample synthesizing space, utilizing their predictive capabilities and BatchNorm statistics to generate exact synthetic images that enhance generality. Additionally, MMA-ViT integrates a unique attention-driven mixup technique for accurate data augmentation during the sample synthesis process, avoiding over-fitting to the networks. The efficacy of MMA-ViT has been demonstrated through extensive experiments and ablation studies on the ImageNet dataset. For example, when Swin-B is quantized to W3/A4, our method achieves a 11.89% top-1 accuracy increase on ImageNet compared to state-of-the-art methods.
Xinrui Chen 0001, Zhuozhen Yu, Shunzhou Wang, Wei Gao 0003
ICME2
2025 A Simple Linear Patch Revives Layer-Pruned Large Language Models
abstract
Layer pruning has emerged as a widely used technique for compressing large language models (LLMs). However, existing layer pruning approaches often incur substantial performance degradation. We identify the majority of this degradation to a single yet previously overlooked issue: \textit{the mismatch of activation magnitudes at the pruning interface}. The pre-interface activations exhibit significantly different scales from the post-interface ones, causing the distributional shift as it propagates through the remaining layers. To address this issue, we introduce \textsc{LinearPatch}, a lightweight and plug-and-play technique that fuses two operations into one matrix multiply at the pruning interface: (i) a Hadamard transformation that suppresses massive outliers at particular tokens and (ii) a channel-wise scaling that aligns activation statistics. On LLaMA-3-8B, \textsc{LinearPatch} preserves up to \textbf{94.15\%} of the original model's performance when pruning 5 out of 32 layers, outperforming the previous state of the art by \textbf{4\%}. The patch can be further refined with 5K unlabeled samples via memory-efficient offline distillation, pushing the retention to 95.16\% within only 30 minutes on a single GPU. Code is available at \url{https://github.com/chenxinrui-tsinghua/LinearPatch}.
Xinrui Chen 0001, Haoli Bai, Ruikang Liu, Xianzhi Yu, Lu Hou 0002, Tian Guan, Yonghong He, Chun Yuan 0003
NeurIPS1
2025 Low-Bit-Width Zero-Shot Quantization With Soft Feature-Infused Hints for IoT Systems
abstract
Quantization has enabled the widespread implementation of deep learning algorithms on resource-constrained Internet of Things (IoT) devices, which compresses neural networks by reducing the bit-width of their parameters. However, most quantization methods invade privacy as they require real training datasets for calibration or fine-tuning. As a solution, zero-shot quantization (ZSQ) has emerged as a paradigm to quantize neural networks without accessing training datasets. Most employ data generation schemes to synthesize calibration data for knowledge transfer from the full-precision networks to the quantized ones. For privacy-protected and resource-constrained IoT devices, achieving optimal deployment necessitates the strategic integration of synthetic data generation and low-bit-width quantization techniques. However, when it comes to the lower bit-width case in ZSQ, we observe that the discrepancy between the full-precision network and the quantized network tends to widen significantly, hindering the knowledge transfer, which is attributed to the three following challenges: 1) hard logits matching with wide discrepancy; 2) unstable feature alignment with huge quantization error; and 3) synthetic data with low diversity. To address these issues, this article presents S-ZSQ, a novel ZSQ framework with two-pronged strategies that enhances both knowledge transfer and synthetic data generation, which enables low-bit-width quantized network to derive more soft feature-infused hints from the full-precision network. We achieve significant improvements on classification tasks, including CIFAR-10/100 and ImageNet-1k, with fewer fine-tuning epochs, particularly in scenarios involving low-bit-width quantization. For example, in the 3-bit ResNet-18/ResNet-50 case, we outperform AdaDFQ by 8.08%/11.16% in top-1 accuracy on ImageNet-1k.
Xinrui Chen 0001, Yizhi Wang 0002, Xitong Ling, Mengkui Li, Ruikang Liu, Minxi Ouyang, Tian Guan, Yonghong He
IEEE Internet Things J.1
2024 HIQ: One-Shot Network Quantization for Histopathological Image Classification
abstract
To deploy neural networks on clinical edge devices, quantization is the most commonly used method to compress the models, which requires a calibration set of hundreds of real images. However, due to privacy concerns, the scarcity of private histopathological images hinders the application of quantization. To address this issue, we develop HIQ, a novel one-shot quantization framework for histopathological image classification networks, which requires only one real image per class for calibration. To compensate for data scarcity, sample BNS alignment is introduced to generate synthetic images with similar distribution to the real ones. To improve the diversity of synthetic images, fine-grained diversity enhancement that provides fine-grained enhancement intensity for different classes and network layers is proposed, based on the observation of the class-wise and layer-wise fine-grained data. Finally, the asymptotic enhancement strategy is highlighted to achieve a trade-off between inter-class distance and intra-class diversity of synthetic images, based on the insight of the smaller inter-class distance of histopathological images than that of natural ones. Extensive experiments on the BRACS dataset show that our method achieves an extremely low accuracy loss even compared to the full precision model in low-bit cases and maintains robustness when missing classes of real images.
Xinrui Chen 0001, Renao Yan, Yizhi Wang 0002, Jiawen Li 0005, Junru Cheng, Tian Guan, Yonghong He
ICASSP1
2024 Agent Aggregator with Mask Denoise Mechanism for Histopathology Whole Slide Image Analysis
abstract
Histopathology analysis is the gold standard for medical diagnosis. Accurate classification of whole slide images (WSIs) and region-of-interests (ROIs) localization can assist pathologists in diagnosis. The gigapixel resolution of WSI and the absence of fine-grained annotations make direct classification and analysis challenging. In weakly supervised learning, multiple instance learning (MIL) presents a promising approach for WSI classification. The prevailing strategy is to use attention mechanisms to measure instance importance for classification. However, attention mechanisms fail to capture inter-instance information, and self-attention causes quadratic computational complexity. To address these challenges, we propose AMD-MIL, an agent aggregator with a mask denoise mechanism. The agent token acts as an intermediate variable between the query and key for computing instance importance. Mask and denoising matrices, mapped from agents-aggregated value, dynamically mask low-contribution representations and eliminate noise. AMD-MIL achieves better attention allocation by adjusting feature representations, capturing micro-metastases in cancer, and improving interpretability. Extensive experiments on CAMELYON-16, CAMELYON-17, TCGA-KIDNEY, and TCGA-LUNG show AMD-MIL's superiority over state-of-the-art methods.
Xitong Ling, Minxi Ouyang, Yizhi Wang 0002, Xinrui Chen 0001, Renao Yan, Hongbo Chu, Junru Cheng, Tian Guan, Sufang Tian, Yonghong He
ACM Multimedia4
2023 ADEQ: Adaptive Diversity Enhancement for Zero-Shot Quantization
Xinrui Chen 0001, Renao Yan, Junru Cheng, Yizhi Wang 0002, Yuqiu Fu, Tian Guan, Yonghong He
ICONIP (1)1
2023 TexQ: Zero-shot Network Quantization with Texture Feature Distribution Calibration
abstract
Quantization is an effective way to compress neural networks. By reducing the bit width of the parameters, the processing efficiency of neural network models at edge devices can be notably improved. Most conventional quantization methods utilize real datasets to optimize quantization parameters and fine-tune. Due to the inevitable privacy and security issues of real samples, the existing real-data-driven methods are no longer applicable. Thus, a natural method is to introduce synthetic samples for zero-shot quantization (ZSQ). However, the conventional synthetic samples fail to retain the detailed texture feature distributions, which severely limits the knowledge transfer and performance of the quantized model. In this paper, a novel ZSQ method, TexQ is proposed to address this issue. We first synthesize a calibration image and extract its calibration center for each class with a texture feature energy distribution calibration method. Then, the calibration centers are used to guide the generator to synthesize samples. Finally, we introduce the mixup knowledge distillation module to diversify synthetic samples for fine-tuning. Extensive experiments on CIFAR10/100 and ImageNet show that TexQ is observed to perform state-of-the-art in ultra-low bit width quantization. For example, when ResNet-18 is quantized to 3-bit, TexQ achieves a 12.18% top-1 accuracy increase on ImageNet compared to state-of-the-art methods. Code at https://github.com/dangsingrue/TexQ.
Xinrui Chen 0001, Yizhi Wang 0002, Renao Yan, Tian Guan, Yonghong He
NeurIPS1