VLDB 2026 Research / reviewers in the wild / expert
Junrui Xiao
dblp:315/0363
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-5256-1091ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion ModelsabstractDiffusion models have achieved great success in image generation tasks. However, the lengthy denoising process and complex neural networks hinder their low-latency applications in real-world scenarios. Quantization can effectively reduce model complexity, and post-training quantization (PTQ), which does not require fine-tuning, is highly promising for compressing and accelerating diffusion models. Unfortunately, we find that due to the highly dynamic activations, existing PTQ methods suffer from distribution mismatch issues at both calibration sample level and reconstruction output level, which makes the performance far from satisfactory. In this paper, we propose EDA-DM, a standardized PTQ method that efficiently addresses the above issues. Specifically, at the calibration sample level, we extract information from the density and diversity of latent space feature maps, which guides the selection of calibration samples to align with the overall sample distribution; and at the reconstruction output level, we theoretically analyze the reasons for previous reconstruction failures and, based on this insight, optimize block reconstruction using the Hessian loss of layers, aligning the outputs of quantized model and full-precision model at different network granularity. Extensive experiments demonstrate that EDA-DM significantly outperforms the existing PTQ methods across various models and datasets. Our method achieves a $1.83\times $ speedup and $4\times $ compression for the popular Stable-Diffusion on MS-COCO, with only a 0.05 loss in CLIP score. Code is available at http://github.com/BienLuky/EDA-DM. Zhikai Li, Junrui Xiao, Mengjuan Chen, Qingyi Gu |
IEEE Trans. Image Process. | 3 |
| 2026 | CoLeQ: Improving Data-Free Quantization via Contrastive LearningabstractModel quantization is an effective approach to reduce the complexity of neural networks, enabling them to be deployed on resource-constrained edge devices. Recently, data-free quantization has been widely investigated, since it does not access the original datasets and can address the widely-held data privacy and security concerns. Its idea is to generate fake data depending on the prior information in the full-precision (FP) model, and then fine-tune the quantized model with them under the supervision of the FP model. The quantization performance relies heavily on the validity of the generated data, however, existing methods suffer from two severe issues: mode collapse and (catastrophic) example forgetting, leading to non-trivial accuracy degradation. In this work, we proposeContrastiveLearningQuantization (CoLeQ), which achieves data diversity enhancement and old knowledge restoration via contrastive learning to address the above issues. Specifically, we introduce the MoCo paradigm that maintains a dynamic momentum queue of the encoded features to data-free quantization. The contrastive learning objective is used to improve data diversity by facilitating the separation of generated samples from the already generated ones in previous mini-batches, thus mitigating the mode collapse problem. Moreover, we design a tied-weight decoder to restore the previous samples from the encoded features in the queue without additional parameters and training, hence cost-effectively preventing the example forgetting problem. Extensive experiments are conducted to evaluate the effectiveness of CoLeQ, and the results demonstrate a consistent superiority compared to state-of-the-art methods. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Multim. | 3 |
| 2025 | BinaryViT: Toward Efficient and Accurate Binary Vision TransformersabstractVision Transformers (ViTs) have emerged as the new fundamental architecture for most computer vision fields. However, the considerable memory and computation costs also hinder their application on resource-limited devices. Currently, binarization has demonstrated remarkable potential as a model compression technique in traditional Convolutional Neural Networks (CNNs), albeit with some accuracy loss. In this paper, we focus on binarization of ViTs, which is still under-studied and suffering a significant performance drop. We start with constructing a strong baseline of binary ViTs, integrating some of the best practices from binary CNNs, which forms the foundation of our exploration. Subsequently, we identify that the severe performance degradation of the baseline is mainly caused by the weight oscillation around the quantization boundary and the information distortion in the activation of ViTs. To address these challenges, we introduce BinaryViT, a precise full binarization framework tailored for Vision Transformers (ViTs), effectively pushing the binarization of ViTs to its limit. Specifically, we propose a novel gradient regularization scheme (GRS), which mitigates oscillations by fostering a smooth moving of latent weights to be away from the quantization boundary during the training process. Additionally, we have devised an Activation Shift Module (ASM) that dynamically adjusts the activation distribution prior to the sign function, thereby minimizing the information distortion stemming from the significant inter-channel variations. Extensive experiments on ImageNet dataset show that our BinaryViT consistently surpasses the strong baseline by 2.05% and improves the accuracy of fully binarized ViTs to a usable level. Furthermore, our method achieves impressive savings of$16.2\times $and$17.7\times $in model size and OPs compared to the full-precision DeiT-S. Junrui Xiao, Zhikai Li, Lianwei Yang, Qingyi Gu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity ReconstructionabstractPost-training quantization (PTQ) efficiently compresses vision models, but unfortunately, it accompanies a certain degree of accuracy degradation. Reconstruction methods aim to enhance model performance by narrowing the gap between the quantized model and the full-precision model, often yielding promising results. However, efforts to significantly improve the performance of PTQ through reconstruction in the Vision Transformer (ViT) have shown limited efficacy. In this paper, we conduct a thorough analysis of the reasons for this limited effectiveness and propose MGRQ (Mixed Granularity Reconstruction Quantization) as a solution to address this issue. Unlike previous reconstruction schemes, MGRQ introduces a mixed granularity reconstruction approach. Specifically, MGRQ enhances the performance of PTQ by introducing Extra-Block Global Supervision and Intra-Block Local Supervision, building upon Optimized Block-wise Reconstruction. Extra-Block Global Supervision considers the relationship between block outputs and the model’s output, aiding block-wise reconstruction through global supervision. Meanwhile, Intra-Block Local Supervision reduces generalization errors by aligning the distribution of outputs at each layer within a block. Subsequently, MGRQ is further optimized for reconstruction through Mixed Granularity Loss Fusion. Extensive experiments conducted on various ViT models illustrate the effectiveness of MGRQ. Notably, MGRQ demonstrates robust performance in low-bit quantization, thereby enhancing the practicality of the quantized model. Lianwei Yang, Zhikai Li, Junrui Xiao, Haisong Gong, Qingyi Gu |
ICIP | 3 |
| 2024 | HTQ: Exploring the High-Dimensional Trade-Off of mixed-precision quantization
Zhikai Li, Xianlei Long, Junrui Xiao, Qingyi Gu |
Pattern Recognit. | 3 |
| 2024 | PSAQ-ViT V2: Toward Accurate and General Data-Free Quantization for Vision TransformersabstractData-free quantization can potentially address data privacy and security concerns in model compression and thus has been widely investigated. Recently, patch similarity aware data-free quantization for vision transformers (PSAQ-ViT) designs a relative value metric, patch similarity, to generate data from pretrained vision transformers (ViTs), achieving the first attempt at data-free quantization for ViTs. In this article, we propose PSAQ-ViT V2, a more accurate and general data-free quantization framework for ViTs, built on top of PSAQ-ViT. More specifically, following the patch similarity metric in PSAQ-ViT, we introduce an adaptive teacher-student strategy, which facilitates the constant cyclic evolution of the generated samples and the quantized model (student) in a competitive and interactive fashion under the supervision of the full-precision (FP) model (teacher), thus significantly improving the accuracy of the quantized model. Moreover, without the auxiliary category guidance, we employ the task- and model-independent prior information, making the general-purpose scheme compatible with a broad range of vision tasks and models. Extensive experiments are conducted on various models on image classification, object detection, and semantic segmentation tasks, and PSAQ-ViT V2, with the naive quantization strategy and without access to real-world data, consistently achieves competitive results, showing potential as a powerful baseline on data-free quantization for ViTs. For instance, with Swin-S as the (backbone) model, 8-bit quantization reaches 82.13 top-1 accuracy on ImageNet, 50.9 box AP and 44.1 mask AP on COCO, and 47.2 mean Intersection over Union (mIoU) on ADE20K. We hope that accurate and general PSAQ-ViT V2 can serve as a potential and practice solution in real-world applications involving sensitive data. Code is released and merged at: https://github.com/zkkli/PSAQ-ViT. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision TransformersabstractPost-training quantization (PTQ), which only requires a tiny dataset for calibration without end-to-end retraining, is a light and practical model compression technique. Recently, several PTQ schemes for vision transformers (ViTs) have been presented; unfortunately, they typically suffer from non-trivial accuracy degradation, especially in low-bit cases. In this paper, we propose RepQ-ViT, a novel PTQ framework for ViTs based on quantization scale reparameterization, to address the above issues. RepQ-ViT decouples the quantization and inference processes, where the former employs complex quantizers and the latter employs scale-reparameterized simplified quantizers. This ensures both accurate quantization and efficient inference, which distinguishes it from existing approaches that sacrifice quantization performance to meet the target hardware. More specifically, we focus on two components with extreme distributions: post-LayerNorm activations with severe inter-channel variation and post-Softmax activations with power-law features, and initially apply channel-wise quantization and log$\sqrt 2 $ quantization, respectively. Then, we reparameterize the scales to hardware-friendly layer-wise quantization and log2 quantization for inference, with only slight accuracy or computational costs. Extensive experiments are conducted on multiple vision tasks with different model variants, proving that RepQ-ViT, without hyperparameters and expensive reconstruction procedures, can outperform existing strong baselines and encouragingly improve the accuracy of 4-bit PTQ of ViTs to a usable level. Code is available at https://github.com/zkkli/RepQ-ViT. Zhikai Li, Junrui Xiao, Lianwei Yang, Qingyi Gu |
ICCV | 2 |
| 2023 | Patch-wise Mixed-Precision Quantization of Vision TransformerabstractAs emerging hardware begins to support mixed bit-width arithmetic computation, mixed-precision quantization is widely used to reduce the complexity of neural networks. However, Vision Transformers (ViT ViTs) require complex self-attention computation to guarantee the learning of powerful feature representations, which makes mixed-precision quantization of ViTs still challenging. In this paper, we propose a novel patch-wise mixed-precision quantization (PMQ) for efficient inference of ViT$s$. Specifically, we design a lightweight global metric, which is faster than existing methods, to measure the sensitivity of each component in ViT$s$to quantization errors. Moreover, we also introduce a pareto frontier approach to automatically allocate the optimal bit-precision according to the sensitivity. To further reduce the computational complexity of self-attention in inference stage, we propose a patch-wise module to reallocate bit-width of patches in each layer. Extensive experiments on the ImageNet dataset shows that our method greatly reduces the search cost and facilitates the application of mixed-precision quantization to ViTs. Junrui Xiao, Zhikai Li, Lianwei Yang, Qingyi Gu |
IJCNN | 1 |
| 2023 | DCIFPN: Deformable cross-scale interaction feature pyramid network for object detectionabstractAbstract Exploiting multi‐scale features is one of the most effective methods to recognize objects of different scales in object detection. Since image pyramid is time‐consuming, Feature Pyramid Network (FPN) becomes the most popular component used for obtaining pyramidal features. Despite its effectiveness, there still exist some intrinsic defects. In this work, it is attributed to insufficient information flow and a Deformable Cross‐scale Interaction Feature Pyramid Network (DCIFPN) is proposed, which aims to promote the information transfer process with content‐aware sampling and dynamic aggregation weights. More specifically, Deformable Semantic Enhancement Module (DSEM) is designed that can construct accurate information flow with dynamic aggregation weights. In addition, Deformable Spatial Refinement Module (DSRM) is proposed to enhance high‐level features with low‐level location details. When DCIFPN is deployed on RetinaNet and FCOS with ResNet‐50, the performance is improved by 1.6 AP and 1.1 AP, respectively, on the challenging MS COCO benchmark. Apart from one‐stage detectors, DCIFPN is also applicable to two‐stage methods such as Faster R‐CNN and Mask R‐CNN. Further experiments on Pascal VOC and CrowdHuman datasets can verify the effectiveness and generalization of the method. Junrui Xiao, Zhikai Li, Qingyi Gu |
IET Image Process. | 1 |
| 2023 | Temporal action detection with dynamic weights based on curriculum learning
Yunze Chen, Junrui Xiao, Qingyi Gu |
Neurocomputing | 3 |
| 2022 | Patch Similarity Aware Data-Free Quantization for Vision Transformers
Zhikai Li, Liping Ma, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
ECCV (11) | 4 |
| 2022 | Dual-discriminator adversarial framework for data-free quantization
Zhikai Li, Liping Ma, Xianlei Long, Junrui Xiao, Qingyi Gu |
Neurocomputing | 4 |
| 2022 | Rethinking prediction alignment in one-stage object detection
Junrui Xiao, Zhikai Li, Qingyi Gu |
Neurocomputing | 1 |