VLDB 2026 Research / reviewers in the wild / expert
Mengjuan Chen
dblp:258/9897
· DBLP profile ↗
9ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-5445-3644ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EDA-DM: Enhanced Distribution Alignment for Post-Training Quantization of Diffusion ModelsabstractDiffusion models have achieved great success in image generation tasks. However, the lengthy denoising process and complex neural networks hinder their low-latency applications in real-world scenarios. Quantization can effectively reduce model complexity, and post-training quantization (PTQ), which does not require fine-tuning, is highly promising for compressing and accelerating diffusion models. Unfortunately, we find that due to the highly dynamic activations, existing PTQ methods suffer from distribution mismatch issues at both calibration sample level and reconstruction output level, which makes the performance far from satisfactory. In this paper, we propose EDA-DM, a standardized PTQ method that efficiently addresses the above issues. Specifically, at the calibration sample level, we extract information from the density and diversity of latent space feature maps, which guides the selection of calibration samples to align with the overall sample distribution; and at the reconstruction output level, we theoretically analyze the reasons for previous reconstruction failures and, based on this insight, optimize block reconstruction using the Hessian loss of layers, aligning the outputs of quantized model and full-precision model at different network granularity. Extensive experiments demonstrate that EDA-DM significantly outperforms the existing PTQ methods across various models and datasets. Our method achieves a $1.83\times $ speedup and $4\times $ compression for the popular Stable-Diffusion on MS-COCO, with only a 0.05 loss in CLIP score. Code is available at http://github.com/BienLuky/EDA-DM. Zhikai Li, Junrui Xiao, Mengjuan Chen, Qingyi Gu |
IEEE Trans. Image Process. | 4 |
| 2026 | CoLeQ: Improving Data-Free Quantization via Contrastive LearningabstractModel quantization is an effective approach to reduce the complexity of neural networks, enabling them to be deployed on resource-constrained edge devices. Recently, data-free quantization has been widely investigated, since it does not access the original datasets and can address the widely-held data privacy and security concerns. Its idea is to generate fake data depending on the prior information in the full-precision (FP) model, and then fine-tune the quantized model with them under the supervision of the FP model. The quantization performance relies heavily on the validity of the generated data, however, existing methods suffer from two severe issues: mode collapse and (catastrophic) example forgetting, leading to non-trivial accuracy degradation. In this work, we proposeContrastiveLearningQuantization (CoLeQ), which achieves data diversity enhancement and old knowledge restoration via contrastive learning to address the above issues. Specifically, we introduce the MoCo paradigm that maintains a dynamic momentum queue of the encoded features to data-free quantization. The contrastive learning objective is used to improve data diversity by facilitating the separation of generated samples from the already generated ones in previous mini-batches, thus mitigating the mode collapse problem. Moreover, we design a tied-weight decoder to restore the previous samples from the encoded features in the queue without additional parameters and training, hence cost-effectively preventing the example forgetting problem. Extensive experiments are conducted to evaluate the effectiveness of CoLeQ, and the results demonstrate a consistent superiority compared to state-of-the-art methods. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Multim. | 2 |
| 2025 | DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight DilationabstractModel quantization is a promising method for accelerating and compressing diffusion models. Nevertheless, since post-training quantization (PTQ) fails catastrophically at low-bit cases, quantization-aware training (QAT) is essential. Unfortunately, the wide range and time-varying activations in diffusion models sharply increase the complexity of quantization, making existing QAT methods inefficient. Equivalent scaling can effectively reduce activation range, but previous methods remain the overall quantization error unchanged. More critically, these methods significantly disrupt the original weight distribution, resulting in poor weight initialization and challenging convergence during QAT training. In this paper, we propose a novel QAT framework for diffusion models, called DilateQuant. Specifically, we propose Weight Dilation (WD) that maximally dilates the unsaturated in-channel weights to a constrained range through equivalent scaling. WD decreases the activation range while preserving the original weight range, which steadily reduces the quantization error and ensures model convergence. To further enhance accuracy and efficiency, we design a Temporal Parallel Quantizer (TPQ) to address the time-varying activations and introduce a Block-wise Knowledge Distillation (BKD) to reduce resource consumption in training. Extensive experiments demonstrate that DilateQuant significantly outperforms existing methods in terms of accuracy and efficiency. Zhikai Li, Mengjuan Chen, Qingyi Gu |
ACM Multimedia | 4 |
| 2024 | PSAQ-ViT V2: Toward Accurate and General Data-Free Quantization for Vision TransformersabstractData-free quantization can potentially address data privacy and security concerns in model compression and thus has been widely investigated. Recently, patch similarity aware data-free quantization for vision transformers (PSAQ-ViT) designs a relative value metric, patch similarity, to generate data from pretrained vision transformers (ViTs), achieving the first attempt at data-free quantization for ViTs. In this article, we propose PSAQ-ViT V2, a more accurate and general data-free quantization framework for ViTs, built on top of PSAQ-ViT. More specifically, following the patch similarity metric in PSAQ-ViT, we introduce an adaptive teacher-student strategy, which facilitates the constant cyclic evolution of the generated samples and the quantized model (student) in a competitive and interactive fashion under the supervision of the full-precision (FP) model (teacher), thus significantly improving the accuracy of the quantized model. Moreover, without the auxiliary category guidance, we employ the task- and model-independent prior information, making the general-purpose scheme compatible with a broad range of vision tasks and models. Extensive experiments are conducted on various models on image classification, object detection, and semantic segmentation tasks, and PSAQ-ViT V2, with the naive quantization strategy and without access to real-world data, consistently achieves competitive results, showing potential as a powerful baseline on data-free quantization for ViTs. For instance, with Swin-S as the (backbone) model, 8-bit quantization reaches 82.13 top-1 accuracy on ImageNet, 50.9 box AP and 44.1 mask AP on COCO, and 47.2 mean Intersection over Union (mIoU) on ADE20K. We hope that accurate and general PSAQ-ViT V2 can serve as a potential and practice solution in real-world applications involving sensitive data. Code is released and merged at: https://github.com/zkkli/PSAQ-ViT. Zhikai Li, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Patch Similarity Aware Data-Free Quantization for Vision Transformers
Zhikai Li, Liping Ma, Mengjuan Chen, Junrui Xiao, Qingyi Gu |
ECCV (11) | 3 |
| 2022 | A Flexible Calibration Algorithm for High-speed Bionic Vision System based on GalvanometerabstractTraditional gimbal-based bionic eye systems usually use a multi-degree-of-freedom mechanical platform to move the camera freely, which makes the structure complex and bulky. The galvanometer-based reflective bionic eye system uses a galvanometer to replace the traditional mechanical rotation structure, which separates the camera from the gimbal system, greatly simplifying the structure. However, there are currently few methods for calibrating such systems, mostly for object detection and tracking. In this paper, a flexible method for high-precision calibration of a galvanometer-based reflective bionic eye system is proposed. In this method, a planar target is used for the calibration of the bionic eye system. The effectiveness and accuracy of the method are evaluated by the reprojection error of the control voltage and the spatial localization of the binocular system. Experiments show that the error of the control voltage after calibration is less than 0.2%. At an indoor distance of about 7 m, the RMSE of spatial visual localization is less than 0.3 cm. Qing Li 0046, Mengjuan Chen, Qingyi Gu, Idaku Ishii |
IROS | 2 |
| 2022 | Class-wise boundary regression by uncertainty in temporal action detectionabstractAbstract Temporal action detection is a crucial aspect of video understanding. It aims to classify the action as well as locate the start and end boundaries of the action in the untrimmed videos. As deep learning is frequently utilized, the accuracy of annotation is crucial to boundary localization. However, it is observed that some annotation instances are ambiguous and the ambiguity varies between categories. To solve the problem above, a Gaussian model is built to estimate the boundary uncertainty for each instance. Based on instance uncertainty, category uncertainty is applied to describe the uncertainty of each category. By combining instance and category uncertainty, the boundaries of the selected proposals are refined and the ranking of candidate proposals is adjusted. Furthermore, overcorrection is avoided for categories with a high level of uncertainty. With the uncertainty approach, state‐of‐the‐art performance is achieved: 57.5% on THUMOS14 ([email protected]) and 35.4% on ActivityNet (mAP@Avg). Yunze Chen, Mengjuan Chen, Qingyi Gu |
IET Image Process. | 2 |
| 2020 | Refinement of Boundary Regression Using Uncertainty in Temporal Action Localization
Yunze Chen, Mengjuan Chen, Jiagang Zhu, Qingyi Gu |
BMVC | 2 |
| 2020 | Natural Scene Facial Expression Recognition with Dimension Reduction NetworkabstractAs an external manifestation of human emotions, expression recognition plays an important role in human-computer interaction. Although existing expression recognition methods performs perfectly on constrained frontal faces, there are still many challenges in expression recognition in natural scenes due to different unrestricted conditions. Expression classification belongs to a pattern recognition problem where intra-class distance is greater than the inter-class distance, which leads to severe over-fitting when using neural networks for expression recognition. This paper proposes a novel net-work structure called Dimension Reduction Network which can effectively reduce generalization error. By adding a data dimension reduction module before the general classification network, a lot of redundant information is filtered, and only useful information is left. This can reduce the interference by irrelevant information when performing classification tasks and reduce generalization error. The proposed method does not require any modification to the classification network, only a small dimension reduction module needs to be added in front of the classification network. However, it can effectively reduce generalization error. We designed big and tiny versions of Dimension Reduction Network, both exceeds our baseline on AffectNet data set. The big version of our proposed method surpassed the state-of-the-art methods by more than 1.2% on AffectNet data set. Our code will open source3when the paper is accepted. Shenhua Hu, Yiming Hu, Xianlei Long, Mengjuan Chen, Qingyi Gu |
ICRA | 5 |