Weilun Feng

dblp:219/1663 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
abstract
Diffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse.
Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Renshuai Tao, Yongjun Xu 0001, Michele Magno
AAAI1
2025 Multi-party Collaborative Attention Control for Image Customization
abstract
The rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leakage or confusion; 3) image-conditioned outputs tend to suffer from inconsistent backgrounds; and 4) high computational costs. To address these issues, this paper introduces Multi-party Collaborative Attention Control (MCA-Ctrl), a tuning-free method that enables high-quality image customization using both text and complex visual conditions. Specifically, MCA-Ctrl leverages two key operations within the self-attention layer to coordinate multiple parallel diffusion processes and guide the target image generation. This approach allows MCA-Ctrl to capture the content and appearance of specific subjects while maintaining semantic consistency with the conditional input. Additionally, to mitigate subject leakage and confusion issues common in complex visual scenarios, we introduce a Subject Localization Module that extracts precise subject and editable image layers based on user instructions. Extensive quantitative and human evaluation experiments show that MCA-Ctrl outperforms existing methods in zero-shot image customization, effectively resolving the mentioned issues.
Chuanguang Yang, Qiuli Wang 0001, Zhulin An, Weilun Feng, Libo Huang 0001, Yongjun Xu 0001
CVPR5
2025 Content-Aware Motion Compensated Temporal Filter for Video Coding
abstract
Video coding achieves efficient compression by exploiting the spatial and temporal correlations within the video signal. However, the noise in the source signal corrupts such correlation and impairs the coding performance. Motion compensated temporal filter (MCTF) [1] is a pre-processing tool that removes certain noise from the source signal, thereby enhancing temporal correlations among adjacent frames. Although numerous efforts have been made to optimize MCTF, MCTF still lacks flexibility in filtering for diverse video content, and its filtering efficiency is still limited. In this paper, we propose the Content-Aware MCTF method (CAMCTF) to enhance the filtering adaptability of MCTF for diverse video contents. The CAMCTF adaptively adjusts the filtering block sizes based on the Sum of Square Error (SSE) and Motion Vector (MV) information calculated during the motion estimation (ME) process in MCTF, and offset weighting method is applied to improve the prediction quality of filtering blocks. Performance was evaluated on top of Versatile Video Coding (VVC) reference software VTM-23.4. The VVC Common Test Conditions (CTC) [2] with QPs 22, 27, 32, 37 are used. As shown in Table 1, CAMCTF achieves an overall of 1.07% and 0.79% luma BD-rate gains in Random Access (RA) and Low Delay (LD) configurations, respectively. The encoder complexity is 112% and 110% for RA and LD configurations, respectively, without decoder complexity increasing.
Yunrui Jian, Meng Lei, Weilun Feng, Zhenan Lin, Chao Zhou 0003
DCC5
2025 Content-Adaptive Motion Compensated Temporal Filter for Versatile Video Coding
abstract
The exploitation of spatial and temporal correlation within video signals is a cornerstone of video coding, which is extensively employed to achieve high coding efficiency. The presence of noise in video signal deteriorates the correlations, leading to a degradation in coding efficiency. Motion Compensated Temporal Filter (MCTF) is a pre-processing tool designed to remove noise from the source signal, thereby enhancing temporal correlations among adjacent frames. In this paper, a Content-Adaptive MCTF method (CAMCTF) is proposed to enhance the filtering adaptability of MCTF for diverse video contents. The proposed CAMCTF method consists of Quadtree Block Partitioning Scheme (QBPS), Offset Block Weighted Compensation (OBWC) and Structural Similarity (SSIM) based Filtering Weight Adjustment (SSFWA). Specifically, QBPS generates size-adaptive Motion Compensation Block (MCB) to better cater to the video content characteristics. Subsequently, OBWC is employed to address the uneven prediction quality of MCBs. Furthermore, the filtering weights of each MCB are fine-tuned by SSFWA based on the SSIM index. The simulation results show that the proposed CAMCTF method can achieve 1.23% and 0.94% BD-rate gains for RA and LD configurations, respectively, on top of Versatile Video Coding (VVC) reference software VTM-23.4.
Yunrui Jian, Xueli Cheng, Weilun Feng, Zhenan Lin, Chao Zhou 0003
ICME5
2025 Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
abstract
Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters. Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present **Q-VDiT**, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the *Token aware Quantization Estimator* (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce *Temporal Maintenance Distillation* (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency score of 23.40, setting a new benchmark and outperforming the current state-of-the-art quantization methods by **1.9$\times$**.
Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Zhulin An, Libo Huang 0001, Boyu Diao, Zixiang Zhao, Yongjun Xu 0001, Michele Magno
ICML1
2025 Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learning
abstract
3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D FSCIL approaches primarily utilize multimodal pre-trained models to extract the semantic features, heavily dependent on meticulously designed high-quality prompts and fine-tuning strategies. To reduce this dependence, this paper proposes a novel method for **3D** **F**SCI**L** with **E**mbedded **G**eometric features (**3D-FLEG**). Specifically, 3D-FLEG develops a point cloud *geometric feature extraction module* to capture category-related geometric characteristics. To address the modality heterogeneity issues that arise from integrating geometric and text features, 3D-FLEG introduces a *geometric feature embedding module*. By augmenting text prompts with spatial geometric features through these modules, 3D-FLEG can learn robust representations of new categories even with limited samples, while mitigating forgetting of the previously learned categories. Experiments conducted on several publicly available 3D point cloud datasets, including ModelNet, ShapeNet, ScanObjectNN, and CO3D, demonstrate 3D-FLEG's superiority over existing state-of-the-art 3D FSCIL methods. Code is available at https://github.com/lixiangqi707/3D-FLEG.
Xiangqi Li, Libo Huang 0001, Zhulin An, Weilun Feng, Chuanguang Yang, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001
ICML4
2025 S2Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Zhulin An, Libo Huang 0001, Michele Magno, Yongjun Xu 0001
NeurIPS1
2024 Reg-PTQ: Regression-specialized Post-training Quantization for Fully Quantized Object Detector
abstract
Although deep learning based object detection is of great significance for various applications, it faces challenges when deployed on edge devices due to the computation and energy limitations. Post-training quantization (PTQ) can improve inference efficiency through integer computing. However, they suffer from severe performance degra-dation when performing full quantization due to overlooking the unique characteristics of regression tasks in ob-ject detection. In this paper, we are the first to explore regression-friendly quantization and conduct full quantization on various detectors. We reveal the intrinsic reason behind the difficulty of quantizing regressors with empir-ical and theoretical justifications, and introduce a novel Regression-specialized Post-Training Quantization (Reg- PTQ) scheme. It includes Filtered Global Loss Integration Calibration to combine the global loss with a two-step fil-tering mechanism, mitigating the adverse impact of false positive bounding boxes, and Learnable Logarithmic-Affine Quantizer tailored for the non-uniform distributed param-eters in regression structures. Extensive experiments on prevalent detectors showcase the effectiveness of the well-designed Reg-PTQ. Notably, our Reg-PTQ achieves 7.6x and 5.4x reduction in computation and storage consumption under INT4 with little performance degradation, which indicates the immense potential of fully quantized detectors in real-world object detection applications.
Yifu Ding 0001, Weilun Feng, Chuyan Chen, Jinyang Guo 0002, Xianglong Liu 0001
CVPR2
2024 Relational Diffusion Distillation for Efficient Image Generation
Weilun Feng, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001
ACM Multimedia1