VLDB 2026 Research / reviewers in the wild / expert
Chuanguang Yang
dblp:241/6325
· DBLP profile ↗
45ranked-venue papers
12as first author
40since 2021 · last 2026
0000-0001-5890-289XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 10 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 9 first-author · 24 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic ConsistencyabstractCross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited availability of paired modalities. To this end, we investigate a general and effective knowledge learning concept under weak semantic consistency, dubbed Asymmetric Cross-modal Knowledge Distillation (ACKD), aiming to bridge modalities with limited semantic overlap. Nevertheless, the shift from strong to weak semantic consistency improves flexibility but exacerbates challenges in knowledge transmission costs, which we rigorously verified based on optimal transport theory. To mitigate the issue, we further propose a framework, namely SemBridge, integrating a Student-Friendly Matching module and a Semantic-aware Knowledge Alignment module. The former leverages self-supervised learning to acquire semantic-based knowledge and provide personalized instruction for each student sample by dynamically selecting the relevant teacher samples. The latter seeks the optimal transport path by employing Lagrangian optimization. To facilitate the research, we curate a benchmark dataset derived from two modalities, namely Multi-Spectral (MS) and asymmetric RGB images, tailored for remote sensing scene classification. Comprehensive experiments exhibit that our framework achieves state-of-the-art performance compared with 7 existing approaches on 6 different model architectures across various datasets. Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang 0039, Zhuoyan Gao, Chao Li 0028 |
AAAI | 3 |
| 2026 | Incentivizing Agentic Reasoning Capability with Outcome Supervision for Knowledge Base Question Answering
Fei Wang 0014, Zixuan Li 0001, Zhao Zhang 0011, Weiwei Ding, Chuanguang Yang, Yongjun Xu 0001, Xiaolong Jin 0001 |
WWW | 6 |
| 2026 | Towards robust medical image segmentation: Spectro-spatial domain generalization with MRAM and DMIR
Junhao Dong 0001, Hansheng Zeng, Fuyan Zhang, Zeyu Dong, Chuanguang Yang, Yingli Tian |
Comput. Vis. Image Underst. | 6 |
| 2025 | MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsabstractDiffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse. Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Renshuai Tao, Yongjun Xu 0001, Michele Magno |
AAAI | 3 |
| 2025 | HSRDiff: A Hierarchical Self-Regulation Diffusion Model for Stochastic Semantic SegmentationabstractIn safety-critical domains such as medical diagnostics and autonomous driving, single-image evidence is sometimes insufficient to reflect the inherent ambiguity of vision problems. Therefore, multiple plausible assumptions that match the image semantics may be needed to reflect the actual distribution of targets and support downstream tasks. However, balancing and improving the diversity and consistency of segmentation predictions under the high-dimensional output spaces and potential multimodal distributions is still challenging. This paper presents Hierarchical Self-Regulation Diffusion (HSRDiff), a unified framework that simulates joint probability distribution over entire labels. Our model self-regulates the balance between the two modes of predicting the label and noise in a novel ``differentiation to unification" pipeline and dynamically fits the optimal path to model the aleatoric uncertainty rooted in observations. In addition, we preserve the high-fidelity reconstruction of the delicate structure in images by leveraging the hierarchical multi-scale condition priors. We validate HSRDiff in three different semantic scenarios. Experimental results show that HSRDiff is superior to the comparison method with a considerable performance gap. Chuanguang Yang, Zhulin An, Libo Huang 0001, Yongjun Xu 0001 |
AAAI | 2 |
| 2025 | Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionabstractMulti-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation strengths among various teachers. Most existing methods often develop weighting strategies from an individual perspective of teacher performance or teacher-student gaps, lacking comprehensive information for guidance. This paper proposes Multi-Teacher Knowledge Distillation with Reinforcement Learning (MTKD-RL) to optimize multi-teacher weights. In this framework, we construct both teacher performance and teacher-student gaps as state information to an agent. The agent outputs the teacher weight and can be updated by the return reward from the student. MTKD-RL reinforces the interaction between the student and teacher using an agent in an RL-based decision mechanism, achieving better matching capability with more meaningful weights. Experimental results on visual recognition tasks, including image classification, object detection, and semantic segmentation tasks, demonstrate that MTKD-RL achieves state-of-the-art performance compared to the existing multi-teacher KD works. Chuanguang Yang, Xinqiang Yu, Zhulin An, Chengqing Yu, Libo Huang 0001, Yongjun Xu 0001 |
AAAI | 1 |
| 2025 | Multi-party Collaborative Attention Control for Image CustomizationabstractThe rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leakage or confusion; 3) image-conditioned outputs tend to suffer from inconsistent backgrounds; and 4) high computational costs. To address these issues, this paper introduces Multi-party Collaborative Attention Control (MCA-Ctrl), a tuning-free method that enables high-quality image customization using both text and complex visual conditions. Specifically, MCA-Ctrl leverages two key operations within the self-attention layer to coordinate multiple parallel diffusion processes and guide the target image generation. This approach allows MCA-Ctrl to capture the content and appearance of specific subjects while maintaining semantic consistency with the conditional input. Additionally, to mitigate subject leakage and confusion issues common in complex visual scenarios, we introduce a Subject Localization Module that extracts precise subject and editable image layers based on user instructions. Extensive quantitative and human evaluation experiments show that MCA-Ctrl outperforms existing methods in zero-shot image customization, effectively resolving the mentioned issues. Chuanguang Yang, Qiuli Wang 0001, Zhulin An, Weilun Feng, Libo Huang 0001, Yongjun Xu 0001 |
CVPR | 2 |
| 2025 | Cross-Layer Graph Knowledge Distillation for Image RecognitionabstractKnowledge Distillation (KD) aims to improve a light-weight student network supervised by a large teacher network. The core idea of KD is to explore valuable knowledge from the teacher. Previous works often extract information from a single sample, but ignore relation modeling among multiple samples between student and teacher. Therefore, we propose Cross-Layer Graph Knowledge Distillation (CLGKD) that conducts graph-augmented feature and relation distillation assisted by graph neural networks. We further propose a meta-learning mechanism to optimize cross-layer matching weights for promoting GKD among all student and teacher layers. Experimental results on image classification and object detection demonstrate that CLGKD achieves state-of-the-art performance compared to other KD methods. Our code is available at https://github.com/cynmzzz/ICASSP2025-CLGKD Jiaming Chu, Yanzhuo Xiang, Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICASSP | 4 |
| 2025 | Prototype-Driven Multi-Feature Generation for Visible-Infrared Person Re-identificationabstractThe primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by varying viewpoints and irregular movements. Existing methods often rely on horizontal partitioning to align part-level features, which can introduce inaccuracies and have limited effectiveness in reducing modality discrepancies. In this paper, we propose a novel Prototype-Driven Multi-feature generation framework (PDM) aimed at mitigating cross-modal discrepancies by constructing diversified features and mining latent semantically similar features for modal alignment. PDM comprises two key components: Multi-Feature Generation Module (MFGM) and Prototype Learning Module (PLM). The MFGM generates diversity features closely distributed from modality-shared features to represent pedestrians. Additionally, the PLM utilizes learnable prototypes to excavate latent semantic similarities among local features between visible and infrared modalities, thereby facilitating cross-modal instance-level alignment. We introduce the cosine heterogeneity loss to enhance prototype diversity for extracting rich local features. Extensive experiments conducted on the SYSU-MM01 and LLCM datasets demonstrate that our approach achieves state-of-the-art performance. Our codes are available at https://github.com/mmunhappy/ICASSP2025-PDM. Zeyu Dong, Chuanguang Yang |
ICASSP | 6 |
| 2025 | OLN++: Improved Object Localization Network for Open-world Object DetectionabstractOpen-world object detection (OWOD) is vital for identifying the new objects not encountered during training. Among the various methods for OWOD, Object Proposals without Learning Classification (OPwLC) stands out, with its Object Localization Network (OLN) stressing the localization features. However, OLN overlooks classification features, leading OPwLC to identify parts of a single object as multiple objects mistakenly. Inspired by the non-maximum suppression (NMS) technique, known for eliminating low-confidence detections, we sought to integrate NMS into OPwLC. However, direct integration of NMS into OPwLC presents a challenge, as OLN does not generate classification confidence scores, which are critical for applying NMS. To address this limitation, we developed a confidence measure module and proposed OLN++, filling the confidence scores gap. OLN++ can be easily implemented with just a few fully connected layers. We evaluated the effectiveness of OLN++ using NMS, Soft-NMS, and the Weighted Box Fusion variant on open-world detection tasks. Experimental results demonstrate that OLN++ significantly outperforms the original OLN. Haonan Mai, Libo Huang 0001, Zhulin An, Jiarui Zhao, Chuanguang Yang, Erhu Zhao, Yongjun Xu 0001 |
ICASSP | 5 |
| 2025 | ECG-guided individual identification via PPGabstractPhotoplethsmography (PPG)-based individual identification aiming at recognizing humans via intrinsic cardiovascular activities has raised extensive attention due to its high security and resistance to mimicry. However, this kind of technology witnesses unpromising results due to the limitation of low information density. To this end, electrocardiogram (ECG) signals have been introduced as a novel modality to enhance the density of input information. Specifically, a novel cross-modal knowledge distillation framework is implemented to propagate discriminate knowledge from ECG modality to PPG modality without incurring additional computational demands at the inference phase. Furthermore, to ensure efficient knowledge propagation, Contrastive Language–Image Pre-training (CLIP)-based knowledge alignment and cross-knowledge assessment modules are proposed respectively. Comprehensive experiments are conducted and results show our framework outperforms the baseline model with the improvement of 2.8% and 3.0% in terms of overall accuracy on seen- and unseen individual recognitions. Riling Wei, Kelu Yao, Chuanguang Yang, Chao Li 0028 |
ICASSP | 4 |
| 2025 | Enhancing Image Generation Fidelity via Progressive PromptsabstractDiffusion transformer (DiT) architecture catches much attention in image generation, which achieves better fidelity, performance, and diversity. However, most existing DiT-based image generation methods are global-aware synthesis and regional prompt control is less explored. In this paper, we propose a coarse-to-fine generation pipeline for regional prompt-following generation. Specifically, we first leverage the powerful large language model (LLM) to generate the high-level description of image (such as content, topic, and objects) and low-level description of image (such as details and style). Then we explore the influence of cross-attention layers in different depths. We discover that deeper layers always responsible for the high-level content control, while the shallow layers handles low-level content control. The various prompts are injected into the proposed regional cross-attention control in order for course-to-fine generation. Using the proposed pipeline, we improve the controllability of DiT-based image generation. Extensive quantitative and qualitative results demonstrate that our pipeline enables to improve the generated performance. Our codes are available at https://github.com/ZhenXiong-dl/ICASSP2025-RCAC. Zhen Xiong, Chuanguang Yang, Tiao Tan, Zhihong Zhu 0001 |
ICASSP | 3 |
| 2025 | Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingabstractSpatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation (termed SDKD), which transfers the multi-scale spatiotemporal representations from a complex teacher model to a more efficient lightweight student network. The teacher model follows an encoder-latent evolution-decoder architecture, where its latent evolution module decouples high-frequency details and low-frequency trends using convolution and Transformer (global low-frequency modeler). However, the multi-layer convolution and deconvolution structures result in slow training and high memory usage. To address these issues, we propose a frequency-aligned knowledge distillation strategy, which extracts multi-scale spectral features from the teacher's latent space, including both high and low frequency components, to guide the lightweight student model in capturing both local fine-grained variations and global evolution patterns. Experimental results show that SDKD significantly improves performance, achieving reductions of up to 81.3% in MSE and in MAE 52.3% on the Navier-Stokes equation dataset. The framework effectively captures both high-frequency variations and long-term trends while reducing computational complexity. Our codes are available at https://github.com/itsnotacie/SDKD Chuanguang Yang, Hansheng Zeng, Zeyu Dong, Zhulin An, Yongjun Xu 0001, Yingli Tian, Hao Wu 0094 |
ICCV | 2 |
| 2025 | Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersabstractDiffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters.
Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present **Q-VDiT**, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the *Token aware Quantization Estimator* (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce *Temporal Maintenance Distillation* (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency score of 23.40, setting a new benchmark and outperforming the current state-of-the-art quantization methods by **1.9$\times$**. Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Zhulin An, Libo Huang 0001, Boyu Diao, Zixiang Zhao, Yongjun Xu 0001, Michele Magno |
ICML | 2 |
| 2025 | Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learningabstract3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D FSCIL approaches primarily utilize multimodal pre-trained models to extract the semantic features, heavily dependent on meticulously designed high-quality prompts and fine-tuning strategies. To reduce this dependence, this paper proposes a novel method for **3D** **F**SCI**L** with **E**mbedded **G**eometric features (**3D-FLEG**). Specifically, 3D-FLEG develops a point cloud *geometric feature extraction module* to capture category-related geometric characteristics. To address the modality heterogeneity issues that arise from integrating geometric and text features, 3D-FLEG introduces a *geometric feature embedding module*. By augmenting text prompts with spatial geometric features through these modules, 3D-FLEG can learn robust representations of new categories even with limited samples, while mitigating forgetting of the previously learned categories. Experiments conducted on several publicly available 3D point cloud datasets, including ModelNet, ShapeNet, ScanObjectNN, and CO3D, demonstrate 3D-FLEG's superiority over existing state-of-the-art 3D FSCIL methods. Code is available at https://github.com/lixiangqi707/3D-FLEG. Xiangqi Li, Libo Huang 0001, Zhulin An, Weilun Feng, Chuanguang Yang, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ICML | 5 |
| 2025 | Merlin: Multi-View Representation Learning for Robust Multivariate Time Series Forecasting with Unfixed Missing RatesabstractMultivariate Time Series Forecasting (MTSF) involves predicting future values of multiple interrelated time series. Recently, deep learning-based MTSF models have gained significant attention for their promising ability to mine semantics (global and local information) within MTS data. However, these models are pervasively susceptible to missing values caused by malfunctioning data collectors. These missing values not only disrupt the semantics of MTS, but their distribution also changes over time. Nevertheless, existing models lack robustness to such issues, leading to suboptimal forecasting performance. To this end, in this paper, we propose Multi-View Representation Learning (Merlin), which can help existing models achieve semantic alignment between incomplete observations with different missing rates and complete observations in MTS. Specifically, Merlin consists of two key modules: offline knowledge distillation and multi-view contrastive learning. The former utilizes a teacher model to guide a student model in mining semantics from incomplete observations, similar to those obtainable from complete observations. The latter improves the student model's robustness by learning from positive/negative data pairs constructed from incomplete observations with different missing rates, ensuring semantic alignment across different missing rates. Therefore, Merlin is capable of effectively enhancing the robustness of existing models against unfixed missing rates while preserving forecasting accuracy. Experiments on four real-world datasets demonstrate the superiority of Merlin. Chengqing Yu, Fei Wang 0014, Chuanguang Yang, Zezhi Shao, Tao Sun 0011, Tangwen Qian, Wei Wei 0002, Zhulin An, Yongjun Xu 0001 |
KDD (2) | 3 |
| 2025 | Accelerating Diffusion Models via Parallel Denoising
Yanming Chen 0002, Zixin Ma, Chuanguang Yang, Zhulin An, Yiwen Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | S2Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Zhulin An, Libo Huang 0001, Michele Magno, Yongjun Xu 0001 |
NeurIPS | 3 |
| 2025 | Achieving fair medical image segmentation in foundation models with adversarial visual prompt tuning
Kai Zhang 0029, Fuyan Zhang, Chuanguang Yang, Zhongliang Guo 0001, Weiping Ding 0001, Tingwen Huang |
Inf. Sci. | 5 |
| 2025 | Enhancing spatiotemporal prediction through the integration of Mamba state space models and Diffusion TransformersabstractThis paper presents an advanced architecture for spatiotemporal prediction MAD , integrating Ma mba modules with D iffusion Transformers for efficient spatiotemporal modeling. The model consists of three phases: encoding, reconstruction, and prediction. Initially, the encoder transforms raw spatiotemporal data into compact latent embeddings. In the reconstruction phase, the Mamba module processes these embeddings through normalization and bidirectional state space models, generating reconstructed representations which are then decoded to restore the input data. The prediction phase utilizes the Diffusion Transformer to model spatiotemporal features, incorporating time embeddings and leveraging self-attention mechanisms to capture complex spatiotemporal dependencies. Finally, the model jointly trains the reconstruction and prediction paths to achieve high-precision spatiotemporal forecasts. Experimental results demonstrate the model’s superior performance across various spatiotemporal prediction tasks, validating its effectiveness and robustness. Our codes are available at https://github.com/Hanson1331/KBS-MAD . Hansheng Zeng, Ruize Niu, Chuanguang Yang, Shiping Wen 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Spatio-temporal masked autoencoder-based phonetic segments classification from ultrasound
Xi Dan, Kele Xu, Yihang Zhou, Chuanguang Yang, Yutao Dou, Cheng Yang 0004 |
Speech Commun. | 4 |
| 2024 | eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented GenerationabstractClass incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) to train a feature extractor incrementally, or store prototypes (features of learned tasks) to estimate the incremental feature distribution. However, the stored exemplars would violate the data privacy concerns, while the fixed prototypes might not reasonably be consistent with the incremental feature distribution, hindering the exploration of real-world CIL applications. In this paper, we propose a data-free CIL method with embedding distillation and Task-oriented generation (eTag), which requires neither exemplar nor prototype. Embedding distillation prevents the feature extractor from forgetting by distilling the outputs from the networks' intermediate blocks. Task-oriented generation enables a lightweight generator to produce dynamic features, fitting the needs of the top incremental classifier. Experimental results confirm that the proposed eTag considerably outperforms state-of-the-art methods on several benchmark datasets. Libo Huang 0001, Yan Zeng 0002, Chuanguang Yang, Zhulin An, Boyu Diao, Yongjun Xu 0001 |
AAAI | 3 |
| 2024 | Class-wise Image Mixture Guided Self-Knowledge Distillation for Image ClassificationabstractWe propose a novel regularization method to effectively train a neural network for avoiding overfitting, thus improving the performance. The core idea is to bridge the gap between predictive distributions derived from two popular image mixture techniques Mixup and CutMix by an ensemble distribution in a class-wise manner. Consistent optimization towards these three distributions is conducted by mutual distillation to guide the model to alleviate over-confidence predictions and robustly learn discriminative features as the classification evidence. Experiments across various image classification tasks show that our method significantly achieves better performance than previous data augmentation Mixup+CutMix and Self-KD methods. Zeyu Dong, Chuanguang Yang, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
CSCWD | 2 |
| 2024 | Online Relational Knowledge Distillation for Image ClassificationabstractExisting online Knowledge Distillation (KD) often perform probability-based predictions from independent data samples for knowledge transfer. However, these online KD methods neglect valuable relational information across multiple networks. To address this problem, we propose Online Relational Knowledge Distillation (ORKD). ORKD includes a discriminative loss to construct meaningful feature space and a relational distillation loss to guide structured knowledge transfer among multiple networks. Beyond feature-level distillation, we further construct an ensemble teacher by aggregating probability predictions from multiple networks. The virtual teacher is used to supervise a specific network to enhance its accuracy and avoid the cohort homogenization problem. Experimental results on CIFAR-100 and ImageNet classification demonstrate that ORKD achieves the best performance among state-of-the-art online KD methods over various network architectures. The qualitative visualization shows that ORKD can help the network to learn a more discriminative feature space, resulting in better classification performance. Yihang Zhou, Chuanguang Yang, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
CSCWD | 2 |
| 2024 | CLIP-KD: An Empirical Study of CLIP Model DistillationabstractContrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by a large teacher CLIP model. We propose several distillation strategies, including relation, feature, gradient and contrastive paradigms, to examine the effectiveness of CLIP-Knowledge Distillation (KD). We show that a simple feature mimicry with Mean Squared Error loss works surprisingly well. Moreover, interactive contrastive learning across teacher and student encoders is also effective in performance improvement. We explain that the success of CLIP-KD can be attributed to maximizing the feature similarity between teacher and student. The unified method is applied to distill several student models trained on CC3M+12M. CLIP-KD improves student CLIP models consistently over zero-shot ImageNet classification and cross-modal retrieval bench-marks. When using ViT-U14 pretrained on Laion-400M as the teacher, CLIP-KD achieves 57.5% and 55.4% zero-shot top-1 ImageNet accuracy over ViT-B/16 and ResNet-50, surpassing the original CLIP without KD by 20.5% and 20.1% margins, respectively. Our code is released on https://github.com/winycg/CLIP-KD. Chuanguang Yang, Zhulin An, Libo Huang 0001, Junyu Bi, Xinqiang Yu, Boyu Diao, Yongjun Xu 0001 |
CVPR | 1 |
| 2024 | DetKDS: Knowledge Distillation Search for Object DetectorsabstractIn this paper, we present DetKDS, the first framework that searches for optimal detection distillation policies. Manual design of detection distillers becomes challenging and time-consuming due to significant disparities in distillation behaviors between detectors with different backbones, paradigms, and label assignments. To tackle these challenges, we leverage search algorithms to discover optimal distillers for homogeneous and heterogeneous student-teacher pairs. Firstly, our search space encompasses global features, foreground-background features, instance features, logits response, and localization response as inputs. Then, we construct omni-directional cascaded transformations and obtain the distiller by selecting the advanced distance function and common weight value options. Finally, we present a divide-and-conquer evolutionary algorithm to handle the explosion of the search space. In this strategy, we first evolve the best distiller formulations of individual knowledge inputs and then optimize the combined weights of these multiple distillation losses. DetKDS automates the distillation process without requiring expert design or additional tuning, effectively reducing the teacher-student gap in various scenarios. Based on the analysis of our search results, we provide valuable guidance that contributes to detection distillation designs. Comprehensive experiments on different detectors demonstrate that DetKDS outperforms state-of-the-art methods in detection and instance segmentation tasks. For instance, DetKDS achieves significant gains than baseline detectors: $+3.7$, $+4.1$, $+4.0$, $+3.7$, and $+3.5$ AP on RetinaNet, Faster-RCNN, FCOS, RepPoints, and GFL, respectively. Code at: https://github.com/lliai/DetKDS. Lujun Li 0001, Yufan Bao, Peijie Dong, Chuanguang Yang, Anggeng Li, Wenhan Luo, Wei Xue 0002, Yike Guo |
ICML | 4 |
| 2024 | CPG: Channel Pruning with DFS Guided Grouping for Efficient Medical Image Segmentation
Xilin Yan, Fuyan Zhang, Boyuan Zhao, Hansheng Zeng, Chuanguang Yang |
ICONIP (4) | 9 |
| 2024 | Online Policy Distillation with Decision-AttentionabstractPolicy Distillation (PD) has become an effective method to improve deep reinforcement learning tasks. The core idea of PD is to distill policy knowledge from a teacher agent to a student agent. However, the teacher-student framework requires a well-trained teacher model which is computationally expensive. In the light of online knowledge distillation, we study the knowledge transfer between different policies that can learn diverse knowledge from the same environment. In this work, we propose Online Policy Distillation (OPD) with Decision-Attention (DA), an online learning framework in which different policies operate in the same environment to learn different perspectives of the environment and transfer knowledge to each other to obtain better performance together. With the absence of a well-performance teacher policy, the group-derived targets play a key role in transferring group knowledge to each student policy. However, naive aggregation functions tend to cause student policies quickly homogenize. To address the challenge, we introduce the Decision-Attention module to the online policies distillation framework. The Decision-Attention module can generate a distinct set of weights for each policy to measure the importance of group members. We use the Atari platform for experiments with various reinforcement learning algorithms, including PPO and DQN. In different tasks, our method can perform better than an independent training policy on both PPO and DQN algorithms. This suggests that our OPD-DA can transfer knowledge between different policies well and help agents obtain more rewards. Xinqiang Yu, Chuanguang Yang, Chengqing Yu, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
IJCNN | 2 |
| 2024 | Relational Diffusion Distillation for Efficient Image Generation
Weilun Feng, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ACM Multimedia | 2 |
| 2024 | Knowledge Distillation Using Hierarchical Self-Supervision Augmented DistributionabstractKnowledge distillation (KD) is an effective framework that aims to transfer meaningful information from a large teacher to a smaller student. Generally, KD often involves how to define and transfer knowledge. Previous KD methods often focus on mining various forms of knowledge, for example, feature maps and refined information. However, the knowledge is derived from the primary supervised task, and thus, is highly task-specific. Motivated by the recent success of self-supervised representation learning, we propose an auxiliary self-supervision augmented task to guide networks to learn more meaningful features. Therefore, we can derive soft self-supervision augmented distributions as richer dark knowledge from this task for KD. Unlike previous knowledge, this distribution encodes joint knowledge from supervised and self-supervised feature learning. Beyond knowledge exploration, we propose to append several auxiliary branches at various hidden layers, to fully take advantage of hierarchical feature maps. Each auxiliary branch is guided to learn self-supervision augmented tasks and distill this distribution from teacher to student. Overall, we call our KD method a hierarchical self-supervision augmented KD (HSSAKD). Experiments on standard image classification show that both offline and online HSSAKD achieves state-of-the-art performance in the field of KD. Further transfer experiments on object detection further verify that HSSAKD can guide the network to learn better features. The code is available at https://github.com/winycg/HSAKD. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingabstractVision-Language Pretraining (VLP) has significantly improved the performance of various vision-language tasks with the matching of images and texts. In this paper, we propose VL-Match, a Vision-Language framework with Enhanced Token-level and Instance-level Matching. At the token level, a Vision-Language Replaced Token Detection task is designed to boost the substantial interaction between text tokens and images, where the text encoder of VLP works as a generator to generate a corrupted text, and the multimodal encoder of VLP works as a discriminator to predict whether each text token in the corrupted text matches the image. At the instance level, in the Image-Text Matching task that judges whether an image-text pair is matched, we propose a novel bootstrapping method to generate hard negative text samples that are different from the positive ones only at the token level. In this way, we can force the network to detect fine-grained differences between images and texts. Notably, with a smaller amount of parameters, VL-Match significantly outperforms previous SOTA on all image-text retrieval tasks. Junyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang, Yuefeng Zhan, Chuanguang Yang, Yujing Wang 0002, Hao Sun 0015, Qi Zhang 0066 |
ICCV | 6 |
| 2023 | Online Knowledge Distillation via Mutual Contrastive Learning for Visual RecognitionabstractThe teacher-free online Knowledge Distillation (KD) aims to train an ensemble of multiple student models collaboratively and distill knowledge from each other. Although existing online KD methods achieve desirable performance, they often focus on class probabilities as the core knowledge type, ignoring the valuable feature representational information. We present a Mutual Contrastive Learning (MCL) framework for online KD. The core idea of MCL is to perform mutual interaction and transfer of contrastive distributions among a cohort of networks in an online manner. Our MCL can aggregate cross-network embedding information and maximize the lower bound to the mutual information between two networks. This enables each network to learn extra contrastive knowledge from others, leading to better feature representations, thus improving the performance of visual recognition tasks. Beyond the final layer, we extend MCL to intermediate layers and perform an adaptive layer-matching mechanism trained by meta-optimization. Experiments on image classification and transfer learning to visual recognition tasks show that layer-wise MCL can lead to consistent performance gains against state-of-the-art online KD approaches. The superiority demonstrates that layer-wise MCL can guide the network to generate better feature representations. Our code is publicly avaliable at https://github.com/winycg/L-MCL. Chuanguang Yang, Zhulin An, Helong Zhou, Fuzhen Zhuang, Yongjun Xu 0001, Qian Zhang 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Prior Gradient Mask Guided Pruning-Aware Fine-TuningabstractWe proposed a Prior Gradient Mask Guided Pruning-aware Fine-Tuning (PGMPF) framework to accelerate deep Convolutional Neural Networks (CNNs). In detail, the proposed PGMPF selectively suppresses the gradient of those ”unimportant” parameters via a prior gradient mask generated by the pruning criterion during fine-tuning. PGMPF has three charming characteristics over previous works: (1) Pruning-aware network fine-tuning. A typical pruning pipeline consists of training, pruning and fine-tuning, which are relatively independent, while PGMPF utilizes a variant of the pruning mask as a prior gradient mask to guide fine-tuning, without complicated pruning criteria. (2) An excellent tradeoff between large model capacity during fine-tuning and stable convergence speed to obtain the final compact model. Previous works preserve more training information of pruned parameters during fine-tuning to pursue better performance, which would incur catastrophic non-convergence of the pruned model for relatively large pruning rates, while our PGMPF greatly stabilizes the fine-tuning phase by gradually constraining the learning rate of those ”unimportant” parameters. (3) Channel-wise random dropout of the prior gradient mask to impose some gradient noise to fine-tuning to further improve the robustness of final compact model. Experimental results on three image classification benchmarks CIFAR10/ 100 and ILSVRC-2012 demonstrate the effectiveness of our method for various CNN architectures, datasets and pruning rates. Notably, on ILSVRC-2012, PGMPF reduces 53.5% FLOPs on ResNet-50 with only 0.90% top-1 accuracy drop and 0.52% top-5 accuracy drop, which has advanced the state-of-the-art with negligible extra computational cost. Linhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan, Yongjun Xu 0001 |
AAAI | 3 |
| 2022 | Mutual Contrastive Learning for Visual Representation LearningabstractWe present a collaborative learning method called Mutual Contrastive Learning (MCL) for general visual representation learning. The core idea of MCL is to perform mutual interaction and transfer of contrastive distributions among a cohort of networks. A crucial component of MCL is Interactive Contrastive Learning (ICL). Compared with vanilla contrastive learning, ICL can aggregate cross-network embedding information and maximize the lower bound to the mutual information between two networks. This enables each network to learn extra contrastive knowledge from others, leading to better feature representations for visual recognition tasks. We emphasize that the resulting MCL is conceptually simple yet empirically powerful. It is a generic framework that can be applied to both supervised and self-supervised representation learning. Experimental results on image classification and transfer learning to object detection show that MCL can lead to consistent performance gains, demonstrating that MCL can guide the network to generate better feature representations. Code is available at https://github.com/winycg/MCL. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
AAAI | 1 |
| 2022 | Cross-Image Relational Knowledge Distillation for Semantic SegmentationabstractCurrent Knowledge Distillation (KD) methods for semantic segmentation often guide the student to mimic the teacher's structured information generated from individual data samples. However, they ignore the global semantic relations among pixels across various images that are valuable for KD. This paper proposes a novel Cross-Image Relational KD (CIRKD), which focuses on transferring structured pixel-to-pixel and pixel-to-region relations among the whole images. The motivation is that a good teacher network could construct a well-structured feature space in terms of global pixel dependencies. CIRKD makes the student mimic better structured semantic relations from the teacher, thus improving the segmentation performance. Experimental results over Cityscapes, CamVid and Pascal VOC datasets demonstrate the effectiveness of our proposed approach against state-of-the-art distillation methods. The code is available at https://github.com/winycg/CIRKD. Chuanguang Yang, Helong Zhou, Zhulin An, Yongjun Xu 0001, Qian Zhang 0009 |
CVPR | 1 |
| 2022 | MixSKD: Self-Knowledge Distillation from Mixup for Image Recognition
Chuanguang Yang, Zhulin An, Helong Zhou, Linhang Cai, Xiang Zhi, Jiwen Wu, Yongjun Xu 0001, Qian Zhang 0009 |
ECCV (24) | 1 |
| 2022 | Localizing Semantic Patches for Accelerating Image ClassificationabstractExisting works often focus on reducing the architecture redundancy for accelerating image classification but ignore the spatial redundancy of the input image. This paper proposes an efficient image classification pipeline to solve this problem. We first pinpoint task-aware regions over the input image by a lightweight patch proposal network called AnchorNet. We then feed these localized semantic patches with much smaller spatial redundancy into a general classification network. Unlike the popular design of deep CNN, we aim to carefully design the Receptive Field of AnchorNet without intermediate convolutional paddings. This ensures the exact mapping from a high-level spatial location to the specific input image patch. The contribution of each patch is interpretable. Moreover, AnchorNet is compatible with any downstream architecture. Experimental results on ImageNet show that our method outperforms SOTA dynamic inference methods with fewer inference costs. Our code is available at https://github.com/winycg/AnchorNet. Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICME | 1 |
| 2021 | Multi-View Contrastive Learning for Online Knowledge DistillationabstractPrevious Online Knowledge Distillation (OKD) often carries out mutually exchanging probability distributions, but neglects the useful representational knowledge. We there-fore propose Multi-view Contrastive Learning (MCL) for OKD to implicitly capture correlations of feature embeddings encoded by multiple peer networks, which provide various views for understanding the input data instances. Benefiting from MCL, we can learn a more discriminative representation space for classification than previous OKD methods. Experimental results on image classification demonstrate that our MCL-OKD outperforms other state-of-the-art OKD methods by large margins without sacrificing additional inference cost. Codes are available at https://github.com/winycg/MCL-OKD. Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICASSP | 1 |
| 2021 | Hierarchical Self-supervised Augmented Knowledge DistillationabstractKnowledge distillation often involves how to define and transfer knowledge from teacher to student effectively. Although recent self-supervised contrastive knowledge achieves the best performance, forcing the network to learn such knowledge may damage the representation learning of the original class recognition task. We therefore adopt an alternative self-supervised augmented task to guide the network to learn the joint distribution of the original recognition task and self-supervised auxiliary task. It is demonstrated as a richer knowledge to improve the representation power without losing the normal classification capability. Moreover, it is incomplete that previous methods only transfer the probabilistic knowledge between the final layers. We propose to append several auxiliary classifiers to hierarchical intermediate feature maps to generate diverse self-supervised knowledge and perform the one-to-one transfer to teach the student network thoroughly. Our method significantly surpasses the previous SOTA SSKD with an average improvement of 2.56% on CIFAR-100 and an improvement of 0.77% on ImageNet across widely used network pairs. Codes are available at https://github.com/winycg/HSAKD. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
IJCAI | 1 |
| 2021 | Soft and Hard Filter Pruning via Dimension ReductionabstractFilter pruning is widely used to reduce the computation of deep learning, enabling the deployment of Deep Neural Networks (DNNs) in resource-limited devices. Conventional Hard Filter Pruning (HFP) method zeroizes pruned filters and stops updating them, thus reducing the search space of the model. On the contrary, Soft Filter Pruning (SFP) simply zeroizes pruned filters, keeping updating them in the following training epochs, thus maintaining the capacity of the network. However, SFP, together with its variants, converges much slower than HFP due to its larger search space. Firstly, we generalize SFP-based methods and HFP to analyze their characteristics. Then we propose a Gradually Hard Filter Pruning (GHFP) method to smoothly switch from SFP-based methods to HFP during training and pruning, thus maintaining a large search space at first, gradually reducing the capacity of the model to ensure a moderate convergence speed. Furthermore, we view filter pruning as dimension reduction and propose a novel dimension reduction block integrated into GHFP to significantly outperform other methods by a moderate margin. Linhang Cai, Zhulin An, Chuanguang Yang, Yongjun Xu 0001 |
IJCNN | 3 |
| 2020 | Gated Convolutional Networks with Hybrid Connectivity for Image ClassificationabstractWe propose a simple yet effective method to reduce the redundancy of DenseNet by substantially decreasing the number of stacked modules by replacing the original bottleneck by our SMG module, which is augmented by local residual. Furthermore, SMG module is equipped with an efficient two-stage pipeline, which aims to DenseNet-like architectures that need to integrate all previous outputs, i.e., squeezing the incoming informative but redundant features gradually by hierarchical convolutions as a hourglass shape and then exciting it by multi-kernel depthwise convolutions, the output of which would be compact and hold more informative multi-scale features. We further develop a forget and an update gate by introducing the popular attention modules to implement the effective fusion instead of a simple addition between reused and new features. Due to the Hybrid Connectivity (nested combination of global dense and local residual) and Gated mechanisms, we called our network as the HCGNet. Experimental results on CIFAR and ImageNet datasets show that HCGNet is more prominently efficient than DenseNet, and can also significantly outperform state-of-the-art networks with less complexity. Moreover, HCGNet also shows the remarkable interpretability and robustness by network dissection and adversarial defense, respectively. On MS-COCO, HCGNet can consistently learn better features than popular backbones. Chuanguang Yang, Zhulin An, Hui Zhu 0002, Kun Zhang 0045, Kaiqiang Xu, Chao Li 0028, Yongjun Xu 0001 |
AAAI | 1 |
| 2020 | DRNet: Dissect and Reconstruct the Convolutional Neural Network via Interpretable MannersabstractConvolutional neural networks (ConvNets) are widely used in real life. People usually use ConvNets which pre-trained on a fixed number of classes. However, for different application scenarios, we usually do not need all of the classes, which means ConvNets are redundant when dealing with these tasks. This paper focuses on the redundancy of ConvNet channels. We proposed a novel idea: using an interpretable manner to find the most important channels for every single class (dissect), and dynamically run channels according to classes in need (reconstruct). For VGG16 pre-trained on CIFAR-10, we only run 11\% parameters for two-classes sub-tasks on average with negligible accuracy loss. For VGG16 pre-trained on ImageNet, our method averagely gains 14.29\% accuracy promotion for two-classes sub-tasks. In addition, analysis show that our method captures some semantic meanings of channels, and uses the context information more targeted for sub-tasks of ConvNets. Zhulin An, Chuanguang Yang, Hui Zhu 0002, Kaiqiang Xu, Yongjun Xu 0001 |
ECAI | 3 |
| 2020 | Softer Pruning, Incremental RegularizationabstractNetwork pruning is widely used to compress Deep Neural Networks (DNNs). The Soft Filter Pruning (SFP) method zeroizes the pruned filters during training while updating them in the next training epoch. Thus the trained information of the pruned filters is completely dropped. To utilize the trained pruned filters, we proposed a SofteR Filter Pruning (SRFP) method and its variant, Asymptotic SofteR Filter Pruning (ASRFP), simply decaying the pruned weights with a monotonic decreasing parameter. Our methods perform well across various networks, datasets and pruning rates, also transferable to weight pruning. On ILSVRC-2012, ASRFP prunes 40% of the parameters on ResNet-34 with 1.63% top-1 and 0.68% top-5 accuracy improvement. In theory, SRFP and ASRFP are an incremental regularization of the pruned filters. Besides, We note that SRFP and ASRFP pursue better results while slowing down the speed of convergence. Linhang Cai, Zhulin An, Chuanguang Yang, Yongjun Xu 0001 |
ICPR | 3 |
| 2020 | Efficient Search for the Number of Channels for Convolutional Neural NetworksabstractLatest algorithms for automatic neural architecture search perform remarkably but few of them can effectively design the number of channels for convolutional neural networks and consume less computational efforts. In this paper, we propose a method for efficient automatic search which is special to the widths of networks instead of the connections within neural architectures. Our method, functionally incremental search based on function-preserving, will explore the number of channels for almost any convolutional neural network rapidly while controlling the number of parameters and even the amount of computations (FLOPs). On CIFAR-10 and CIFAR-100 classification, our method using minimal computational resources (0.41 ~ 1.29 GPU-days) can discover more effective rules of the widths of networks to improve the accuracy (a ~ 1.08 on CIFAR-10 and b ~ 2.33 on CIFAR-100) with fewer number of parameters. Hui Zhu 0002, Zhulin An, Chuanguang Yang, Kaiqiang Xu, Yongjun Xu 0001 |
IJCNN | 3 |
| 2019 | Multi-objective Pruning for CNNs Using Genetic Algorithm
Chuanguang Yang, Zhulin An, Chao Li 0028, Boyu Diao, Yongjun Xu 0001 |
ICANN (2) | 1 |