EDBT 2026 Demo / reviewers in the wild / expert
Zhulin An
dblp:85/734
· DBLP profile ↗
64ranked-venue papers
0as first author
49since 2021 · last 2026
0000-0002-7593-8293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 22 since 2021Databases, data management, data science and information retrieval · 9 · 8 since 2021Systems, architecture and hardware · 5 · 3 since 2021Computer networks · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AF-YOLO: Asymptotic Feature Extraction and Fusion for Aerial Object DetectionabstractAerial object detection plays a vital role in applications such as natural disaster prevention and urban traffic management, thanks to its ability to handle wide coverage areas and diverse objects. As a leading method for this task, You Only Look Once (YOLO) leverages multi-scale feature extraction to detect objects of various sizes. However, most YOLO-based methods focus on feature extraction and fusion from adjacent scales, neglecting the potential collaboration between non-adjacent scales. This limitation leads to redundant parameters and suboptimal detection performance. To address these issues, this paper proposes AF-YOLO (Asymptotic Feature Extraction and Fusion YOLO), a novel approach tailored for aerial object detection. AF-YOLO introduces two lightweight modules: SCC2f and PAFFN. SCC2f, an optimized version of cross-stage partial bottleneck with spatial and channel reconstruction convolution layers, reduces redundancy and enables efficient multi-scale feature extraction. PAFFN, a parallel asymptotic feature fusion network, facilitates enhanced interaction and fusion of non-adjacent scale features. Additionally, AF-YOLO incorporates a P2 layer to improve small object detection and removes YOLO’s P5 layer for a more lightweight design, specifically optimized for aerial detection tasks. Experimental results demonstrate AF-YOLO’s significant improvements across multiple benchmarks: on the VisDrone dataset, it achieves a 6.1% higher mAP0.5compared to recent baselines while using only 41.8% of their parameters; on the DIOR dataset, it shows a 3.3% accuracy improvement over YOLOv8. These quantitative results are further supported by its superior performance on the DOTA and FAIR1M datasets, with additional validation on HazyDet confirming its robustness in adverse weather conditions. Collectively, these achievements highlight AF-YOLO’s exceptional generalization capability and efficient lightweight design, establishing a new state-of-the-art for aerial object detection systems. Lve Huang, Huabiao Yan, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsabstractDiffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse. Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Renshuai Tao, Yongjun Xu 0001, Michele Magno |
AAAI | 4 |
| 2025 | HSRDiff: A Hierarchical Self-Regulation Diffusion Model for Stochastic Semantic SegmentationabstractIn safety-critical domains such as medical diagnostics and autonomous driving, single-image evidence is sometimes insufficient to reflect the inherent ambiguity of vision problems. Therefore, multiple plausible assumptions that match the image semantics may be needed to reflect the actual distribution of targets and support downstream tasks. However, balancing and improving the diversity and consistency of segmentation predictions under the high-dimensional output spaces and potential multimodal distributions is still challenging. This paper presents Hierarchical Self-Regulation Diffusion (HSRDiff), a unified framework that simulates joint probability distribution over entire labels. Our model self-regulates the balance between the two modes of predicting the label and noise in a novel ``differentiation to unification" pipeline and dynamically fits the optimal path to model the aleatoric uncertainty rooted in observations. In addition, we preserve the high-fidelity reconstruction of the delicate structure in images by leveraging the hierarchical multi-scale condition priors. We validate HSRDiff in three different semantic scenarios. Experimental results show that HSRDiff is superior to the comparison method with a considerable performance gap. Chuanguang Yang, Zhulin An, Libo Huang 0001, Yongjun Xu 0001 |
AAAI | 3 |
| 2025 | Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionabstractMulti-teacher Knowledge Distillation (KD) transfers diverse knowledge from a teacher pool to a student network. The core problem of multi-teacher KD is how to balance distillation strengths among various teachers. Most existing methods often develop weighting strategies from an individual perspective of teacher performance or teacher-student gaps, lacking comprehensive information for guidance. This paper proposes Multi-Teacher Knowledge Distillation with Reinforcement Learning (MTKD-RL) to optimize multi-teacher weights. In this framework, we construct both teacher performance and teacher-student gaps as state information to an agent. The agent outputs the teacher weight and can be updated by the return reward from the student. MTKD-RL reinforces the interaction between the student and teacher using an agent in an RL-based decision mechanism, achieving better matching capability with more meaningful weights. Experimental results on visual recognition tasks, including image classification, object detection, and semantic segmentation tasks, demonstrate that MTKD-RL achieves state-of-the-art performance compared to the existing multi-teacher KD works. Chuanguang Yang, Xinqiang Yu, Zhulin An, Chengqing Yu, Libo Huang 0001, Yongjun Xu 0001 |
AAAI | 4 |
| 2025 | A Nonlinear Hash-Based Optimization Method for SpMV on GPUsabstractSparse matrix-vector multiplication (SpMV) is a fundamental operation with a wide range of applications in scientific computing and artificial intelligence. However, the large scale and sparsity of sparse matrix often make it a performance bottleneck. In this paper, we highlight the effectiveness of hash-based techniques in optimizing sparse matrix reordering, introducing the Hash-based Partition (HBP) format, a lightweight SpMV approach. HBP retains the performance benefits of the 2Dpartitioning method while leveraging the hash transformation's ability to group similar elements, thereby accelerating the preprocessing phase of sparse matrix reordering. Additionally, we achieve parallel load balancing across matrix blocks through a competitive method. Our experiments, conducted on both Nvidia Jetson AGX Orin and Nvidia RTX 4090, show that in the preprocessing step, our method offers an average speedup of 3.53 times compared to the sorting approach and 3.67 times compared to the dynamic programming method employed in Regu2D. Furthermore, in SpMV, our method achieves a maximum speedup of 3.32 times on Orin and 3.01 times on RTX4090 against the CSR format in sparse matrices from the University of Florida Sparse Matrix Collection. Boyu Diao, Hangda Liu, Zhulin An, Yongjun Xu 0001 |
CCGrid | 4 |
| 2025 | Multi-party Collaborative Attention Control for Image CustomizationabstractThe rapid advancement of diffusion models has increased the need for customized image generation. However, current customization methods face several limitations: 1) typically accept either image or text conditions alone; 2) customization in complex visual scenarios often leads to subject leakage or confusion; 3) image-conditioned outputs tend to suffer from inconsistent backgrounds; and 4) high computational costs. To address these issues, this paper introduces Multi-party Collaborative Attention Control (MCA-Ctrl), a tuning-free method that enables high-quality image customization using both text and complex visual conditions. Specifically, MCA-Ctrl leverages two key operations within the self-attention layer to coordinate multiple parallel diffusion processes and guide the target image generation. This approach allows MCA-Ctrl to capture the content and appearance of specific subjects while maintaining semantic consistency with the conditional input. Additionally, to mitigate subject leakage and confusion issues common in complex visual scenarios, we introduce a Subject Localization Module that extracts precise subject and editable image layers based on user instructions. Extensive quantitative and human evaluation experiments show that MCA-Ctrl outperforms existing methods in zero-shot image customization, effectively resolving the mentioned issues. Chuanguang Yang, Qiuli Wang 0001, Zhulin An, Weilun Feng, Libo Huang 0001, Yongjun Xu 0001 |
CVPR | 4 |
| 2025 | IOR: Inversed Objects Replay for Incremental Object DetectionabstractExisting Incremental Object Detection (IOD) methods partially alleviate catastrophic forgetting when incrementally detecting new objects in real-world scenarios. However, many of these methods rely on the assumption that unlabeled old-class objects may co-occur with labeled new-class objects in the incremental data. When unlabeled old-class objects are absent, the performance of existing methods tends to degrade. The absence can be mitigated by generating old-class samples, but it incurs high costs. This paper argues that previous generation-based IOD suffers from redundancy, both in the use of generative models, which require additional training and storage, and in the overproduction of generated samples, many of which do not contribute significantly to performance improvements. To eliminate the redundancy, we propose Inversed Objects Replay (IOR). Specifically, we generate old-class samples by inversing the original detectors, thus eliminating the necessity of training and storing additional generative models. We propose augmented replay to reuse the objects in generated samples, reducing redundant generations. Moreover, we propose high-value knowledge distillation focusing on the positions of old-class objects overwhelmed by the background, which transfers the knowledge to the incremental detector. Extensive experiments conducted on MS COCO 2017 demonstrate that our method can efficiently improve detection performance in IOD scenarios with the absence of old-class objects. Zijia An, Boyu Diao, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
ICASSP | 5 |
| 2025 | Cross-Layer Graph Knowledge Distillation for Image RecognitionabstractKnowledge Distillation (KD) aims to improve a light-weight student network supervised by a large teacher network. The core idea of KD is to explore valuable knowledge from the teacher. Previous works often extract information from a single sample, but ignore relation modeling among multiple samples between student and teacher. Therefore, we propose Cross-Layer Graph Knowledge Distillation (CLGKD) that conducts graph-augmented feature and relation distillation assisted by graph neural networks. We further propose a meta-learning mechanism to optimize cross-layer matching weights for promoting GKD among all student and teacher layers. Experimental results on image classification and object detection demonstrate that CLGKD achieves state-of-the-art performance compared to other KD methods. Our code is available at https://github.com/cynmzzz/ICASSP2025-CLGKD Jiaming Chu, Yanzhuo Xiang, Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICASSP | 5 |
| 2025 | OLN++: Improved Object Localization Network for Open-world Object DetectionabstractOpen-world object detection (OWOD) is vital for identifying the new objects not encountered during training. Among the various methods for OWOD, Object Proposals without Learning Classification (OPwLC) stands out, with its Object Localization Network (OLN) stressing the localization features. However, OLN overlooks classification features, leading OPwLC to identify parts of a single object as multiple objects mistakenly. Inspired by the non-maximum suppression (NMS) technique, known for eliminating low-confidence detections, we sought to integrate NMS into OPwLC. However, direct integration of NMS into OPwLC presents a challenge, as OLN does not generate classification confidence scores, which are critical for applying NMS. To address this limitation, we developed a confidence measure module and proposed OLN++, filling the confidence scores gap. OLN++ can be easily implemented with just a few fully connected layers. We evaluated the effectiveness of OLN++ using NMS, Soft-NMS, and the Weighted Box Fusion variant on open-world detection tasks. Experimental results demonstrate that OLN++ significantly outperforms the original OLN. Haonan Mai, Libo Huang 0001, Zhulin An, Jiarui Zhao, Chuanguang Yang, Erhu Zhao, Yongjun Xu 0001 |
ICASSP | 3 |
| 2025 | Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingabstractSpatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation (termed SDKD), which transfers the multi-scale spatiotemporal representations from a complex teacher model to a more efficient lightweight student network. The teacher model follows an encoder-latent evolution-decoder architecture, where its latent evolution module decouples high-frequency details and low-frequency trends using convolution and Transformer (global low-frequency modeler). However, the multi-layer convolution and deconvolution structures result in slow training and high memory usage. To address these issues, we propose a frequency-aligned knowledge distillation strategy, which extracts multi-scale spectral features from the teacher's latent space, including both high and low frequency components, to guide the lightweight student model in capturing both local fine-grained variations and global evolution patterns. Experimental results show that SDKD significantly improves performance, achieving reductions of up to 81.3% in MSE and in MAE 52.3% on the Navier-Stokes equation dataset. The framework effectively captures both high-frequency variations and long-term trends while reducing computational complexity. Our codes are available at https://github.com/itsnotacie/SDKD Chuanguang Yang, Hansheng Zeng, Zeyu Dong, Zhulin An, Yongjun Xu 0001, Yingli Tian, Hao Wu 0094 |
ICCV | 5 |
| 2025 | Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersabstractDiffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters.
Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present **Q-VDiT**, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the *Token aware Quantization Estimator* (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce *Temporal Maintenance Distillation* (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency score of 23.40, setting a new benchmark and outperforming the current state-of-the-art quantization methods by **1.9$\times$**. Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Zhulin An, Libo Huang 0001, Boyu Diao, Zixiang Zhao, Yongjun Xu 0001, Michele Magno |
ICML | 6 |
| 2025 | Geometric Feature Embedding for Effective 3D Few-Shot Class Incremental Learningabstract3D few-shot class incremental learning (FSCIL) aims to learn new point cloud categories from limited samples while preventing the forgetting of previously learned categories. This research area significantly enhances the capabilities of self-driving vehicles and computer vision systems. Existing 3D FSCIL approaches primarily utilize multimodal pre-trained models to extract the semantic features, heavily dependent on meticulously designed high-quality prompts and fine-tuning strategies. To reduce this dependence, this paper proposes a novel method for **3D** **F**SCI**L** with **E**mbedded **G**eometric features (**3D-FLEG**). Specifically, 3D-FLEG develops a point cloud *geometric feature extraction module* to capture category-related geometric characteristics. To address the modality heterogeneity issues that arise from integrating geometric and text features, 3D-FLEG introduces a *geometric feature embedding module*. By augmenting text prompts with spatial geometric features through these modules, 3D-FLEG can learn robust representations of new categories even with limited samples, while mitigating forgetting of the previously learned categories. Experiments conducted on several publicly available 3D point cloud datasets, including ModelNet, ShapeNet, ScanObjectNN, and CO3D, demonstrate 3D-FLEG's superiority over existing state-of-the-art 3D FSCIL methods. Code is available at https://github.com/lixiangqi707/3D-FLEG. Xiangqi Li, Libo Huang 0001, Zhulin An, Weilun Feng, Chuanguang Yang, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ICML | 3 |
| 2025 | Classification-Based False Alarm Suppression for SAR Target DetectionabstractFalse alarm suppression is becoming increasingly important as it directly impacts the reliability and efficiency of synthetic aperture radar (SAR) image detection systems. previous methods for false alarm suppression have focused primarily on identifying the motion properties of targets and removing the embedded noise. However, SAR images are captured in a single band, which means they lack continuous bands and dynamic information. In addition, noise removal often results in a significant loss of detail in the image. In this paper, we innovatively propose a classification-based false alarm suppression framework for SAR object detection, avoiding the need for motion identification and noise removal. In practice, we first train a classification network to categorize the SAR image slices into ocean, land, and offshore scenes. Based on the classification results, we then dynamically adjust the Intersection Over Union (IoU) threshold of Non-Maximum Suppression (NMS) in different scenes. Experimental results on a newly large multi-class target SAR dataset, MSAR-1.0, show that the false alarm rate decreased from 21% to 13%. Libo Huang 0001, Zhulin An, Yongjun Xu 0001, Xia Hong 0002, Bingo Wing-Kuen Ling |
ISCAS | 3 |
| 2025 | Merlin: Multi-View Representation Learning for Robust Multivariate Time Series Forecasting with Unfixed Missing RatesabstractMultivariate Time Series Forecasting (MTSF) involves predicting future values of multiple interrelated time series. Recently, deep learning-based MTSF models have gained significant attention for their promising ability to mine semantics (global and local information) within MTS data. However, these models are pervasively susceptible to missing values caused by malfunctioning data collectors. These missing values not only disrupt the semantics of MTS, but their distribution also changes over time. Nevertheless, existing models lack robustness to such issues, leading to suboptimal forecasting performance. To this end, in this paper, we propose Multi-View Representation Learning (Merlin), which can help existing models achieve semantic alignment between incomplete observations with different missing rates and complete observations in MTS. Specifically, Merlin consists of two key modules: offline knowledge distillation and multi-view contrastive learning. The former utilizes a teacher model to guide a student model in mining semantics from incomplete observations, similar to those obtainable from complete observations. The latter improves the student model's robustness by learning from positive/negative data pairs constructed from incomplete observations with different missing rates, ensuring semantic alignment across different missing rates. Therefore, Merlin is capable of effectively enhancing the robustness of existing models against unfixed missing rates while preserving forecasting accuracy. Experiments on four real-world datasets demonstrate the superiority of Merlin. Chengqing Yu, Fei Wang 0014, Chuanguang Yang, Zezhi Shao, Tao Sun 0011, Tangwen Qian, Wei Wei 0002, Zhulin An, Yongjun Xu 0001 |
KDD (2) | 8 |
| 2025 | Accelerating Diffusion Models via Parallel Denoising
Yanming Chen 0002, Zixin Ma, Chuanguang Yang, Zhulin An, Yiwen Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | S2Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Zhulin An, Libo Huang 0001, Michele Magno, Yongjun Xu 0001 |
NeurIPS | 7 |
| 2025 | Selective Learning for Deep Time Series ForecastingabstractBenefiting from high capacity for capturing complex temporal patterns, deep learning (DL) has significantly advanced time series forecasting (TSF). However, deep models tend to suffer from severe overfitting due to the inherent vulnerability of time series to noise and anomalies. The prevailing DL paradigm uniformly optimizes all timesteps through the MSE loss and learns those uncertain and anomalous timesteps without difference, ultimately resulting in overfitting. To address this, we propose a novel selective learning strategy for deep TSF. Specifically, selective learning screens a subset of the whole timesteps to calculate the MSE loss in optimization, guiding the model to focus on generalizable timesteps while disregarding non-generalizable ones. Our framework introduces a dual-mask mechanism to target timesteps: (1) an uncertainty mask leveraging residual entropy to filter uncertain timesteps, and (2) an anomaly mask employing residual lower bound estimation to exclude anomalous timesteps. Extensive experiments across eight real-world datasets demonstrate that selective learning can significantly improve the predictive performance for typical state-of-the-art deep models, including 37.4% MSE reduction for Informer, 8.4% for TimesNet, and 6.5% for iTransformer. Yisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li 0008, Zhulin An, Cheems Wang, Yongjun Xu 0001, Fei Wang 0014 |
NeurIPS | 5 |
| 2025 | On the Integration of Spatial-Temporal Knowledge: A Lightweight Approach to Atmospheric Time Series ForecastingabstractTransformers have gained attention in atmospheric time series forecasting (ATSF) for their ability to capture global spatial-temporal correlations. However, their complex architectures lead to excessive parameter counts and extended training times, limiting their scalability to large-scale forecasting. In this paper, we revisit ATSF from a theoretical perspective of atmospheric dynamics and uncover a key insight: spatial-temporal position embedding (STPE) can inherently model spatial-temporal correlations even without attention mechanisms. Its effectiveness arises from integrating geographical coordinates and temporal features, which are intrinsically linked to atmospheric dynamics. Based on this, we propose **STELLA**, a **S**patial-**T**emporal knowledge **E**mbedded **L**ightweight mode**L** for ASTF, utilizing only STPE and an MLP architecture in place of Transformer layers. With 10k parameters and one hour of training, STELLA achieves superior performance on five datasets compared to other advanced methods. The paper emphasizes the effectiveness of spatial-temporal knowledge integration over complex architectures, providing novel insights for ATSF. Yisong Fu, Fei Wang 0014, Zezhi Shao, Boyu Diao, Lin Wu 0006, Zhulin An, Chengqing Yu, Yujie Li 0008, Yongjun Xu 0001 |
NeurIPS | 6 |
| 2025 | Low-redundancy distillation for continual learning
Boyu Diao, Libo Huang 0001, Zijia An, Hangda Liu, Zhulin An, Yongjun Xu 0001 |
Pattern Recognit. | 6 |
| 2025 | AdaE: Knowledge Graph Embedding With Adaptive Embedding SizesabstractKnowledge Graph Embedding (KGE) aims to learn dense embeddings as the representations for entities and relations in KGs. Indeed, the entities in existing KGs suffer from the data imbalance issue, i.e., there exists a substantial disparity in the occurrence frequencies among various entities. Existing KGE models pre-define a unified and fixed dimension size for all entity embeddings. However, embedding sizes of entities are highly desired for their frequencies, while a uniform embedding size may result in inadequate expression of entities, i.e., leading to overfitting for low-frequency entities and underfitting for high-frequency ones. A straight-forward idea is to set the embedding sizes for each entity before KGE training. However, manually selecting different embedding sizes is labor-intensive and time-consuming, which is difficult to achieve in real-world scenarios. To tackle this problem, we propose AdaE, which adaptively learns KG embeddings with different embedding sizes during training. In particular, AdaE is capable of selecting appropriate dimension sizes for each entity from a continuous integer space. To this end, we specially tailor bilevel optimization for the KGE task, which alternately learns representations and embedding sizes of entities. Moreover, it is worth noting that our framework is general and flexible, which is suitable for various existing KGE models. Extensive experiments demonstrate the effectiveness and compatibility of AdaE. Zhanpeng Guan, Zhao Zhang 0011, Fuzhen Zhuang, Fei Wang 0014, Zhulin An, Yongjun Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | GinAR+: A Robust End-to-End Framework for Multivariate Time Series Forecasting With Missing ValuesabstractSpatial-Temporal Graph Neural Networks (STGNNs) have been widely utilized in multivariate time series forecasting (MTSF), but they rely on the assumption of data completeness. In practice, due to factors such as natural disaster, STGNNs frequently encounter the challenge of missing data resulting from numerous malfunctioning data collectors. In this case, on the one hand, due to the presence of missing values, STGNNs easily generate incorrect spatial correlations, leading to the performance degradation. On the other hand, STGNNs require separate training of models for different missing rates, limiting their robustness. To address these challenges, we first propose two important components (interpolation attention and adaptive graph convolution), which utilize normal values to recover missing values into reliable representations and reconstruct spatial correlations. Then, we replace the fully connected layers in simple recursive units with these two components and propose Graph Interpolation Attention Recursive Network (GinAR), aiming to recursively correct spatial correlations and achieve end-to-end MTSF with missing values. Finally, we use data with different missing rates as positive and negative data pairs. By employing contrastive learning to train GinAR, we propose GinAR+ and enhance its robustness to data with different missing rates. Experiments validate the superiority of GinAR+ and our motivation. Chengqing Yu, Fei Wang 0014, Zezhi Shao, Tangwen Qian, Zhao Zhang 0011, Wei Wei 0002, Zhulin An, Qi Wang 0025, Yongjun Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | ID-centric Pre-training for RecommendationabstractClassical sequential recommendation models generally adopt ID embeddings to store knowledge learned from user historical behaviors and represent items. However, these unique IDs are challenging to be transferred to new domains. With the thriving of pre-trained language model (PLM), some pioneer works adopt PLM for pre-trained recommendation, where modality information is considered universal across domains via PLM. Unfortunately, the behavioral information in ID embeddings is verified to currently dominate in recommendation compared to modality information and thus limits these models’ performance. In this work, we propose a novel ID-centric recommendation pre-training paradigm (IDP), which directly transfers informative ID embeddings learned in pre-training domains to item representations in new domains. Specifically, in pre-training stage, besides the ID-based sequential recommendation model, we also build a Cross-domain ID-matcher (CDIM) learned by both behavioral and modality information. In the tuning stage, modality information of new domain items is regarded as a cross-domain bridge built by CDIM. They first adopted to retrieve behaviorally and semantically similar items from pre-training domains using CDIM. Next, these retrieved items’ pre-trained ID embeddings are directly adopted to generate downstream new items’ embeddings. Through extensive experiments on real-world datasets, we demonstrate that our proposed model significantly outperforms all baselines. Yiqing Wu, Ruobing Xie, Zhao Zhang 0011, Xu Zhang 0028, Fuzhen Zhuang, Leyu Lin, Zhanhui Kang, Zhulin An, Yongjun Xu 0001 |
ACM Trans. Inf. Syst. | 8 |
| 2024 | eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented GenerationabstractClass incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) to train a feature extractor incrementally, or store prototypes (features of learned tasks) to estimate the incremental feature distribution. However, the stored exemplars would violate the data privacy concerns, while the fixed prototypes might not reasonably be consistent with the incremental feature distribution, hindering the exploration of real-world CIL applications. In this paper, we propose a data-free CIL method with embedding distillation and Task-oriented generation (eTag), which requires neither exemplar nor prototype. Embedding distillation prevents the feature extractor from forgetting by distilling the outputs from the networks' intermediate blocks. Task-oriented generation enables a lightweight generator to produce dynamic features, fitting the needs of the top incremental classifier. Experimental results confirm that the proposed eTag considerably outperforms state-of-the-art methods on several benchmark datasets. Libo Huang 0001, Yan Zeng 0002, Chuanguang Yang, Zhulin An, Boyu Diao, Yongjun Xu 0001 |
AAAI | 4 |
| 2024 | Class-wise Image Mixture Guided Self-Knowledge Distillation for Image ClassificationabstractWe propose a novel regularization method to effectively train a neural network for avoiding overfitting, thus improving the performance. The core idea is to bridge the gap between predictive distributions derived from two popular image mixture techniques Mixup and CutMix by an ensemble distribution in a class-wise manner. Consistent optimization towards these three distributions is conducted by mutual distillation to guide the model to alleviate over-confidence predictions and robustly learn discriminative features as the classification evidence. Experiments across various image classification tasks show that our method significantly achieves better performance than previous data augmentation Mixup+CutMix and Self-KD methods. Zeyu Dong, Chuanguang Yang, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
CSCWD | 5 |
| 2024 | Online Relational Knowledge Distillation for Image ClassificationabstractExisting online Knowledge Distillation (KD) often perform probability-based predictions from independent data samples for knowledge transfer. However, these online KD methods neglect valuable relational information across multiple networks. To address this problem, we propose Online Relational Knowledge Distillation (ORKD). ORKD includes a discriminative loss to construct meaningful feature space and a relational distillation loss to guide structured knowledge transfer among multiple networks. Beyond feature-level distillation, we further construct an ensemble teacher by aggregating probability predictions from multiple networks. The virtual teacher is used to supervise a specific network to enhance its accuracy and avoid the cohort homogenization problem. Experimental results on CIFAR-100 and ImageNet classification demonstrate that ORKD achieves the best performance among state-of-the-art online KD methods over various network architectures. The qualitative visualization shows that ORKD can help the network to learn a more discriminative feature space, resulting in better classification performance. Yihang Zhou, Chuanguang Yang, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
CSCWD | 5 |
| 2024 | CLIP-KD: An Empirical Study of CLIP Model DistillationabstractContrastive Language-Image Pre-training (CLIP) has become a promising language-supervised visual pre-training framework. This paper aims to distill small CLIP models supervised by a large teacher CLIP model. We propose several distillation strategies, including relation, feature, gradient and contrastive paradigms, to examine the effectiveness of CLIP-Knowledge Distillation (KD). We show that a simple feature mimicry with Mean Squared Error loss works surprisingly well. Moreover, interactive contrastive learning across teacher and student encoders is also effective in performance improvement. We explain that the success of CLIP-KD can be attributed to maximizing the feature similarity between teacher and student. The unified method is applied to distill several student models trained on CC3M+12M. CLIP-KD improves student CLIP models consistently over zero-shot ImageNet classification and cross-modal retrieval bench-marks. When using ViT-U14 pretrained on Laion-400M as the teacher, CLIP-KD achieves 57.5% and 55.4% zero-shot top-1 ImageNet accuracy over ViT-B/16 and ResNet-50, surpassing the original CLIP without KD by 20.5% and 20.1% margins, respectively. Our code is released on https://github.com/winycg/CLIP-KD. Chuanguang Yang, Zhulin An, Libo Huang 0001, Junyu Bi, Xinqiang Yu, Boyu Diao, Yongjun Xu 0001 |
CVPR | 2 |
| 2024 | Online Policy Distillation with Decision-AttentionabstractPolicy Distillation (PD) has become an effective method to improve deep reinforcement learning tasks. The core idea of PD is to distill policy knowledge from a teacher agent to a student agent. However, the teacher-student framework requires a well-trained teacher model which is computationally expensive. In the light of online knowledge distillation, we study the knowledge transfer between different policies that can learn diverse knowledge from the same environment. In this work, we propose Online Policy Distillation (OPD) with Decision-Attention (DA), an online learning framework in which different policies operate in the same environment to learn different perspectives of the environment and transfer knowledge to each other to obtain better performance together. With the absence of a well-performance teacher policy, the group-derived targets play a key role in transferring group knowledge to each student policy. However, naive aggregation functions tend to cause student policies quickly homogenize. To address the challenge, we introduce the Decision-Attention module to the online policies distillation framework. The Decision-Attention module can generate a distinct set of weights for each policy to measure the importance of group members. We use the Atari platform for experiments with various reinforcement learning algorithms, including PPO and DQN. In different tasks, our method can perform better than an independent training policy on both PPO and DQN algorithms. This suggests that our OPD-DA can transfer knowledge between different policies well and help agents obtain more rewards. Xinqiang Yu, Chuanguang Yang, Chengqing Yu, Libo Huang 0001, Zhulin An, Yongjun Xu 0001 |
IJCNN | 5 |
| 2024 | Unified Dual-Intent Translation for Joint Modeling of Search and RecommendationabstractRecommendation systems, which assist users in discovering their preferred items among numerous options, have served billions of users across various online platforms. Intuitively, users' interactions with items are highly driven by their unchanging inherent intents (e.g., always preferring high-quality items) and changing demand intents (e.g., wanting a T-shirt in summer but a down jacket in winter). However, both types of intents are implicitly expressed in recommendation scenario, posing challenges in leveraging them for accurate intent-aware recommendations. Fortunately, in search scenario, often found alongside recommendation on the same online platform, users express their demand intents explicitly through their query words. Intuitively, in both scenarios, a user shares the same inherent intent and the interactions may be influenced by the same demand intent. It is therefore feasible to utilize the interaction data from both scenarios to reinforce the dual intents for joint intent-aware modeling. But the joint modeling should deal with two problems: 1) accurately modeling users' implicit demand intents in recommendation; 2) modeling the relation between the dual intents and the interactive items. To address these problems, we propose a novel model named Unified Dual-Intents Translation for joint modeling of Search and Recommendation (UDITSR). To accurately simulate users' demand intents in recommendation, we utilize real queries from search data as supervision information to guide its generation. To explicitly model the relation among the triplet , we propose a dual-intent translation propagation mechanism to learn the triplet in the same semantic space via embedding translations. Extensive experiments demonstrate that UDITSR outperforms SOTA baselines both in search and recommendation tasks. Yuting Zhang 0010, Yiqing Wu, Ruidong Han, Ying Sun 0006, Yongchun Zhu, Xiang Li 0067, Wei Lin 0022, Fuzhen Zhuang, Zhulin An, Yongjun Xu 0001 |
KDD | 9 |
| 2024 | Relational Diffusion Distillation for Efficient Image Generation
Weilun Feng, Chuanguang Yang, Zhulin An, Libo Huang 0001, Boyu Diao, Fei Wang 0014, Yongjun Xu 0001 |
ACM Multimedia | 3 |
| 2024 | Tag Tree-Guided Multi-grained Alignment for Multi-Domain Short Video Recommendation
Yuting Zhang 0010, Zhao Zhang 0011, Yiqing Wu, Ying Sun 0006, Fuzhen Zhuang, Lantao Hu, Han Li 0005, Kun Gai, Zhulin An, Yongjun Xu 0001 |
ACM Multimedia | 10 |
| 2024 | Continual Learning in the Frequency DomainabstractContinual learning (CL) is designed to learn new tasks while preserving existing knowledge. Replaying samples from earlier tasks has proven to be an effective method to mitigate the forgetting of previously acquired knowledge. However, the current research on the training efficiency of rehearsal-based methods is insufficient, which limits the practical application of CL systems in resource-limited scenarios. The human visual system (HVS) exhibits varying sensitivities to different frequency components, enabling the efficient elimination of visually redundant information. Inspired by HVS, we propose a novel framework called Continual Learning in the Frequency Domain (CLFD). To our knowledge, this is the first study to utilize frequency domain features to enhance the performance and efficiency of CL training on edge devices. For the input features of the feature extractor, CLFD employs wavelet transform to map the original input image into the frequency domain, thereby effectively reducing the size of input feature maps. Regarding the output features of the feature extractor, CLFD selectively utilizes output features for distinct classes for classification, thereby balancing the reusability and interference of output features based on the frequency domain similarity of the classes across various tasks. Optimizing only the input and output features of the feature extractor allows for seamless integration of CLFD with various rehearsal-based methods. Extensive experiments conducted in both cloud and edge environments demonstrate that CLFD consistently improves the performance of state-of-the-art (SOTA) methods in both precision and training efficiency. Specifically, CLFD can increase the accuracy of the SOTA CL method by up to 6.83% and reduce the training time by 2.6×. Boyu Diao, Libo Huang 0001, Zijia An, Zhulin An, Yongjun Xu 0001 |
NeurIPS | 5 |
| 2024 | Open-category referring expression comprehension via multi-modal knowledge transfer
Wenyu Mi, Jianji Wang 0001, Fuzhen Zhuang, Zhulin An |
Neurocomputing | 4 |
| 2024 | Multi-scale conditional reconstruction generative adversarial network
Yanming Chen 0002, Zhulin An, Fuzhen Zhuang |
Image Vis. Comput. | 3 |
| 2024 | Sketch-fusion: A gradient compression method with multi-layer fusion for communication-efficient distributed training
Li-Rong Dai 0001, Luqi Gong, Zhulin An, Yongjun Xu 0001, Boyu Diao |
J. Parallel Distributed Comput. | 3 |
| 2024 | Lung Nodule Segmentation and Uncertain Region Prediction With an Uncertainty-Aware Attention MechanismabstractRadiologists possess diverse training and clinical experiences, leading to variations in the segmentation annotations of lung nodules and resulting in segmentation uncertainty. Conventional methods typically select a single annotation as the learning target or attempt to learn a latent space comprising multiple annotations. However, these approaches fail to leverage the valuable information inherent in the consensus and disagreements among the multiple annotations. In this paper, we propose an Uncertainty-Aware Attention Mechanism (UAAM) that utilizes consensus and disagreements among multiple annotations to facilitate better segmentation. To this end, we introduce the Multi-Confidence Mask (MCM), which combines a Low-Confidence (LC) Mask and a High-Confidence (HC) Mask. The LC mask indicates regions with low segmentation confidence, where radiologists may have different segmentation choices. Following UAAM, we further design an Uncertainty-Guide Multi-Confidence Segmentation Network (UGMCS-Net), which contains three modules: a Feature Extracting Module that captures a general feature of a lung nodule, an Uncertainty-Aware Module that produces three features for the annotations' union, intersection, and annotation set, and an Intersection-Union Constraining Module that uses distances between the three features to balance the predictions of final segmentation and MCM. To comprehensively demonstrate the performance of our method, we propose a Complex-Nodule Validation on LIDC-IDRI, which tests UGMCS-Net's segmentation performance on lung nodules that are difficult to segment using common methods. Experimental results demonstrate that our method can significantly improve the segmentation performance on nodules that are difficult to segment using conventional methods. Qiuli Wang 0001, Yue Zhang 0042, Zhulin An, Chen Liu 0026, Xiaohong Zhang 0002, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Knowledge Distillation Using Hierarchical Self-Supervision Augmented DistributionabstractKnowledge distillation (KD) is an effective framework that aims to transfer meaningful information from a large teacher to a smaller student. Generally, KD often involves how to define and transfer knowledge. Previous KD methods often focus on mining various forms of knowledge, for example, feature maps and refined information. However, the knowledge is derived from the primary supervised task, and thus, is highly task-specific. Motivated by the recent success of self-supervised representation learning, we propose an auxiliary self-supervision augmented task to guide networks to learn more meaningful features. Therefore, we can derive soft self-supervision augmented distributions as richer dark knowledge from this task for KD. Unlike previous knowledge, this distribution encodes joint knowledge from supervised and self-supervised feature learning. Beyond knowledge exploration, we propose to append several auxiliary branches at various hidden layers, to fully take advantage of hierarchical feature maps. Each auxiliary branch is guided to learn self-supervision augmented tasks and distill this distribution from teacher to student. Overall, we call our KD method a hierarchical self-supervision augmented KD (HSSAKD). Experiments on standard image classification show that both offline and online HSSAKD achieves state-of-the-art performance in the field of KD. Further transfer experiments on object detection further verify that HSSAKD can guide the network to learn better features. The code is available at https://github.com/winycg/HSAKD. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Triple Dual Learning for Opinion-based Explainable RecommendationabstractRecently, with the aim of enhancing the trustworthiness of recommender systems, explainable recommendation has attracted much attention from the research community. Intuitively, users’ opinions toward different aspects of an item determine their ratings (i.e., users’ preferences) for the item. Therefore, rating prediction from the perspective of opinions can realize personalized explanations at the level of item aspects and user preferences. However, there are several challenges in developing an opinion-based explainable recommendation: (1) The complicated relationship between users’ opinions and ratings. (2) The difficulty of predicting the potential (i.e., unseen) user-item opinions because of the sparsity of opinion information. To tackle these challenges, we propose an overall preference-aware opinion-based explainable rating prediction model by jointly modeling the multiple observations of user-item interaction (i.e., review, opinion, rating). To alleviate the sparsity problem and raise the effectiveness of opinion prediction, we further propose a triple dual learning-based framework with a novelly designed triple dual constraint . Finally, experiments on three popular datasets show the effectiveness and great explanation performance of our framework. Yuting Zhang 0010, Ying Sun 0006, Fuzhen Zhuang, Yongchun Zhu, Zhulin An, Yongjun Xu 0001 |
ACM Trans. Inf. Syst. | 5 |
| 2023 | Modeling Dual Period-Varying Preferences for Takeaway RecommendationabstractTakeaway recommender systems, which aim to accurately provide stores that offer foods meeting users' interests, have served billions of users in our daily life. Different from traditional recommendation, takeaway recommendation faces two main challenges: (1) Dual Interaction-Aware Preference Modeling. Traditional recommendation commonly focuses on users' single preferences for items while takeaway recommendation needs to comprehensively consider users' dual preferences for stores and foods. (2) Period-Varying Preference Modeling. Conventional recommendation generally models continuous changes in users' preferences from a session-level or day-level perspective. However, in practical takeaway systems, users' preferences vary significantly during the morning, noon, night, and late night periods of the day. To address these challenges, we propose a Dual Period-Varying Preference modeling (DPVP) for takeaway recommendation. Specifically, we design a dual interaction-aware module, aiming to capture users' dual preferences based on their interactions with stores and foods. Moreover, to model various preferences in different time periods of the day, we propose a time-based decomposition module as well as a time-aware gating mechanism. Extensive offline and online experiments demonstrate that our model outperforms state-of-the-art methods on real-world datasets and it is capable of modeling the dual period-varying preferences. Moreover, our model has been deployed online on Meituan Takeaway platform, leading to an average improvement in GMV (Gross Merchandise Value) of 0.70%. Yuting Zhang 0010, Yiqing Wu, Ran Le, Yongchun Zhu, Fuzhen Zhuang, Ruidong Han, Xiang Li 0067, Wei Lin 0022, Zhulin An, Yongjun Xu 0001 |
KDD | 9 |
| 2023 | Weighted Knowledge Graph EmbeddingabstractKnowledge graph embedding (KGE) aims to project both entities and relations in a knowledge graph (KG) into low-dimensional vectors. Indeed, existing KGs suffer from the data imbalance issue, i.e., entities and relations conform to a long-tail distribution, only a small portion of entities and relations occur frequently, while the vast majority of entities and relations only have a few training samples. Existing KGE methods assign equal weights to each entity and relation during the training process. Under this setting, long-tail entities and relations are not fully trained during training, leading to unreliable representations. In this paper, we propose WeightE, which attends differentially to different entities and relations. Specifically, WeightE is able to endow lower weights to frequent entities and relations, and higher weights to infrequent ones. In such manner, WeightE is capable of increasing the weights of long-tail entities and relations, and learning better representations for them. In particular, WeightE tailors bilevel optimization for the KGE task, where the inner level aims to learn reliable entity and relation embeddings, and the outer level attempts to assign appropriate weights for each entity and relation. Moreover, it is worth noting that our technique of applying weights to different entities and relations is general and flexible, which can be applied to a number of existing KGE models. Finally, we extensively validate the superiority of WeightE against various state-of-the-art baselines. Zhao Zhang 0011, Zhanpeng Guan, Fuzhen Zhuang, Zhulin An, Fei Wang 0014, Yongjun Xu 0001 |
SIGIR | 5 |
| 2023 | Online Knowledge Distillation via Mutual Contrastive Learning for Visual RecognitionabstractThe teacher-free online Knowledge Distillation (KD) aims to train an ensemble of multiple student models collaboratively and distill knowledge from each other. Although existing online KD methods achieve desirable performance, they often focus on class probabilities as the core knowledge type, ignoring the valuable feature representational information. We present a Mutual Contrastive Learning (MCL) framework for online KD. The core idea of MCL is to perform mutual interaction and transfer of contrastive distributions among a cohort of networks in an online manner. Our MCL can aggregate cross-network embedding information and maximize the lower bound to the mutual information between two networks. This enables each network to learn extra contrastive knowledge from others, leading to better feature representations, thus improving the performance of visual recognition tasks. Beyond the final layer, we extend MCL to intermediate layers and perform an adaptive layer-matching mechanism trained by meta-optimization. Experiments on image classification and transfer learning to visual recognition tasks show that layer-wise MCL can lead to consistent performance gains against state-of-the-art online KD approaches. The superiority demonstrates that layer-wise MCL can guide the network to generate better feature representations. Our code is publicly avaliable at https://github.com/winycg/L-MCL. Chuanguang Yang, Zhulin An, Helong Zhou, Fuzhen Zhuang, Yongjun Xu 0001, Qian Zhang 0009 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Prior Gradient Mask Guided Pruning-Aware Fine-TuningabstractWe proposed a Prior Gradient Mask Guided Pruning-aware Fine-Tuning (PGMPF) framework to accelerate deep Convolutional Neural Networks (CNNs). In detail, the proposed PGMPF selectively suppresses the gradient of those ”unimportant” parameters via a prior gradient mask generated by the pruning criterion during fine-tuning. PGMPF has three charming characteristics over previous works: (1) Pruning-aware network fine-tuning. A typical pruning pipeline consists of training, pruning and fine-tuning, which are relatively independent, while PGMPF utilizes a variant of the pruning mask as a prior gradient mask to guide fine-tuning, without complicated pruning criteria. (2) An excellent tradeoff between large model capacity during fine-tuning and stable convergence speed to obtain the final compact model. Previous works preserve more training information of pruned parameters during fine-tuning to pursue better performance, which would incur catastrophic non-convergence of the pruned model for relatively large pruning rates, while our PGMPF greatly stabilizes the fine-tuning phase by gradually constraining the learning rate of those ”unimportant” parameters. (3) Channel-wise random dropout of the prior gradient mask to impose some gradient noise to fine-tuning to further improve the robustness of final compact model. Experimental results on three image classification benchmarks CIFAR10/ 100 and ILSVRC-2012 demonstrate the effectiveness of our method for various CNN architectures, datasets and pruning rates. Notably, on ILSVRC-2012, PGMPF reduces 53.5% FLOPs on ResNet-50 with only 0.90% top-1 accuracy drop and 0.52% top-5 accuracy drop, which has advanced the state-of-the-art with negligible extra computational cost. Linhang Cai, Zhulin An, Chuanguang Yang, Yangchun Yan, Yongjun Xu 0001 |
AAAI | 2 |
| 2022 | Mutual Contrastive Learning for Visual Representation LearningabstractWe present a collaborative learning method called Mutual Contrastive Learning (MCL) for general visual representation learning. The core idea of MCL is to perform mutual interaction and transfer of contrastive distributions among a cohort of networks. A crucial component of MCL is Interactive Contrastive Learning (ICL). Compared with vanilla contrastive learning, ICL can aggregate cross-network embedding information and maximize the lower bound to the mutual information between two networks. This enables each network to learn extra contrastive knowledge from others, leading to better feature representations for visual recognition tasks. We emphasize that the resulting MCL is conceptually simple yet empirically powerful. It is a generic framework that can be applied to both supervised and self-supervised representation learning. Experimental results on image classification and transfer learning to object detection show that MCL can lead to consistent performance gains, demonstrating that MCL can guide the network to generate better feature representations. Code is available at https://github.com/winycg/MCL. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
AAAI | 2 |
| 2022 | Cross-Image Relational Knowledge Distillation for Semantic SegmentationabstractCurrent Knowledge Distillation (KD) methods for semantic segmentation often guide the student to mimic the teacher's structured information generated from individual data samples. However, they ignore the global semantic relations among pixels across various images that are valuable for KD. This paper proposes a novel Cross-Image Relational KD (CIRKD), which focuses on transferring structured pixel-to-pixel and pixel-to-region relations among the whole images. The motivation is that a good teacher network could construct a well-structured feature space in terms of global pixel dependencies. CIRKD makes the student mimic better structured semantic relations from the teacher, thus improving the segmentation performance. Experimental results over Cityscapes, CamVid and Pascal VOC datasets demonstrate the effectiveness of our proposed approach against state-of-the-art distillation methods. The code is available at https://github.com/winycg/CIRKD. Chuanguang Yang, Helong Zhou, Zhulin An, Yongjun Xu 0001, Qian Zhang 0009 |
CVPR | 3 |
| 2022 | MixSKD: Self-Knowledge Distillation from Mixup for Image Recognition
Chuanguang Yang, Zhulin An, Helong Zhou, Linhang Cai, Xiang Zhi, Jiwen Wu, Yongjun Xu 0001, Qian Zhang 0009 |
ECCV (24) | 2 |
| 2022 | FPAR: Filter Pruning Via Attention and Rank EnhancementabstractIn recent years, deep convolutional neural networks (CNNs) have become larger than ever, thus their deployment on edge devices becomes difficult. There are numerous popular methods to accelerate the networks; however, many of these methods consider only the importance of a single filter to the network and neglect the coorelation between filters. To solve this problem, we propose a novel filter pruning method, called Filter Pruning via Attention and Rank Enhancement (FPAR), based on the attention mechanism and rank of feature maps. Moreover, the inspiration for it comes from a discovery: For a network with attention modules, irrespective of the batch of input images, the mean of channel-wise weights of the attention module is almost constant. Thus, we can use a few batches of input data to obtain this indicator to guide pruning. With extensive experiments on various datasets, demonstrate that our method outperforms the most advanced methods with similar accuracy. For example, using VGG-16, we removed 62.8% of floating-point operations (FLOPs) even with a 0.24% of the accuracy increase compared with the unpruned network. Yanming Chen 0002, Mingrui Shuai, Shubin Lou, Zhulin An, Yiwen Zhang 0001 |
ICME | 4 |
| 2022 | Localizing Semantic Patches for Accelerating Image ClassificationabstractExisting works often focus on reducing the architecture redundancy for accelerating image classification but ignore the spatial redundancy of the input image. This paper proposes an efficient image classification pipeline to solve this problem. We first pinpoint task-aware regions over the input image by a lightweight patch proposal network called AnchorNet. We then feed these localized semantic patches with much smaller spatial redundancy into a general classification network. Unlike the popular design of deep CNN, we aim to carefully design the Receptive Field of AnchorNet without intermediate convolutional paddings. This ensures the exact mapping from a high-level spatial location to the specific input image patch. The contribution of each patch is interpretable. Moreover, AnchorNet is compatible with any downstream architecture. Experimental results on ImageNet show that our method outperforms SOTA dynamic inference methods with fewer inference costs. Our code is available at https://github.com/winycg/AnchorNet. Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICME | 2 |
| 2021 | Multi-View Contrastive Learning for Online Knowledge DistillationabstractPrevious Online Knowledge Distillation (OKD) often carries out mutually exchanging probability distributions, but neglects the useful representational knowledge. We there-fore propose Multi-view Contrastive Learning (MCL) for OKD to implicitly capture correlations of feature embeddings encoded by multiple peer networks, which provide various views for understanding the input data instances. Benefiting from MCL, we can learn a more discriminative representation space for classification than previous OKD methods. Experimental results on image classification demonstrate that our MCL-OKD outperforms other state-of-the-art OKD methods by large margins without sacrificing additional inference cost. Codes are available at https://github.com/winycg/MCL-OKD. Chuanguang Yang, Zhulin An, Yongjun Xu 0001 |
ICASSP | 2 |
| 2021 | Hierarchical Self-supervised Augmented Knowledge DistillationabstractKnowledge distillation often involves how to define and transfer knowledge from teacher to student effectively. Although recent self-supervised contrastive knowledge achieves the best performance, forcing the network to learn such knowledge may damage the representation learning of the original class recognition task. We therefore adopt an alternative self-supervised augmented task to guide the network to learn the joint distribution of the original recognition task and self-supervised auxiliary task. It is demonstrated as a richer knowledge to improve the representation power without losing the normal classification capability. Moreover, it is incomplete that previous methods only transfer the probabilistic knowledge between the final layers. We propose to append several auxiliary classifiers to hierarchical intermediate feature maps to generate diverse self-supervised knowledge and perform the one-to-one transfer to teach the student network thoroughly. Our method significantly surpasses the previous SOTA SSKD with an average improvement of 2.56% on CIFAR-100 and an improvement of 0.77% on ImageNet across widely used network pairs. Codes are available at https://github.com/winycg/HSAKD. Chuanguang Yang, Zhulin An, Linhang Cai, Yongjun Xu 0001 |
IJCAI | 2 |
| 2021 | Soft and Hard Filter Pruning via Dimension ReductionabstractFilter pruning is widely used to reduce the computation of deep learning, enabling the deployment of Deep Neural Networks (DNNs) in resource-limited devices. Conventional Hard Filter Pruning (HFP) method zeroizes pruned filters and stops updating them, thus reducing the search space of the model. On the contrary, Soft Filter Pruning (SFP) simply zeroizes pruned filters, keeping updating them in the following training epochs, thus maintaining the capacity of the network. However, SFP, together with its variants, converges much slower than HFP due to its larger search space. Firstly, we generalize SFP-based methods and HFP to analyze their characteristics. Then we propose a Gradually Hard Filter Pruning (GHFP) method to smoothly switch from SFP-based methods to HFP during training and pruning, thus maintaining a large search space at first, gradually reducing the capacity of the model to ensure a moderate convergence speed. Furthermore, we view filter pruning as dimension reduction and propose a novel dimension reduction block integrated into GHFP to significantly outperform other methods by a moderate margin. Linhang Cai, Zhulin An, Chuanguang Yang, Yongjun Xu 0001 |
IJCNN | 2 |
| 2020 | Gated Convolutional Networks with Hybrid Connectivity for Image ClassificationabstractWe propose a simple yet effective method to reduce the redundancy of DenseNet by substantially decreasing the number of stacked modules by replacing the original bottleneck by our SMG module, which is augmented by local residual. Furthermore, SMG module is equipped with an efficient two-stage pipeline, which aims to DenseNet-like architectures that need to integrate all previous outputs, i.e., squeezing the incoming informative but redundant features gradually by hierarchical convolutions as a hourglass shape and then exciting it by multi-kernel depthwise convolutions, the output of which would be compact and hold more informative multi-scale features. We further develop a forget and an update gate by introducing the popular attention modules to implement the effective fusion instead of a simple addition between reused and new features. Due to the Hybrid Connectivity (nested combination of global dense and local residual) and Gated mechanisms, we called our network as the HCGNet. Experimental results on CIFAR and ImageNet datasets show that HCGNet is more prominently efficient than DenseNet, and can also significantly outperform state-of-the-art networks with less complexity. Moreover, HCGNet also shows the remarkable interpretability and robustness by network dissection and adversarial defense, respectively. On MS-COCO, HCGNet can consistently learn better features than popular backbones. Chuanguang Yang, Zhulin An, Hui Zhu 0002, Kun Zhang 0045, Kaiqiang Xu, Chao Li 0028, Yongjun Xu 0001 |
AAAI | 2 |
| 2020 | DRNet: Dissect and Reconstruct the Convolutional Neural Network via Interpretable MannersabstractConvolutional neural networks (ConvNets) are widely used in real life. People usually use ConvNets which pre-trained on a fixed number of classes. However, for different application scenarios, we usually do not need all of the classes, which means ConvNets are redundant when dealing with these tasks. This paper focuses on the redundancy of ConvNet channels. We proposed a novel idea: using an interpretable manner to find the most important channels for every single class (dissect), and dynamically run channels according to classes in need (reconstruct). For VGG16 pre-trained on CIFAR-10, we only run 11\% parameters for two-classes sub-tasks on average with negligible accuracy loss. For VGG16 pre-trained on ImageNet, our method averagely gains 14.29\% accuracy promotion for two-classes sub-tasks. In addition, analysis show that our method captures some semantic meanings of channels, and uses the context information more targeted for sub-tasks of ConvNets. Zhulin An, Chuanguang Yang, Hui Zhu 0002, Kaiqiang Xu, Yongjun Xu 0001 |
ECAI | 2 |
| 2020 | Towards More Efficient And Effective Inference: The Joint Decision Of Multi-ParticipantsabstractExisting approaches to improve the performances of convolutional neural networks by optimizing the local architectures or deepening the networks tend to increase the size of models significantly. In order to deploy and apply the neural networks to edge devices which are in great demand, reducing the scale of networks is quite crucial. However, It is easy to degrade the performance of image processing by compressing the networks. In this paper, we propose a method which is suitable for edge devices while improving the efficiency and effectiveness of inference. The joint decision of multiparticipants, mainly contain multi-layers and multi-networks, can achieve higher classification accuracy (0.26% on CFAR-10 and 4.49% on CFAR-100 at most) with similar total number of parameters for classical convolutional neural networks. Hui Zhu 0002, Zhulin An, Kaiqiang Xu, Yongjun Xu 0001 |
ICIP | 2 |
| 2020 | Softer Pruning, Incremental RegularizationabstractNetwork pruning is widely used to compress Deep Neural Networks (DNNs). The Soft Filter Pruning (SFP) method zeroizes the pruned filters during training while updating them in the next training epoch. Thus the trained information of the pruned filters is completely dropped. To utilize the trained pruned filters, we proposed a SofteR Filter Pruning (SRFP) method and its variant, Asymptotic SofteR Filter Pruning (ASRFP), simply decaying the pruned weights with a monotonic decreasing parameter. Our methods perform well across various networks, datasets and pruning rates, also transferable to weight pruning. On ILSVRC-2012, ASRFP prunes 40% of the parameters on ResNet-34 with 1.63% top-1 and 0.68% top-5 accuracy improvement. In theory, SRFP and ASRFP are an incremental regularization of the pruned filters. Besides, We note that SRFP and ASRFP pursue better results while slowing down the speed of convergence. Linhang Cai, Zhulin An, Chuanguang Yang, Yongjun Xu 0001 |
ICPR | 2 |
| 2020 | HLNet: Modeling High and Low Frequencies for Scene ParsingabstractIn this paper we propose to model high and low frequencies of segmentation map, based on the observation that the map can be seen as a mixture of different frequencies. Based on the sparsity of high frequencies and local similarity of low frequencies, we design special building blocks and further a novel High and Low frequency Network (HLNet) with two branches based on FCN to predict high and low frequencies of the segmentation map, respectively. Specifically, we design a high frequency branch with a small kernel size and high-resolution features to predict a sparse high frequency component. Mean-while, a low frequency branch with similarity computing and low-resolution features is employed to predict a low frequency component. On top of two branches, we combine two different frequency components to generate final result for scene parsing. We empirically demonstrate that the designed model achieves superior performance 44.07% on ADE20K, and 80.14% mIoU on Cityscapes datasets. Kaiqiang Xu, Zhulin An, Hui Zhu 0002, Yongjun Xu 0001 |
IJCNN | 2 |
| 2020 | Efficient Search for the Number of Channels for Convolutional Neural NetworksabstractLatest algorithms for automatic neural architecture search perform remarkably but few of them can effectively design the number of channels for convolutional neural networks and consume less computational efforts. In this paper, we propose a method for efficient automatic search which is special to the widths of networks instead of the connections within neural architectures. Our method, functionally incremental search based on function-preserving, will explore the number of channels for almost any convolutional neural network rapidly while controlling the number of parameters and even the amount of computations (FLOPs). On CIFAR-10 and CIFAR-100 classification, our method using minimal computational resources (0.41 ~ 1.29 GPU-days) can discover more effective rules of the widths of networks to improve the accuracy (a ~ 1.08 on CIFAR-10 and b ~ 2.33 on CIFAR-100) with fewer number of parameters. Hui Zhu 0002, Zhulin An, Chuanguang Yang, Kaiqiang Xu, Yongjun Xu 0001 |
IJCNN | 2 |
| 2019 | Multi-objective Pruning for CNNs Using Genetic Algorithm
Chuanguang Yang, Zhulin An, Chao Li 0028, Boyu Diao, Yongjun Xu 0001 |
ICANN (2) | 2 |
| 2017 | TDMA Versus CSMA/CA for Wireless Multihop Communications: A Stochastic Worst-Case Delay AnalysisabstractWireless networks have become a very attractive solution for soft real-time data transport in the industry. For such technologies to carry real-time traffic, reliable bounds on end-to-end communication delays have to be ascertained to warrant a proper system behavior. As for legacy wired embedded and real-time networks, two main wireless multiple access methods can be leveraged: one is time division multiple access (TDMA), which follows a time-triggered paradigm, and the other is carrier sense multiple access with collision avoidance (CSMA/CA), which follows an event-triggered paradigm. This paper proposes an analytical comparison of the time behavior of two representative TDMA and CSMA/CA protocols in terms of the worst-case end-to-end delay. This worst-case delay is expressed in a probabilistic manner because our analytical framework captures the versatility of the wireless medium. Analytical delay bounds are obtained from delay distributions, which are compared to fine-grained simulation results. Exhibited study cases show that TDMA can offer smaller or larger worst-case bounds than CSMA/CA depending on its settings. Qi Wang 0025, Katia Jaffrès-Runser, Yongjun Xu 0001, Jean-Luc Scharbarg, Zhulin An, Christian Fraboul |
IEEE Trans. Ind. Informatics | 5 |
| 2016 | A Reliable Depth-Based Routing Protocol with Network Coding for Underwater Sensor NetworksabstractWith the rapid development of marine technology, underwater sensor networks (UWSNs) are gradually evolving from research to practice in recent years. Practicability and reliability are two major concerns for routing protocols in UWSNs. As localization is not necessary in depth-based routing protocol (DBR), it has an outstanding practicability than other geographic routing protocols. However, the reliability is not well ensured. In this paper, we propose an innovative depth-based routing with network coding improving routing reliability while preserving the intrinsic distributed manner of DBR and introducing little time delay and energy cost. Moreover, a simple analytical performance model where ideal MAC is assumed is proposed to derive the analytical delivery ratio for our DBR-NC and DBR protocols. This analytical model is validated by simulation results. The extensive simulation results show that the proposed DBR-NC protocol outperforms (over 15%) the state of art DBR protocols in terms of packet delivery ratio. We also show that our DBR-NC will not introduce much extra delay and energy consumptions. Boyu Diao, Yongjun Xu 0001, Qi Wang 0025, Zhao Chen 0007, Chao Li 0028, Zhulin An, Guangjie Han |
ICPADS | 6 |
| 2016 | TDMA versus CSMA/CA for wireless multi-hop communications: A comparison for soft real-time networkingabstractWireless networks have become a very attractive solution for soft real-time data transport in the industry. For such technologies to carry real-time traffic, reliable bounds on end-to-end communication delays have to be ascertained to warrant a proper system behavior. As for legacy wired embedded and real-time networks, two main wireless multiple access methods can be leveraged: (i) time division multiple access (TDMA), which follows a time-triggered paradigm and (ii) Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA), which follows an event-triggered paradigm. This paper proposes an analytical comparison of the time behavior of two representative TDMA and CSMA/CA protocols in terms of worst-case end-to-end delay. This worst-case delay is expressed in a probabilistic manner because our analytical framework captures the versatility of the wireless medium. Analytical delay bounds are obtained from delay distributions, which are compared to fine-grained simulation results. Exhibited study cases show that TDMA can offer smaller or larger worst-case bounds than CSMA/CA depending on its settings. Qi Wang 0025, Katia Jaffrès-Runser, Yongjun Xu 0001, Jean-Luc Scharbarg, Zhulin An, Christian Fraboul |
WFCS | 5 |
| 2016 | Optimal scheduling for energy harvesting mobile sensing devices
Weiwei Fang, Zhulin An, Qiang Liu 0014 |
Comput. Commun. | 5 |
| 2015 | Vessel trajectory partitioning based on hierarchical fusion of position data
Xianbin Wu, Lin Wu 0006, Yongjun Xu 0001, Zhulin An, Boyu Diao |
FUSION | 4 |
| 2014 | Achieving optimal admission control with dynamic scheduling in energy constrained network systems
Weiwei Fang, Zhulin An, Lei Shu 0001, Yongjun Xu 0001 |
J. Netw. Comput. Appl. | 2 |
| 2013 | Poster abstract: human tracking based on LRF and wearable IMU data fusionabstractHuman tracking is one of the most important requirements for service mobile robots. Cameras and Laser Ranger Finders (LRFs) are usually used together for human tracking. But these kinds of solutions are too computationally expensive for most embedded processors on these robots as complex computer vision algorithms are needed to process large number of pixels. In this paper, we describe a method combining kinematic measurements from LRF mounted on the robot and Inertial Measurement Unit (IMU) carried by the target. These two types of sensors can calculate human's velocity and position independently, which are used as information for both indentifying and tracking the target. As pixels observed by LRF and IMU are 1D rather than 2D, our method requires much less computation and memory resources and can be implemented with low-performance embedded processors. Lin Wu 0006, Zhulin An, Yongjun Xu 0001 |
IPSN | 2 |
| 2013 | Odometer in the pocketabstractSome previous work has shown the feasibility of Pedestrian Dead Reckoning (PDR) using a mobile phone, but the estimation of length walked is still a big challenge. In this paper, we propose a formula for estimating walking velocity, which can be integrated to get distance. Lin Wu 0006, Yongjun Xu 0001, Zhulin An, Chaonong Xu, Fei Wang 0014 |
MobiSys | 3 |