Chengchao Shen

dblp:217/3551 · DBLP profile ↗
← Back
24ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-0249-764XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 MCAF: Improving Mortality Risk Prediction Using Multimodal Learning with Balanced Pre-training and Correlation-Aware Fusion
Abdulrahman Al-badwi, Chengchao Shen, Abdulrahman Al-Dailami, Raeed Alsabri, Hulin Kuang, Jianxin Wang 0001
ISBRA (1)2
2026 SecPredNet: Joint Modeling of PLM Embeddings and Physicochemical Features for Signal Peptide Secretion Efficiency Prediction
Chengchao Shen, Jianxin Wang 0001
ISBRA (1)3
2026 Inter-instance similarity modeling for contrastive learning
Chengchao Shen
Pattern Recognit.1
2026 Multiple object stitching for unsupervised representation learning
Chengchao Shen, Jianxin Wang 0001
Pattern Recognit.1
2025 A unified Personalized Federated Learning framework ensuring Domain Generalization
Yuan Liu 0038, Chengchao Shen, Yixiong Liang, Jianxin Wang 0001
Expert Syst. Appl.4
2025 Multi-grained contrast for data-efficient unsupervised representation learning
Chengchao Shen
Pattern Recognit.1
2025 Asymmetric patch sampling for contrastive learning
Chengchao Shen, Hulin Kuang, Jin Liu 0012, Jianxin Wang 0001
Pattern Recognit.1
2025 Data-efficient multi-scale fusion vision transformer
Chengchao Shen
Pattern Recognit.3
2023 Improving Medical Image Denoising via a Lightweight Plug-and-play Module
abstract
Medical image denoising, as a part of medical image processing, is significant for the assessment and diagnosis of diseases. To improve the medical image denoising performance of existing deep learning methods, we propose a lightweight plug-and-play module (LP2M) with low complexity, which can be plugged before current image denoising methods. Specifically, the proposed LP2M contains three stacked Convolutional Neural Network (CNN) based blocks: a image receptor block, an adaptive receptive field selection block and a high-low frequency processing block. The image receptor block can perceive color or grayscale images and perform preliminary processing. The adaptive receptive field selection block includes two parallel paths with different receptive fields (i.e., convolution kernel sizes) and the adaptive weighting operation, which can process multi-scale information in the image. The high-low frequency processing block consists of a low frequency pathway using convolutional layers with large kernel sizes, and a high frequency pathway using covolutional layers with small kernel sizes, which can process the low and high frequency components in images. Extensive validation experiments are performed on five state-of-the-art denoising methods on multiple medical image datasets for three different medical image denoising tasks (X-ray image denoising, magnetic resonance image denoising and dermoscopic image denoising). Experimental results show that our proposed LP2M can effectively improve the results of these five state-of-the-art methods for three denoising tasks with only increasing 0.996K parameters and 63.112M FLOPs, and it can provide potential direction for improving image denoising.
Hulin Kuang, Jin Liu 0012, Chengchao Shen, Jianxin Wang 0001
BIBM4
2023 How to Prevent the Poor Performance Clients for Personalized Federated Learning?
abstract
Personalized federated learning (pFL) collaboratively trains personalized models, which provides a customized model solution for individual clients in the presence of heterogeneous distributed local data. Although many recent studies have applied various algorithms to enhance personalization in pFL, they mainly focus on improving the performance from averaging or top perspective. However, part of the clients may fall into poor performance and are not clearly discussed. Therefore, how to prevent these poor clients should be considered critically. Intuitively, these poor clients may come from biased universal information shared with others. To address this issue, we propose a novel pFL strategy, called Personalize Locally, Generalize Universally (PLGU). PLGU generalizes the fine-grained universal information and moderates its biased performance by designing a Layer-Wised Sharpness Aware Minimization (LWSAM) algorithm while keeping the personalization local. Specifically, we embed our proposed PLGU strategy into two pFL schemes concluded in this paper: with/without a global model, and present the training procedures in detail. Through in-depth study, we show that the proposed PLGU strategy achieves competitive generalization bounds on both considered pFL schemes. Our extensive experimental results show that all the proposed PLGU based-algorithms achieve state-of-the-art performance.
Chengchao Shen, Lixing Chen
CVPR5
2023 Modeling global distribution for federated learning with label distribution skew
Chengchao Shen, Yuan Liu 0038, Yeyu Ou, Yixiong Liang, Jianxin Wang 0001
Pattern Recognit.2
2022 MVSF: Multi-View Signature Fusion Network for Noninvasively Predicting Ki67 Status
abstract
Ki67 is a promising molecular biomarker for the diagnosis of lung adenocarcinoma. However, previous methods to determine Ki67 status often require tumor tissue sampling, which is invasive for patients. This study proposes a multi-view signature fusion network (MVSF), combining deep learning encoded (DLE) signatures, handcrafted radiomics (HCR) signatures, and clinical information to noninvasively predict Ki67 status. Multi-view signatures are combined through a tensor fusion network to obtain potentially high-dimensional signatures. Finally, a cooperative game theory-based approach is applied to quantitatively interpret the contribution of signatures to decision-making. The proposed MVSF is evaluated on a retrospectively collected dataset of 661 patients. Experimental results show that the MVSF achieves encouraging performance, with an area under the receiver operating characteristic curve of 0.80 and an accuracy of 0.78, outperforming several state-of-the-art Ki67 status prediction methods, which implies that our proposed method could provide potential support for Ki67 status prediction.
Jianhong Cheng, Jin Liu 0012, Hulin Kuang, Chengchao Shen, Jianxin Wang 0001
BIBM5
2021 Progressive Network Grafting for Few-Shot Knowledge Distillation
abstract
Knowledge distillation has demonstrated encouraging performances in deep model compression. Most existing approaches, however, require massive labeled data to accomplish the knowledge transfer, making the model compression a cumbersome and costly process. In this paper, we investigate the practical few-shot knowledge distillation scenario, where we assume only a few samples without human annotations are available for each category. To this end, we introduce a principled dual-stage distillation scheme tailored for few-shot data. In the first step, we graft the student blocks one by one onto the teacher, and learn the parameters of the grafted block intertwined with those of the other teacher blocks. In the second step, the trained student blocks are progressively connected and then together grafted onto the teacher network, allowing the learned student blocks to adapt themselves to each other and eventually replace the teacher network. Experiments demonstrate that our approach, with only a few unlabeled samples, achieves gratifying results on CIFAR10,CIFAR100, and ILSVRC-2012. On CIFAR10 and CIFAR100, our performances are even on par with those of knowledge distillation schemes that utilize the full datasets. The source code is available at https://github.com/zju-vipa/NetGraft.
Chengchao Shen, Xinchao Wang, Youtan Yin, Jie Song 0011, Sihui Luo 0001, Mingli Song
AAAI1
2021 Training Generative Adversarial Networks in One Stage
abstract
Generative Adversarial Networks (GANs) have demonstrated unprecedented success in various image generation tasks. The encouraging results, however, come at the price of a cumbersome training process, during which the generator and discriminator are alternately updated in two stages. In this paper, we investigate a general training scheme that enables training GANs efficiently in only one stage. Based on the adversarial losses of the generator and discriminator, we categorize GANs into two classes, Symmetric GANs and Asymmetric GANs, and introduce a novel gradient decomposition method to unify the two, allowing us to train both classes in one stage and hence alleviate the training effort. We also computationally analyze the efficiency of the proposed method, and empirically demonstrate that, the proposed method yields a solid 1.5× acceleration across various datasets and network architectures. Furthermore, we show that the proposed method is readily applicable to other adversarial-training scenarios, such as data-free knowledge distillation. The code is available at https://github.com/zju-vipa/OSGAN.
Chengchao Shen, Youtan Yin, Xinchao Wang, Xubin Li, Jie Song 0011, Mingli Song
CVPR1
2021 Contrastive Model Invertion for Data-Free Knolwedge Distillation
abstract
Model inversion, whose goal is to recover training data from a pre-trained model, has been recently proved feasible. However, existing inversion methods usually suffer from the mode collapse problem, where the synthesized instances are highly similar to each other and thus show limited effectiveness for downstream tasks, such as knowledge distillation. In this paper, we propose Contrastive Model Inversion (CMI), where the data diversity is explicitly modeled as an optimizable objective, to alleviate the mode collapse issue. Our main observation is that, under the constraint of the same amount of data, higher data diversity usually indicates stronger instance discrimination. To this end, we introduce in CMI a contrastive learning objective that encourages the synthesizing instances to be distinguishable from the already synthesized ones in previous batches. Experiments of pre-trained models on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that CMI not only generates more visually plausible instances than the state of the arts, but also achieves significantly superior performance when the generated data are used for knowledge distillation. Code is available at https://github.com/zju-vipa/DataFree.
Gongfan Fang, Jie Song 0011, Xinchao Wang, Chengchao Shen, Xingen Wang, Mingli Song
IJCAI4
2021 Mosaicking to Distill: Knowledge Distillation from Out-of-Domain Data
abstract
Knowledge distillation~(KD) aims to craft a compact student model that imitates the behavior of a pre-trained teacher in a target domain. Prior KD approaches, despite their gratifying results, have largely relied on the premise that \emph{in-domain} data is available to carry out the knowledge transfer. Such an assumption, unfortunately, in many cases violates the practical setting, since the original training data or even the data domain is often unreachable due to privacy or copyright reasons. In this paper, we attempt to tackle an ambitious task, termed as \emph{out-of-domain} knowledge distillation~(OOD-KD), which allows us to conduct KD using only OOD data that can be readily obtained at a very low cost. Admittedly, OOD-KD is by nature a highly challenging task due to the agnostic domain gap. To this end, we introduce a handy yet surprisingly efficacious approach, dubbed as~\textit{MosaicKD}. The key insight behind MosaicKD lies in that, samples from various domains share common local patterns, even though their global semantic may vary significantly; these shared local patterns, in turn, can be re-assembled analogous to mosaic tiling, to approximate the in-domain data and to further alleviating the domain discrepancy. In MosaicKD, this is achieved through a four-player min-max game, in which a generator, a discriminator, a student network, are collectively trained in an adversarial manner, partially under the guidance of a pre-trained teacher. We validate MosaicKD over {classification and semantic segmentation tasks} across various benchmarks, and demonstrate that it yields results much superior to the state-of-the-art counterparts on OOD data. Our code is available at \url{https://github.com/zju-vipa/MosaicKD}.
Gongfan Fang, Yifan Bao, Jie Song 0011, Xinchao Wang, Donglin Xie, Chengchao Shen, Mingli Song
NeurIPS6
2020 DEPARA: Deep Attribution Graph for Deep Knowledge Transferability
abstract
Exploring the intrinsic interconnections between the knowledge encoded in PRe-trained Deep Neural Networks (PR-DNNs) of heterogeneous tasks sheds light on their mutual transferability, and consequently enables knowledge transfer from one task to another so as to reduce the training effort of the latter. In this paper, we propose the DEeP Attribution gRAph (DEPARA) to investigate the transferability of knowledge learned from PR-DNNs. In DEPARA, nodes correspond to the inputs and are represented by their vectorized attribution maps with regards to the outputs of the PR-DNN. Edges denote the relatedness between inputs and are measured by the similarity of their features extracted from the PR-DNN. The knowledge transferability of two PR-DNNs is measured by the similarity of their corresponding DEPARAs. We apply DEPARA to two important yet under-studied problems in transfer learning: pre-trained model selection and layer selection. Extensive experiments are conducted to demonstrate the effectiveness and superiority of the proposed method in solving both these problems. Code, data and models reproducing the results in this paper are available at https://github.com/zju-vipa/DEPARA.
Jie Song 0011, Jingwen Ye, Xinchao Wang, Chengchao Shen, Feng Mao, Mingli Song
CVPR5
2019 Amalgamating Knowledge towards Comprehensive Classification
abstract
With the rapid development of deep learning, there have been an unprecedentedly large number of trained deep network models available online. Reusing such trained models can significantly reduce the cost of training the new models from scratch, if not infeasible at all as the annotations used for the training original networks are often unavailable to public. We propose in this paper to study a new model-reusing task, which we term as knowledge amalgamation. Given multiple trained teacher networks, each of which specializes in a different classification problem, the goal of knowledge amalgamation is to learn a lightweight student model capable of handling the comprehensive classification. We assume no other annotations except the outputs from the teacher models are available, and thus focus on extracting and amalgamating knowledge from the multiple teachers. To this end, we propose a pilot two-step strategy to tackle the knowledge amalgamation task, by learning first the compact feature representations from teachers and then the network parameters in a layer-wise manner so as to build the student model. We apply this approach to four public datasets and obtain very encouraging results: even without any human annotation, the obtained student model is competent to handle the comprehensive classification task and in most cases outperforms the teachers in individual sub-tasks.
Chengchao Shen, Xinchao Wang, Jie Song 0011, Mingli Song
AAAI1
2019 Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
abstract
A massive number of well-trained deep networks have been released by developers online. These networks may focus on different tasks and in many cases are optimized for different datasets. In this paper, we study how to exploit such heterogeneous pre-trained networks, known as teachers, so as to train a customized student network that tackles a set of selective tasks defined by the user. We assume no human annotations are available, and each teacher may be either single- or multi-task. To this end, we introduce a dual-step strategy that first extracts the task-specific knowledge from the heterogeneous teachers sharing the same sub-task, and then amalgamates the extracted knowledge to build the student network. To facilitate the training, we employ a selective learning scheme where, for each unlabelled sample, the student learns adaptively from only the teacher with the least prediction ambiguity. We evaluate the proposed approach on several datasets and the experimental results demonstrate that the student, learned by such adaptive knowledge amalgamation, achieves performances even better than those of the teachers.
Chengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song 0011, Mingli Song
ICCV1
2019 Deep Model Transferability from Attribution Maps
abstract
Exploring the transferability between heterogeneous tasks sheds light on their intrinsic interconnections, and consequently enables knowledge transfer from one task to another so as to reduce the training effort of the latter. In this paper, we propose an embarrassingly simple yet very efficacious approach to estimating the transferability of deep networks, especially those handling vision tasks. Unlike the seminal work of \emph{taskonomy} that relies on a large number of annotations as supervision and is thus computationally cumbersome, the proposed approach requires no human annotations and imposes no constraints on the architectures of the networks. This is achieved, specifically, via projecting deep networks into a \emph{model space}, wherein each network is treated as a point and the distances between two points are measured by deviations of their produced attribution maps. The proposed approach is several-magnitude times faster than taskonomy, and meanwhile preserves a task-wise topological structure highly similar to the one obtained by taskonomy. Code is available at \url{https://github.com/zju-vipa/TransferbilityFromAttributionMaps}.
Jie Song 0011, Xinchao Wang, Chengchao Shen, Mingli Song
NeurIPS4
2018 Transductive Unbiased Embedding for Zero-Shot Learning
abstract
Most existing Zero-Shot Learning (ZSL) methods have the strong bias problem, in which instances of unseen (target) classes tend to be categorized as one of the seen (source) classes. So they yield poor performance after being deployed in the generalized ZSL settings. In this paper, we propose a straightforward yet effective method named Quasi-Fully Supervised Learning (QFSL) to alleviate the bias problem. Our method follows the way of transductive learning, which assumes that both the labeled source images and unlabeled target images are available for training. In the semantic embedding space, the labeled source images are mapped to several fixed points specified by the source categories, and the unlabeled target images are forced to be mapped to other points specified by the target categories. Experiments conducted on AwA2, CUB and SUN datasets demonstrate that our method outperforms existing state-of-the-art approaches by a huge margin of 9.3 ~ 24.5% following generalized ZSL settings, and by a large margin of 0.2 ~ 16.2% following conventional ZSL settings.
Jie Song 0011, Chengchao Shen, Yezhou Yang, Yang Liu 0212, Mingli Song
CVPR2
2018 Selective Zero-Shot Classification with Augmented Attributes
Jie Song 0011, Chengchao Shen, Jie Lei 0002, Anxiang Zeng, Kairi Ou, Dacheng Tao, Mingli Song
ECCV (9)2
2018 DeepSIC: Deep Semantic Image Compression
Sihui Luo 0001, Yezhou Yang, Yanling Yin, Chengchao Shen, Mingli Song
ICONIP (1)4
2018 Intra-class Structure Aware Networks for Screen Defect Detection
Chengchao Shen, Jie Song 0011, Shuyi Song, Sihui Luo 0001, Mingli Song
ICONIP (4)1