EDBT 2026 Demo / reviewers in the wild / expert
Yihong Zhuang
dblp:274/2364
· DBLP profile ↗
17ranked-venue papers
1as first author
17since 2021 · last 2023
0000-0002-3176-3217ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Self-Supervised Image Denoising Using Implicit Deep Denoiser PriorabstractWe devise a new regularization for denoising with self-supervised learning. The regularization uses a deep image prior learned by the network, rather than a traditional predefined prior. Specifically, we treat the output of the network as a ``prior'' that we again denoise after ``re-noising.'' The network is updated to minimize the discrepancy between the twice-denoised image and its prior. We demonstrate that this regularization enables the network to learn to denoise even if it has not seen any clean images. The effectiveness of our method is based on the fact that CNNs naturally tend to capture low-level image statistics. Since our method utilizes the image prior implicitly captured by the deep denoising CNN to guide denoising, we refer to this training strategy as an Implicit Deep Denoiser Prior (IDDP). IDDP can be seen as a mixture of learning-based methods and traditional model-based denoising methods, in which regularization is adaptively formulated using the output of the network. We apply IDDP to various denoising tasks using only observed corrupted data and show that it achieves better denoising results than other self-supervised denoising methods. Huangxing Lin, Yihong Zhuang, Xinghao Ding, Delu Zeng, Yue Huang 0001, Xiaotong Tu, John W. Paisley |
AAAI | 2 |
| 2023 | A Two-Stage Federated Learning Framework for Class Imbalance in Aerial Scene Classification
Zhengpeng Lv, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
PRCV (4) | 2 |
| 2023 | Interclass Similarity Transfer for Imbalanced Aerial Scene ClassificationabstractImbalanced class distributions widely exist in real-world aerial images, which brings a significant challenge to aerial scene classification due to the undesirable bias toward the majority classes as well as overfitting for the minority classes. Although the similarity between different scene classes may be inconsistent, they can be measured by the mean of feature statistics. This motivates us to transfer the statistics of the majority class to the minority class having similar feature statistics. Specifically, based on the observation that the feature statistics of each class may follow the Gaussian distribution, the similarity across different classes would thus be described by the mean of feature statistics. The distributions of minority classes would afterward be calibrated by statistical transfer via interclass similarity (STAIRS), and a sufficient number of features could hence be generated for the minority class to improve its performance in classifier learning. We demonstrate the effectiveness of the proposed method for imbalanced aerial scene classification on the imbalanced aerial image dataset (AID) and NWPU-RESISC45 datasets. The proposed method outperforms alternatives by a large margin in both overall performance and minority classification performance of imbalanced aerial scenes. Changxing Jing, Lexing Huang, Senlin Cai, Yihong Zhuang, Zhenlong Xiao, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Exploring personalization via federated representation Learning on non-IID data
Changxing Jing, Yan Huang 0032, Yihong Zhuang, Liyan Sun, Zhenlong Xiao, Yue Huang 0001, Xinghao Ding |
Neural Networks | 3 |
| 2023 | Unsupervised Video-Based Action Recognition With Imagining Motion and Perceiving AppearanceabstractVideo-based action recognition is a challenging task, which demands carefully considering the temporal property of videos in addition to the appearance attributes. Particularly, the temporal domain of raw videos usually contains significantly more redundant or irrelevant information than still images. For that, this paper proposes an unsupervised video-based action recognition approach with imagining motion and perceiving appearance, called IMPA, by comprehensively learning the spatio-temporal characteristics inherited in videos, with a particular emphasis on the moving object for action recognition. Specifically, a self-supervised Motion Extracting Block (MEB) is designed to extract the principal motion features by focusing on the large movement of the moving object, based on the observation that humans can infer complete motion trajectories from partial moving objects. To further take the indispensable appearance attribute in videos into account, an unsupervised Appearance Learning Block (ALB) is developed to perceive the static appearance, thus in combination with the MEB to recognize actions. Extensive validation experiments and ablation studies on multiple datasets demonstrate that our proposed IMPA approach obtains superior performance and surpasses other classical and state-of-the-art unsupervised action recognition methods. Wei Lin 0021, Yihong Zhuang, Xinghao Ding, Xiaotong Tu, Yue Huang 0001, Huanqiang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Unpaired Speckle Extraction for SAR DespecklingabstractSpeckle suppression is a critical step in synthetic aperture radar (SAR) imaging. Since speckle-free SAR images are inaccessible, supervised denoising methods are not suitable for this task. To exploit the strong capabilities of convolutional neural networks (CNNs), we propose Unpaired Speckle Extraction (SAR-USE), an unsupervised method for SAR despeckling. Our method utilizes unpaired SAR and clean optical images to extract “real” speckle for learning despeckling. First, a CNN that has never seen clean SAR images is employed to extract speckle from the SAR image. Then, the extracted speckle is multiplied with a random optical image to synthesize paired data for learning speckle removal. Through a Siamese network, speckle extraction and learning despeckling are performed alternately and promote each other. To make the extracted speckle more visually and statistically realistic, it is constrained by a noise correction module to be unit mean while maintaining spatial correlation. After convergence, the CNN is a good denoiser that can effectively extract speckle from SAR images. Experiments on synthetic datasets show that the denoising ability of the proposed method is as good as its supervised counterpart. More importantly, SAR-USE is very efficient for removing the spatially correlated speckle in real data that supervised learning methods cannot. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Learning Rate DropoutabstractOptimization algorithms are of great importance to efficiently and effectively train a deep neural network. However, the existing optimization algorithms show unsatisfactory convergence behavior, either slowly converging or not seeking to avoid bad local optima. Learning rate dropout (LRD) is a new gradient descent technique to motivate faster convergence and better generalization. LRD aids the optimizer to actively explore in the parameter space by randomly dropping some learning rates (to 0); at each iteration, only parameters whose learning rate is not 0 are updated. Since LRD reduces the number of parameters to be updated for each iteration, the convergence becomes easier. For parameters that are not updated, their gradients are accumulated (e.g., momentum) by the optimizer for the next update. Accumulating multiple gradients at fixed parameter positions gives the optimizer more energy to escape from the saddle point and bad local optima. Experiments show that LRD is surprisingly effective in accelerating training while preventing overfitting. Huangxing Lin, Weihong Zeng, Yihong Zhuang, Xinghao Ding, Yue Huang 0001, John W. Paisley |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Knowledge Condensation Distillation
Chenxin Li, Mingbao Lin, Zhiyuan Ding, Nie Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Liujuan Cao |
ECCV (11) | 5 |
| 2022 | A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene RecognitionabstractIn real-world scenarios, aerial image datasets are generally class imbalanced, where the majority classes have rich samples, while the minority classes only have a few samples. Such class imbalanced datasets bring great challenges to aerial scene recognition. In this paper, we explore a novel two-stage contrastive learning framework, which aims to take care of representation learning and classifier learning, thereby boosting aerial scene recognition. Specifically, in the representation learning stage, we design a data augmentation policy to improve the potential of contrastive learning according to the characteristics of aerial images. And we employ supervised contrastive learning to learn the association between aerial images of the same scene. In the classification learning stage, we fix the encoder to maintain good representation and use the re-balancing strategy to train a less biased classifier. A variety of experimental results on the imbalanced aerial image datasets show the advantages of the proposed two-stage contrastive learning framework for the imbalanced aerial scene recognition. Lexing Huang, Senlin Cai, Yihong Zhuang, Changxing Jing, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICASSP | 3 |
| 2022 | A Simple Siamese Framework for Vibration Signal RepresentationsabstractSiamese networks are widely used in various contrastive learning methods for recognition tasks, with few labeled data and abundant unlabeled data. In the field of fault diagnosis, it is universal to face the problem that large collections of common fault data and few catastrophic fault samples result in the imbalanced distribution of fault data collection. In this paper, a simple Siamese framework is proposed to learn meaningful signal representations using the differently augmented views of the signals only in the time domain. The industrial fault diagnosis including class balanced and imbalanced motor fault diagnosis is performed to verify the validity of the signal representations. The results demonstrate that the proposed method can significantly balance the representations of both the major and minor classes, which proves the capability of the Siamese framework for class imbalanced classification. Guanxing Zhou, Yihong Zhuang, Xinghao Ding, Yue Huang 0001, Saqlain Abbas, Xiaotong Tu |
ICIP | 2 |
| 2022 | A Hybrid Framework Based on Classifier Calibration for Imbalanced Aerial Scene Recognition
Yihong Zhuang, Changxing Jing, Senlin Cai, Lexing Huang, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICONIP (3) | 1 |
| 2022 | Self-Supervised SAR Despeckling Powered by Implicit Deep Denoiser PriorabstractSpeckle removal is an important preprocessing step for synthetic aperture radar (SAR) imaging. Since speckle-free SAR images do not exist, supervised methods are not applicable. In this letter, we propose implicit deep denoiser prior (SAR-IDDP), a self-supervised method for SAR despeckling. SAR-IDDP uses a deep image prior (DIP) implicitly captured by the convolutional neural network (CNN) to formulate regularization instead of traditional hand-crafted priors. Specifically, we treat the output of the CNN as a “prior” that we denoise again after “renoising.” The CNN is updated to maximize the similarity between the again denoised image and its prior. The renoising procedure is designed based on the assumption of unit mean noise, while the spatial correlation of speckle is also involved. The despeckling ability of our method stems from CNN’s natural tendency to capture low-level image statistics. Experiments show that SAR-IDDP achieves significant improvements over existing model-based and self-supervised despeckling methods on both synthetic and real SAR images. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | One-Shot HRRP Generation for Radar Target RecognitionabstractInsufficient data of a noncooperative target seriously affect the performance of radar automatic target recognition (RATR) using the high-resolution range profile (HRRP), especially when the noncooperative target has only one sample. To this end, we propose an unsupervised data generation method to generate noncooperative HRRP signals. We utilize the pretrained generative adversarial networks (GANs) model to learn the HRRP general probability distribution. To emphasize the representative and discriminative power of generated HRRP signals, a joint optimization method is proposed to preserve category information. Moreover, a feature diversification method is proposed to make the generated samples have sufficient aspect characteristics to further fit the probability distribution of the noncooperative target. Thus, the generated HRRP signals can effectively improve the recognition performance of noncooperative target. Extensive experiments on HRRP data sets demonstrate the superior performance of our method over other state-of-the-art methods. Liangchao Shi, Zhehan Liang, Yi Wen 0003, Yihong Zhuang, Yue Huang 0001, Xinghao Ding |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Dual Domain Multi-Task Model for Vehicle Re-IdentificationabstractVehicle re-identification (re-id) is an essential task in the field of intelligent transportation systems (ITS). The main goal of re-id is to find the same vehicle in different scenarios, which can is still a challenging task in both ITS and computer vision (CV). The existing vehicle re-identification methods simply combine the coarse-grained and the fine-grained attributes together with multi-task training. However, such combination may still have limited performance in vehicles with trivial appearance differences, or with rare models and colors. To solve this problem, we propose a simple yet effective framework, called dual domain multi-task model (DDM), that divides the vehicle images into two domains based on the frequency. And then two parallel branches are proposed to recover the two domains. Furthermore, a multi-task method is proposed, which combines the classification loss in color and model together with triplet loss for fine-grained distance measurement. Besides, a progressive strategy is used in the training process. Two public datasets, PKU VehicleID and VeRi are used to validate the proposed DDM. The experimental results demonstrate that the proposed approach outperforms the existing methods on both datasets. Yue Huang 0001, Borong Liang, Weiping Xie, Yinghao Liao, Zhenyu Kuang, Yihong Zhuang, Xinghao Ding |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Harmonizing Pathological and Normal Pixels for Pseudo-Healthy SynthesisabstractSynthesizing a subject-specific pathology-free image from a pathological image is valuable for algorithm development and clinical practice. In recent years, several approaches based on the Generative Adversarial Network (GAN) have achieved promising results in pseudo-healthy synthesis. However, the discriminator (i.e., a classifier) in the GAN cannot accurately identify lesions and further hampers from generating admirable pseudo-healthy images. To address this problem, we present a new type of discriminator, the segmentor, to accurately locate the lesions and improve the visual quality of pseudo-healthy images. Then, we apply the generated images into medical image enhancement and utilize the enhanced results to cope with the low contrast problem existing in medical image segmentation. Furthermore, a reliable metric is proposed by utilizing two attributes of label noise to measure the health of synthetic images. Comprehensive experiments on the T2 modality of BraTS demonstrate that the proposed method substantially outperforms the state-of-the-art methods. The method achieves better performance than the existing methods with only 30% of the training data. The effectiveness of the proposed method is also demonstrated on the LiTS and the T1 modality of BraTS. The code and the pre-trained model of this study are publicly available at https://github.com/Au3C2/Generator-Versus-Segmentor. Yihong Zhuang, Liyan Sun, Yue Huang 0001, Xinghao Ding, Guisheng Wang, Lin Yang 0002, Yizhou Yu |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Noise2Grad: Extract Image Noise to DenoiseabstractIn many image denoising tasks, the difficulty of collecting noisy/clean image pairs limits the application of supervised CNNs. We consider such a case in which paired data and noise statistics are not accessible, but unpaired noisy and clean images are easy to collect. To form the necessary supervision, our strategy is to extract the noise from the noisy image to synthesize new data. To ease the interference of the image background, we use a noise removal module to aid noise extraction. The noise removal module first roughly removes noise from the noisy image, which is equivalent to excluding much background information. A noise approximation module can therefore easily extract a new noise map from the removed noise to match the gradient of the noisy input. This noise map is added to a random clean image to synthesize a new data pair, which is then fed back to the noise removal module to correct the noise removal process. These two modules cooperate to extract noise finely. After convergence, the noise removal module can remove noise without damaging other background details, so we use it as our final denoising network. Experiments show that the denoising performance of the proposed method is competitive with other supervised CNNs. Huangxing Lin, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
IJCAI | 2 |
| 2021 | Generator Versus Segmentor: Pseudo-healthy Synthesis
Chenxin Li, Liyan Sun, Yihong Zhuang, Yue Huang 0001, Xinghao Ding, Yizhou Yu |
MICCAI (6) | 5 |