VLDB 2026 Research / reviewers in the wild / expert
Yinghuan Shi
dblp:30/7184
· DBLP profile ↗
135ranked-venue papers
11as first author
84since 2021 · last 2026
0000-0003-4534-7318ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 8 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 5 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 29 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRAabstractContinual learning (CL) in vision-language models (VLMs) faces significant challenges in improving task adaptation and avoiding catastrophic forgetting. Existing methods usually have heavy inference burden or rely on external knowledge, while Low-Rank Adaptation (LoRA) has shown potential in reducing these issues by enabling parameter-efficient tuning. However, considering directly using LoRA to alleviate the catastrophic forgetting problem is non-trivial, we introduce a novel framework that restructures a single LoRA module as a decomposable Rank-1 Expert Pool. Our method learns to dynamically compose a sparse, task-specific update by selecting from this expert pool, guided by the semantics of the [CLS] token. In addition, we propose an Activation-Guided Orthogonal (AGO) loss that orthogonalizes critical parts of LoRA weights across tasks. This sparse composition and orthogonalization enable fewer parameter updates, resulting in domain-aware learning while minimizing inter-task interference and maintaining downstream task performance. Extensive experiments across multiple settings demonstrate state-of-the-art results in all metrics, surpassing zero-shot upper bounds in generalization. Notably, it reduces trainable parameters by 96.7% compared to the baseline method, eliminating reliance on external datasets or task-ID discriminators. The merged LoRAs retain less weights and incur no inference latency, making our method computationally lightweight. Zhan Fa, Yue Duan, Jian Zhang 0090, Lei Qi 0001, Wanqi Yang, Yinghuan Shi |
AAAI | 6 |
| 2026 | Diffusion-Based Data Augmentation for Image Recognition: A Systematic Analysis and Evaluation
Zekun Li 0010, Yinghuan Shi, Yang Gao 0001 |
Int. J. Comput. Vis. | 2 |
| 2026 | Toward the Connection Between Activation Sparsity and Flat MinimaabstractThe observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing performance. To theoretically explain this phenomenon, existing works have shown that activation sparsity does not result from the data properties or data fitting but from the implicit bias of the training process. However, these connections are obtained with strong assumptions (e.g., shallow networks, a small number of training steps, and special training techniques), which cannot be applied to deep models standardly trained with a large number of steps. Different from these works, we find that the flatness of loss landscapes is also closely related to the MLP activation sparsity and can serve as a weaker assumption because it naturally emerges in the standard training of deep networks without the above strong assumptions. Specifically, we find that 1) the MLP activation sparsity equals a ratio between "augmented flatness" (a weighted sum of flatness measures) and the product of the input norm and activation gradient of the MLP. We empirically find that this ratio decreases during training, leading to sparse activations. 2) We also propose the notion of derivative sparsity, which reduces to activation sparsity under $\operatorname{ReLU}$ReLU, but further enables pruning in the backward propagation and is more stable than activation sparsity. With the theoretical findings, we can further encourage activation sparsity by decreasing the numerator and increasing the denominator of the ratio: 1) To improve (lower) the flatness, we add different bias vectors to input tokens of MLP blocks to strengthen stochastic gradient noise that drives the model to a flat area. 2) We restrict the lower bound of affine parameters in LayerNorm to increase the input norm of MLPs. 3) To increase the activation sparsity, we propose an activation function $\operatorname{JSReLU}$JSReLU to encourage the search of parameters with sparse derivatives and sparse activations. These plug-and-play modifications can effectively reduce the ratio and produce sparser activations. Experiments on ImageNet-1K and C4 demonstrate relative improvements of at least 36% on inference sparsity and at least 50% on training sparsity over vanilla Transformers, indicating further potential cost reduction in both inference and training. Ze Peng 0001, Jian Zhang 0090, Lei Qi 0001, Yang Gao 0001, Yinghuan Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | A hybrid dual-augmentation constraint framework for single-source domain generalization in medical image segmentation
Jintao Guo, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi |
Pattern Recognit. | 6 |
| 2026 | Diversity-Enhanced Collaborative Mamba for Semi-Supervised Medical Image SegmentationabstractAcquiring high-quality annotated data for medical image segmentation is tedious and costly. Semi-supervised segmentation techniques alleviate this burden by leveraging unlabeled data to generate pseudo labels. Recently, advanced state space models, represented by Mamba, have shown efficient handling of long-range dependencies. This drives us to explore their potential in semi-supervised medical image segmentation. In this paper, we propose a novel Diversity-enhanced Collaborative Mamba framework (namely DCMamba) for semi-supervised medical image segmentation, which explores and utilizes the diversity from data, network, and feature perspectives. Firstly, from the data perspective, we develop patch-level weak-strong mixing augmentation with Mamba's scanning modeling characteristics. Moreover, from the network perspective, we introduce a diverse-scan collaboration module, which could benefit from the prediction discrepancies arising from different scanning directions. Furthermore, from the feature perspective, we adopt an uncertainty-weighted contrastive learning mechanism to enhance the diversity of feature representation. Experiments demonstrate that our DCMamba significantly outperforms other semi-supervised medical image segmentation methods, e.g., yielding the latest SSM-based method by 6.69% on the Synapse dataset with 20% labeled data. The code is available at https://github.com/ShumengLI/DCMamba. Shumeng Li, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIPabstractContrastive Language-Image Pretraining (CLIP) has been widely used in vision tasks. Notably, CLIP has demonstrated promising performance in few-shot learning (FSL). However, existing CLIP-based methods in training-free FSL (i.e., without the requirement of additional training) mainly learn different modalities independently, leading to two essential issues: 1) severe anomalous match in image modality; 2) varying quality of generated text prompts. To address these issues, we build a mutual guidance mechanism, that introduces an Image-Guided-Text (IGT) component to rectify varying quality of text prompts through image representations, and a Text-Guided-Image (TGI) component to mitigate the anomalous match of image modality through text representations. By integrating IGT and TGI, we adopt a perspective of Text-Image Mutual guidance Optimization, proposing TIMO. Extensive experiments show that TIMO significantly outperforms the state-of-the-art (SOTA) training-free method. Additionally, by exploring the extent of mutual guidance, we propose an enhanced variant, TIMO-S, which even surpasses the best training-required methods by 0.33% with approximately ×100 less time cost. Yayuan Li, Jintao Guo, Lei Qi 0001, Wenbin Li 0006, Yinghuan Shi |
AAAI | 5 |
| 2025 | Steady Progress Beats Stagnation: Mutual Aid of Foundation and Conventional Models in Mixed Domain Semi-Supervised Medical Image SegmentationabstractLarge pretrained visual foundation models exhibit impressive general capabilities. However, the extensive prior knowledge inherent in these models can sometimes be a double-edged sword when adapting them to downstream tasks in specific domains. In the context of semi-supervised medical image segmentation with domain shift, foundation models like MedSAM tend to make overconfident predictions, some of which are incorrect. The error accumulation hinders the effective utilization of unlabeled data and limits further improvements. In this paper, we introduce a Synergistic training framework for Foundation and Conventional models (SynFoC) to address the issue. We observe that a conventional model trained from scratch has the ability to correct the high-confidence mispredictions of the foundation model, while the foundation model can supervise it with high-quality pseudo-labels in the early training stages. Furthermore, to enhance the collaborative training effectiveness of both models and promote reliable convergence towards optimization, the consensus-divergence consistency regularization is proposed. We demonstrate the superiority of our method across four public multi-domain datasets. In particular, our method improves the Dice score by 10.31% on the Prostate dataset. Our code is available at https://github.com/MQinghe/SynFoC. Qinghe Ma, Jian Zhang 0090, Zekun Li 0010, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi |
CVPR | 6 |
| 2025 | Balanced Direction from Multifarious Choices: Arithmetic Meta-Learning for Domain GeneralizationabstractDomain generalization is proposed to address distribution shift, arising from statistical disparities between training source and unseen target domains. The widely used first-order meta-learning algorithms demonstrate strong performance for domain generalization by leveraging the gradient matching theory, which aims to establish balanced parameters across source domains to reduce overfitting to any particular domain. However, our analysis reveals that there are actually numerous directions to achieve gradient matching, with current methods representing just one possible path. These methods actually overlook another critical factor that the balanced parameters should be close to the centroid of optimal parameters of each source domain. To address this, we propose a simple yet effective arithmetic meta-learning with arithmetic-weighted gradients. This approach, while adhering to the principles of gradient matching, promotes a more precise balance by estimating the centroid between domain-specific optimal parameters. Experimental results validate the effectiveness of our strategy. Our code is available at https://github.com/zzwdx/ARITH. Xiran Wang, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
CVPR | 4 |
| 2025 | Taste More, Taste Better: Diverse Data and Strong Model Boost Semi-Supervised Crowd CountingabstractSemi-supervised crowd counting is crucial for addressing the high annotation costs of densely populated scenes. Although several methods based on pseudo-labeling have been proposed, it remains challenging to effectively and accurately utilize unlabeled data. In this paper, we propose a novel framework called Taste More Taste Better (TMTB), which emphasizes both data and model aspects. Firstly, we explore a data augmentation technique well-suited for the crowd counting task. By inpainting the background regions, this technique can effectively enhance data diversity while preserving the fidelity of the entire scenes. Secondly, we introduce the Visual State Space Model as backbone to capture the global context information from crowd scenes, which is crucial for extremely crowded, lowlight, and adverse weather scenarios. In addition to the traditional regression head for exact prediction, we employ an Anti-Noise classification head to provide less exact but more accurate supervision, since the regression head is sensitive to noise in manual annotations. We conduct extensive experiments on four benchmark datasets and show that our method outperforms state-of-the-art methods by a large margin. Code is publicly available on https://github.com/syhien/taste_more_taste_better. Maochen Yang, Zekun Li 0010, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
CVPR | 5 |
| 2025 | Divide-And-Conquer for Enhancing Unlabeled Learning, Stability, and Plasticity in Semi-Supervised Continual Learning
Yue Duan, Taicai Chen, Lei Qi 0001, Yinghuan Shi |
ICCV | 4 |
| 2025 | Correspondence as Video: Test-Time Adaption on SAM2 for Reference Segmentation in the WildabstractLarge vision models like the Segment Anything Model (SAM) exhibit significant limitations when applied to downstream tasks in the wild. Consequently, reference segmentation, which leverages reference images and their corresponding masks to impart novel knowledge to the model, emerges as a promising new direction for adapting vision models. However, existing reference segmentation approaches predominantly rely on meta-learning, which still necessitates an extensive meta-training process and brings massive data and computational cost. In this study, we propose a novel approach by representing the inherent correspondence between reference-target image pairs as a pseudo video. This perspective allows the latest version of SAM, known as SAM2, which is equipped with interactive video object segmentation (iVOS) capabilities, to be adapted to downstream tasks in a lightweight manner. We term this approach Correspondence As Video for SAM (CAV-SAM). CAV-SAM comprises two key modules: the Diffusion-Based Semantic Transition (DBST) module employs a diffusion model to construct a semantic transformation sequence, while the Test-Time Geometric Alignment (TTGA) module aligns the geometric changes within this sequence through test-time fine-tuning. We evaluated CAVSAM on widely-used datasets, achieving segmentation performance improvements exceeding 5% over SOTA methods. Implementation is provided in the supplementary materials. Zekun Li 0010, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
ICCV | 5 |
| 2025 | Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGDabstractInformation-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so that one can exploit the properties of data and algorithm to derive tighter bounds. However, we observe that although the flatness bias is crucial for SGD’s generalization, these bounds fail to capture the improved generalization under better flatness and are also numerically loose. This is caused by the inadequate leverage of SGD's flatness bias in existing IT bounds. This paper derives a more flatness-leveraging IT bound for the flatness-favoring SGD. The bound indicates the learned models generalize better if the large-variance directions of the final weight covariance have small local curvatures in the loss landscape. Experiments on deep neural networks show our bound not only correctly reflects the better generalization when flatness is improved, but is also numerically much tighter. This is achieved by a flexible technique called "omniscient trajectory". When applied to Gradient Descent’s minimax excess risk on convex-Lipschitz-Bounded problems, it improves representative IT bounds’ $\Omega(1)$ rates to $O(1/\sqrt{n})$. It also implies a by-pass of memorization-generalization trade-offs. Codes are available at [https://github.com/peng-ze/omniscient-bounds](https://github.com/peng-ze/omniscient-bounds). Ze Peng 0001, Jian Zhang 0090, Yisen Wang 0001, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ICLR | 5 |
| 2025 | GA-SAM: Geometry-Aware SAM Adaptation with Sparse Annotation-Driven Point Cloud Completion
Shumeng Li, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
MICCAI (8) | 4 |
| 2025 | Fusing Dual Encoders: Single-Source Domain Generalization with Extremely Few Annotations
Jintao Guo, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
MICCAI (6) | 5 |
| 2025 | Towards Perfection: Building Inter-component Mutual Correction for Retinex-based Low-light Image EnhancementabstractIn low-light image enhancement, Retinex-based deep learning methods have garnered significant attention due to their exceptional interpretability. These methods decompose images into mutually independent illumination and reflectance components, allows each component to be enhanced separately. In fact, achieving perfect decomposition of illumination and reflectance components proves to be quite challenging, with some residuals still existing after decomposition. In this paper, we formally name these residuals as inter-component residuals (ICR), which has been largely underestimated by previous methods. In our investigation, ICR not only affects the accuracy of the decomposition but also causes enhanced components to deviate from the ideal outcome, ultimately reducing the final synthesized image quality. To address this issue, we propose a novel Inter-correction Retinex model (IRetinex) to alleviate ICR during the decomposition and enhancement stage. In the decomposition stage, we leverage inter-component residual reduction module to reduce the feature similarity between illumination and reflectance components. In the enhancement stage, we utilize the feature similarity between the two components to detect and mitigate the impact of ICR within each enhancement unit. Extensive experiments on three low-light benchmark datasets demonstrated that by reducing ICR, our method outperforms state-of-the-art approaches both qualitatively and quantitatively. Our code is available at: https://github.com/caoluyang0830/IRetinex.git. Luyang Cao, Han Xu 0001, Jian Zhang 0090, Lei Qi 0001, Jiayi Ma 0001, Yinghuan Shi, Yang Gao 0001 |
ACM Multimedia | 6 |
| 2025 | Patch-Aware Batch Normalization for Improving Cross-Domain RobustnessabstractDespite the significant success of deep learning in computer vision tasks, cross-domain tasks still present a challenge in which the model’s performance will degrade when the training set and the test set follow different distributions. Most existing methods employ adversarial learning or instance normalization for achieving data augmentation to solve this task. In contrast, considering that the batch normalization (BN) layer may not be robust for unseen domains and there exist the differences between local patches of an image, we propose a novel method called patch-aware batch normalization (PBN). To be specific, we first split feature maps of a batch into non-overlapping patches along the spatial dimension, and then independently normalize each patch to jointly optimize the shared BN parameter at each iteration. By exploiting the differences between local patches of an image, our proposed PBN can effectively enhance the robustness of the model’s parameters. Besides, considering the statistics from each patch may be inaccurate due to their smaller size compared to the global feature maps, we incorporate the globally accumulated statistics with the statistics from each batch to obtain the final statistics for normalizing each patch. Since the proposed PBN can replace the typical BN, it can be integrated into most existing state-of-the-art methods. Extensive experiments and analysis demonstrate the effectiveness of our PBN in multiple computer vision tasks, including classification, object detection, instance retrieval, and semantic segmentation. Lei Qi 0001, Dongjia Zhao, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Background Matters: A Cross-View Bidirectional Modeling Framework for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation (SSMIS) leverages unlabeled data to reduce reliance on manually annotated images. However, current SOTA approaches predominantly focus on foreground-oriented modeling (i.e., segmenting only the foreground region) and have largely overlooked the potential benefits of explicitly modeling the background region. Our study theoretically and empirically demonstrates that highly certain predictions in background modeling enhance the confidence of corresponding foreground modeling. Building on this insight, we propose the Cross-view Bidirectional Modeling (CVBM) framework, which introduces a novel perspective by incorporating background modeling to improve foreground modeling performance. Within CVBM, background modeling serves as an auxiliary perspective, providing complementary supervisory signals to enhance the confidence of the foreground model. Additionally, CVBM introduces an innovative bidirectional consistency mechanism, which ensures mutual alignment between foreground predictions and background-guided predictions. Extensive experiments demonstrate that our approach achieves SOTA performance on the LA, Pancreas, ACDC, and HRF datasets. Notably, on the Pancreas dataset, CVBM outperforms fully supervised methods (i.e., DSC: 84.57% vs. 83.89%) while utilizing only 20% of the labeled data. Our code is publicly available at https://github.com/caoluyang0830/CVBM.git. Luyang Cao, Yinghuan Shi |
IEEE Trans. Image Process. | 3 |
| 2025 | An Adaptor for Triggering Semi-Supervised Learning to Out-of-Box Serve Deep Image ClusteringabstractRecently, some works integrate SSL techniques into deep clustering frameworks to enhance image clustering performance. However, they all need pretraining, clustering learning, or a trained clustering model as prerequisites, limiting the flexible and out-of-box application of SSL learners in the image clustering task. This work introduces ASD, an adaptor that enables the cold-start of SSL learners for deep image clustering without any prerequisites. Specifically, we first randomly sample pseudo-labeled data from all unlabeled data, and set an instance-level classifier to learn them with semantically aligned instance-level labels. With the ability of instance-level classification, we track the class transitions of predictions on unlabeled data to extract high-level similarities of instance-level classes, which can be utilized to assign cluster-level labels to pseudo-labeled data. Finally, we use the pseudo-labeled data with assigned cluster-level labels to trigger a general SSL learner trained on the unlabeled data for image clustering. We show the superior performance of ASD across various benchmarks against the latest deep image clustering approaches and very slight accuracy gaps compared to SSL methods using ground-truth, e.g., only 1.33% on CIFAR-10. Moreover, ASD can also further boost the performance of existing SSL-embedded deep image clustering methods. Yue Duan, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Mamba-Sea: A Mamba-Based Framework With Global-to-Local Sequence Augmentation for Generalizable Medical Image SegmentationabstractTo segment medical images with distribution shifts, domain generalization (DG) has emerged as a promising setting to train models on source domains that can generalize to unseen target domains. Existing DG methods are mainly based on CNN or ViT architectures. Recently, advanced state space models, represented by Mamba, have shown promising results in various supervised medical image segmentation. The success of Mamba is primarily owing to its ability to capture long-range dependencies while keeping linear complexity with input sequence length, making it a promising alternative to CNNs and ViTs. Inspired by the success, in the paper, we explore the potential of the Mamba architecture to address distribution shifts in DG for medical image segmentation. Specifically, we propose a novel Mamba-based framework, Mamba-Sea, incorporating global-to-local sequence augmentation to improve the model's generalizability under domain shift issues. Our Mamba-Sea introduces a global augmentation mechanism designed to simulate potential variations in appearance across different sites, aiming to suppress the model's learning of domain-specific information. At the local level, we propose a sequence-wise augmentation along input sequences, which perturbs the style of tokens within random continuous sub-sequences by modeling and resampling style statistics associated with domain shifts. To our best knowledge, Mamba-Sea is the first work to explore the generalization of Mamba for medical image segmentation, providing an advanced and promising Mamba-based architecture with strong robustness to domain shifts. Remarkably, our proposed method is the first to surpass a Dice coefficient of 90% on the Prostate dataset, which exceeds previous SOTA of 88.61%. The code is available at https://github.com/orange-czh/Mamba-Sea. Zihan Cheng 0001, Jintao Guo, Jian Zhang 0090, Lei Qi 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Stitching, Fine-Tuning, and Re-Training: A SAM-Enabled Framework for Semi-Supervised 3D Medical Image SegmentationabstractSegment Anything Model (SAM) fine-tuning has shown remarkable performance in medical image segmentation in a fully supervised manner, but requires precise annotations. To reduce the annotation cost and maintain satisfactory performance, in this work, we leverage the capabilities of SAM for establishing semi-supervised medical image segmentation models. Rethinking the requirements of effectiveness, efficiency, and compatibility, we propose a three-stage framework, i.e., Stitching, Fine-tuning, and Re-training (SFR). The current fine-tuning approaches mostly involve 2D slice-wise fine-tuning that disregards the contextual information between adjacent slices. Our stitching strategy mitigates the mismatch between natural and 3D medical images. The stitched images are then used for fine-tuning SAM, providing robust initialization of pseudo-labels. Afterwards, we train a 3D semi-supervised segmentation model while maintaining the same parameter size as the conventional segmenter such as V-Net. Our SFR framework is plug-and-play, and easily compatible with various popular semi-supervised methods. We also develop an extended framework SFR+ with selective fine-tuning and re-training through confidence estimation. Extensive experiments validate that our SFR and SFR+ achieve significant improvements in both moderate annotation and scarce annotation across five datasets. In particular, SFR framework improves the Dice score of Mean Teacher from 29.68% to 74.40% with only one labeled data of LA dataset. The code is available at https://github.com/ShumengLI/SFR. Shumeng Li, Lei Qi 0001, Qian Yu 0007, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Unleashing the Power of Intermediate Domains for Mixed Domain Semi-Supervised Medical Image SegmentationabstractBoth limited annotation and domain shift are prevalent challenges in medical image segmentation. Traditional semi-supervised segmentation and unsupervised domain adaptation methods address one of these issues separately. However, the coexistence of limited annotation and domain shift is quite common, which motivates us to introduce a novel and challenging scenario: Mixed Domain Semi-supervised medical image Segmentation (MiDSS), where limited labeled data from a single domain and a large amount of unlabeled data from multiple domains. To tackle this issue, we propose the UST-RUN framework, which fully leverages intermediate domain information to facilitate knowledge transfer. We employ Unified Copy-paste (UCP) to construct intermediate domains, and propose a Symmetric GuiDance training strategy (SymGD) to supervise unlabeled data by merging pseudo-labels from intermediate samples. Subsequently, we introduce a Training Process aware Random Amplitude MixUp (TP-RAM) to progressively incorporate style-transition components into intermediate samples. To generate more diverse intermediate samples, we further select reliable samples with high-quality pseudo-labels, which are then mixed with other unlabeled data. Additionally, we generate sophisticated intermediate samples with high-quality pseudo-labels for unreliable samples, ensuring effective knowledge transfer for them. Extensive experiments on four public datasets demonstrate the superiority of UST-RUN. Notably, UST-RUN achieves a 12.94% improvement in Dice score on the Prostate dataset. Our code is available at https://github.com/MQinghe/UST-RUN. Qinghe Ma, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | WeakMedSAM: Weakly-Supervised Medical Image Segmentation via SAM With Sub-Class Exploration and Prompt Affinity MiningabstractWe have witnessed remarkable progress in foundation models in vision tasks. Currently, several recent works have utilized the segmenting anything model (SAM) to boost the segmentation performance in medical images, where most of them focus on training an adaptor for fine-tuning a large amount of pixel-wise annotated medical images following a fully supervised manner. In this paper, to reduce the labeling cost, we investigate a novel weakly-supervised SAM-based segmentation model, namely WeakMedSAM. Specifically, our proposed WeakMedSAM contains two modules: 1) to mitigate severe co-occurrence in medical images, a sub-class exploration module is introduced to learn accurate feature representations. 2) to improve the quality of the class activation maps, our prompt affinity mining module utilizes the prompt capability of SAM to obtain an affinity map for random-walk refinement. Our method can be applied to any SAM-like backbone, and we conduct experiments with SAMUS and EfficientSAM. The experimental results on three popularly-used benchmark datasets, i.e., BraTS 2019, AbdomenCT-1K, and MSD Cardiac dataset, show the promising results of our proposed WeakMedSAM. Our code is available at https://github.com/wanghr64/WeakMedSAM. Lian Huai, Wenbin Li 0006, Lei Qi 0001, Xingqun Jiang, Yinghuan Shi |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Balancing Multi-Target Semi-Supervised Medical Image Segmentation With Collaborative Generalist and SpecialistsabstractDespite the promising performance achieved by current semi-supervised models in segmenting individual medical targets, many of these models suffer a notable decrease in performance when tasked with the simultaneous segmentation of multiple targets. A vital factor could be attributed to the imbalanced scales among different targets: during simultaneously segmenting multiple targets, large targets dominate the loss, leading to small targets being misclassified as larger ones. To this end, we propose a novel method, which consists of a Collaborative Generalist and several Specialists, termed CGS. It is centered around the idea of employing a specialist for each target class, thus avoiding the dominance of larger targets. The generalist performs conventional multi-target segmentation, while each specialist is dedicated to distinguishing a specific target class from the remaining target classes and the background. Based on a theoretical insight, we demonstrate that CGS can achieve a more balanced training. Moreover, we develop cross-consistency losses to foster collaborative learning between the generalist and the specialists. Lastly, regarding their intrinsic relation that the target class of any specialized head should belong to the remaining classes of the other heads, we introduce an inter-head error detection module to further enhance the quality of pseudo-labels. Experimental results on three popular benchmarks showcase its superior performance compared to state-of-the-art methods. Our code is available at https://github.com/wangyou0804/CGS. Zekun Li 0010, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | PG-LBO: Enhancing High-Dimensional Bayesian Optimization with Pseudo-Label and Gaussian Process GuidanceabstractVariational Autoencoder based Bayesian Optimization (VAE-BO) has demonstrated its excellent performance in addressing high-dimensional structured optimization problems. However, current mainstream methods overlook the potential of utilizing a pool of unlabeled data to construct the latent space, while only concentrating on designing sophisticated models to leverage the labeled data. Despite their effective usage of labeled data, these methods often require extra network structures, additional procedure, resulting in computational inefficiency. To address this issue, we propose a novel method to effectively utilize unlabeled data with the guidance of labeled data. Specifically, we tailor the pseudo-labeling technique from semi-supervised learning to explicitly reveal the relative magnitudes of optimization objective values hidden within the unlabeled data. Based on this technique, we assign appropriate training weights to unlabeled data to enhance the construction of a discriminative latent space. Furthermore, we treat the VAE encoder and the Gaussian Process (GP) in Bayesian optimization as a unified deep kernel learning process, allowing the direct utilization of labeled data, which we term as Gaussian Process guidance. This directly and effectively integrates the goal of improving GP accuracy into the VAE training, thereby guiding the construction of the latent space. The extensive experiments demonstrate that our proposed method outperforms existing VAE-BO algorithms in various optimization scenarios. Our code will be published at https://github.com/TaicaiChen/PG-LBO. Taicai Chen, Yue Duan, Dong Li 0016, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 5 |
| 2024 | Roll with the Punches: Expansion and Shrinkage of Soft Label Selection for Semi-supervised Fine-Grained LearningabstractWhile semi-supervised learning (SSL) has yielded promising results, the more realistic SSL scenario remains to be explored, in which the unlabeled data exhibits extremely high recognition difficulty, e.g., fine-grained visual classification in the context of SSL (SS-FGVC). The increased recognition difficulty on fine-grained unlabeled data spells disaster for pseudo-labeling accuracy, resulting in poor performance of the SSL model. To tackle this challenge, we propose Soft Label Selection with Confidence-Aware Clustering based on Class Transition Tracking (SoC) by reconstructing the pseudo-label selection process by jointly optimizing Expansion Objective and Shrinkage Objective, which is based on a soft label manner. Respectively, the former objective encourages soft labels to absorb more candidate classes to ensure the attendance of ground-truth class, while the latter encourages soft labels to reject more noisy classes, which is theoretically proved to be equivalent to entropy minimization. In comparisons with various state-of-the-art methods, our approach demonstrates its superior performance in SS-FGVC. Checkpoints and source code are available at https://github.com/NJUyued/SoC4SS-FGVC. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
AAAI | 6 |
| 2024 | Constructing and Exploring Intermediate Domains in Mixed Domain Semi-supervised Medical Image SegmentationabstractBoth limited annotation and domain shift are preva-lent challenges in medical image segmentation. Traditional semi-supervised segmentation and unsupervised do-main adaptation methods address one of these issues sepa-rately. However, the coexistence of limited annotation and domain shift is quite common, which motivates us to in-troduce a novel and challenging scenario: Mixed Domain Semi-supervised medical image Segmentation (MiDSS). In this scenario, we handle data from multiple medical cen-ters, with limited annotations available for a single do-main and a large amount of unlabeled data from multi-ple domains. We found that the key to solving the prob-lem lies in how to generate reliable pseudo labels for the unlabeled data in the presence of domain shift with la-beled data. To tackle this issue, we employ Unified Copy-Paste (UCP) between images to construct intermediate do-mains, facilitating the knowledge transfer from the do-main of labeled data to the domains of unlabeled data. To fully utilize the information within the intermediate do-main, we propose a symmetric Guidance training strategy (SymGD), which additionally offers direct guidance to un-labeled data by merging pseudo labels from intermediate samples. Subsequently, we introduce a Training Process aware Random Amplitude MixUp (TP-RAM) to progres-sively incorporate style-transition components into inter-mediate samples. Compared with existing state-of-the-art approaches, our method achieves a notable 13.57% im-provement in Dice score on Prostate dataset, as demon-strated on three public datasets. Our code is available at https://github.com/MQinghe/MiDSS Qinghe Ma, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
CVPR | 5 |
| 2024 | Learn to Preserve and Diversify: Parameter-Efficient Group with Orthogonal Regularization for Domain Generalization
Jiajun Hu, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ECCV (51) | 4 |
| 2024 | The Devil Is in the Statistics: Mitigating and Exploiting Statistics Difference for Generalizable Semi-supervised Medical Image Segmentation
Muyang Qiu, Jian Zhang 0090, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
ECCV (54) | 5 |
| 2024 | PC2: Pseudo-Classification Based Pseudo-Captioning for Noisy Correspondence Learning in Cross-Modal RetrievalabstractIn the realm of cross-modal retrieval, seamlessly integrating diverse modalities within multimedia remains a formidable challenge, especially given the complexities introduced by noisy correspondence learning (NCL). Such noise often stems from mismatched data pairs, which is a significant obstacle distinct from traditional noisy labels. This paper introduces Pseudo-Classification based Pseudo-Captioning (PC$^2$) framework to address this challenge. PC$^2$ offers a threefold strategy: firstly, it establishes an auxiliary "pseudo-classification" task that interprets captions as categorical labels, steering the model to learn image-text semantic similarity through a non-contrastive mechanism. Secondly, unlike prevailing margin-based techniques, capitalizing on PC$^2$'s pseudo-classification capability, we generate pseudo-captions to provide more informative and tangible supervision for each mismatched pair. Thirdly, the oscillation of pseudo-classification is borrowed to assistant the correction of correspondence. In addition to technical contributions, we develop a realistic NCL dataset called Noise of Web (NoW), which could be a new powerful NCL benchmark where noise exists naturally. Empirical evaluations of PC$^2$ showcase marked improvements over existing state-of-the-art robust cross-modal retrieval techniques on both simulated and realistic datasets with various NCL settings. The contributed dataset and source code are released at https://github.com/alipay/PC2-NoiseofWeb. Yue Duan, Zhangxuan Gu, Zhenzhe Ying, Lei Qi 0001, Changhua Meng, Yinghuan Shi |
ACM Multimedia | 6 |
| 2024 | START: A Generalized State Space Model with Saliency-Driven Token-Aware TransformationabstractDomain Generalization (DG) aims to enable models to generalize to unseen target domains by learning from multiple source domains. Existing DG methods primarily rely on convolutional neural networks (CNNs), which inherently learn texture biases due to their limited receptive fields, making them prone to overfitting source domains. While some works have introduced transformer-based methods (ViTs) for DG to leverage the global receptive field, these methods incur high computational costs due to the quadratic complexity of self-attention. Recently, advanced state space models (SSMs), represented by Mamba, have shown promising results in supervised learning tasks by achieving linear complexity in sequence length during training and fast RNN-like computation during inference. Inspired by this, we investigate the generalization ability of the Mamba model under domain shifts and find that input-dependent matrices within SSMs could accumulate and amplify domain-specific features, thus hindering model generalization. To address this issue, we propose a novel SSM-based architecture with saliency-based token-aware transformation (namely START), which achieves state-of-the-art (SOTA) performances and offers a competitive alternative to CNNs and ViTs. Our START can selectively perturb and suppress domain-specific features in salient tokens within the input-dependent matrices of SSMs, thus effectively reducing the discrepancy between different domains. Extensive experiments on five benchmarks demonstrate that START outperforms existing SOTA DG methods with efficient linear complexity. Our code is available at https://github.com/lingeringlight/START. Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
NeurIPS | 3 |
| 2024 | Domain generalization via Inter-domain Alignment and Intra-domain Expansion
Jiajun Hu, Lei Qi 0001, Jian Zhang 0090, Yinghuan Shi |
Pattern Recognit. | 4 |
| 2024 | Generalizable Metric Network for Cross-Domain Person Re-IdentificationabstractPerson Re-identification (Re-ID) is a crucial technique for public security and has made significant progress in supervised settings. However, the cross-domain (i.e., domain generalization) scene presents a challenge in Re-ID tasks due to unseen test domains and domain-shift between the training and test sets. To tackle this challenge, most existing methods aim to learn domain-invariant or robust features for all domains. In this paper, we observe that the data-distribution gap between the training and test sets is smaller in the sample-pair space than in the sample-instance space. Based on this observation, we propose a Generalizable Metric Network (GMN) to further explore sample similarity in the sample-pair space. Specifically, we add a Metric Network (M-Net) after the main network and train it on positive and negative sample-pair features, which is then employed during the test stage. Additionally, we introduce the Dropout-based Perturbation (DP) module to enhance the generalization capability of the metric network by enriching the sample-pair diversity. Moreover, we develop a Pair-Identity Center (PIC) loss to enhance the model’s discrimination by ensuring that sample-pair features with the same pair-identity are consistent. We validate the effectiveness of our proposed method through a lot of experiments on multiple benchmark datasets and confirm the value of each module in our GMN. Lei Qi 0001, Ziang Liu 0013, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | A Novel Cross-Perturbation for Single Domain GeneralizationabstractSingle domain generalization aims to enhance the ability of the model to generalize to unknown domains when trained on a single source domain. However, the limited diversity in the training data hampers the learning of domain-invariant features, resulting in compromised generalization performance. To address this, data perturbation (augmentation) has emerged as a crucial method to increase data diversity. Nevertheless, existing perturbation methods often focus on either image-level or feature-level perturbations independently, neglecting their synergistic effects. To overcome these limitations, we propose CPerb, a simple yet effective cross-perturbation method. Specifically, CPerb utilizes both horizontal and vertical operations. Horizontally, it applies image-level and feature-level perturbations to enhance the diversity of the training data, mitigating the issue of limited diversity in single-source domains. Vertically, it introduces multi-route perturbation to learn domain-invariant features from different perspectives of samples with the same semantic category, thereby enhancing the generalization capability of the model. Additionally, we propose MixPatch, a novel feature-level perturbation method that exploits local image style information to further diversify the training data. Extensive experiments on various benchmark datasets validate the effectiveness of our method. Dongjia Zhao, Lei Qi 0001, Xiao Shi 0001, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | SETA: Semantic-Aware Edge-Guided Token Augmentation for Domain GeneralizationabstractDomain generalization (DG) aims to enhance the model robustness against domain shifts without accessing target domains. A prevalent category of methods for DG is data augmentation, which focuses on generating virtual samples to simulate domain shifts. However, existing augmentation techniques in DG are mainly tailored for convolutional neural networks (CNNs), with limited exploration in token-based architectures, i.e., vision transformer (ViT) and multi-layer perceptrons (MLP) models. In this paper, we study the impact of prior CNN-based augmentation methods on token-based models, revealing their performance is suboptimal due to the lack of incentivizing the model to learn holistic shape information. To tackle the issue, we propose the Semantic-aware Edge-guided Token Augmentation (SETA) method. SETA transforms token features by perturbing local edge cues while preserving global shape features, thereby enhancing the model learning of shape information. To further enhance the generalization ability of the model, we introduce two stylized variants of our method combined with two state-of-the-art (SOTA) style augmentation methods in DG. We provide a theoretical insight into our method, demonstrating its effectiveness in reducing the generalization risk bound. Comprehensive experiments on five benchmarks prove that our method achieves SOTA performances across various ViT and MLP architectures. Our code is available at https://github.com/lingeringlight/SETA. Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | NormAUG: Normalization-Guided Augmentation for Domain GeneralizationabstractDeep learning has made significant advancements in supervised learning. However, models trained in this setting often face challenges due to domain shift between training and test sets, resulting in a significant drop in performance during testing. To address this issue, several domain generalization methods have been developed to learn robust and domain-invariant features from multiple training domains that can generalize well to unseen test domains. Data augmentation plays a crucial role in achieving this goal by enhancing the diversity of the training data. In this paper, inspired by the observation that normalizing an image with different statistics generated by different batches with various domains can perturb its feature, we propose a simple yet effective method called NormAUG (Normalization-guided Augmentation). Our method includes two paths: the main path and the auxiliary (augmented) path. During training, the auxiliary path includes multiple sub-paths, each corresponding to batch normalization for a single domain or a random combination of multiple domains. This introduces diverse information at the feature level and improves the generalization of the main path. Moreover, our NormAUG method effectively reduces the existing upper boundary for generalization based on theoretical perspectives. During the test stage, we leverage an ensemble strategy to combine the predictions from the auxiliary path of our model, further boosting performance. Extensive experiments are conducted on multiple benchmark datasets to validate the effectiveness of our proposed method. Lei Qi 0001, Hongpeng Yang, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Learning Generalizable Models via Disentangling Spurious and Enhancing Potential CorrelationsabstractDomain generalization (DG) intends to train a model on multiple source domains to ensure that it can generalize well to an arbitrary unseen target domain. The acquisition of domain-invariant representations is pivotal for DG as they possess the ability to capture the inherent semantic information of the data, mitigate the influence of domain shift, and enhance the generalization capability of the model. Adopting multiple perspectives, such as the sample and the feature, proves to be effective. The sample perspective facilitates data augmentation through data manipulation techniques, whereas the feature perspective enables the extraction of meaningful generalization features. In this paper, we focus on improving the generalization ability of the model by compelling it to acquire domain-invariant representations from both the sample and feature perspectives by disentangling spurious correlations and enhancing potential correlations. 1) From the sample perspective, we develop a frequency restriction module, guiding the model to focus on the relevant correlations between object features and labels, thereby disentangling spurious correlations. 2) From the feature perspective, the simple Tail Interaction module implicitly enhances potential correlations among all samples from all source domains, facilitating the acquisition of domain-invariant representations across multiple domains for the model. The experimental results show that Convolutional Neural Networks (CNNs) or Multi-Layer Perceptrons (MLPs) with a strong baseline embedded with these two modules can achieve superior results, e.g., an average accuracy of 92.30% on Digits-DG. Source code is available at https://github.com/RubyHoho/DGeneralization. Lei Qi 0001, Jintao Guo, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | ERNet: Edge Regularization Network for Cerebral Vessel Segmentation in Digital Subtraction Angiography ImagesabstractStroke is a leading cause of disability and fatality in the world, with ischemic stroke being the most common type. Digital Subtraction Angiography images, the gold standard in the operation process, can accurately show the contours and blood flow of cerebral vessels. The segmentation of cerebral vessels in DSA images can effectively help physicians assess the lesions. However, due to the disturbances in imaging parameters and changes in imaging scale, accurate cerebral vessel segmentation in DSA images is still a challenging task. In this paper, we propose a novel Edge Regularization Network (ERNet) to segment cerebral vessels in DSA images. Specifically, ERNet employs the erosion and dilation processes on the original binary vessel annotation to generate pseudo-ground truths of False Negative and False Positive, which serve as constraints to refine the coarse predictions based on their mapping relationship with the original vessels. In addition, we exploit a Hybrid Fusion Module based on convolution and transformers to extract local features and build long-range dependencies. Moreover, to support and advance the open research in the field of ischemic stroke, we introduce FPDSA, the first pixel-level semantic segmentation dataset for cerebral vessels. Extensive experiments on FPDSA illustrate the leading performance of our ERNet. Weijin Xu, Yinghuan Shi, Tao Tan 0002, Wentao Liu 0004, Xipeng Pan, Yiming Deng, Ruisheng Su |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Open-Domain Semi-Supervised Learning via Glocal Cluster Structure ExploitationabstractSemi-supervised learning (SSL) aims to reduce the heavy reliance of current deep models on costly manual annotation by leveraging a large amount of unlabeled data in combination with a much smaller set of labeled data. However, most existing SSL methods assume that all labeled and unlabeled data are drawn from the same feature distribution, which can be impractical in real-world applications. In this study, we take the initial step to systematically investigate the open-domain semi-supervised learning setting, where a feature distribution mismatch exists between labeled and unlabeled data. In pursuit of an effective solution for open-domain SSL, we propose a novel framework calledGlocalMatch, which aims to exploit bothglobal and local(i.e., glocal) cluster structure of open-domain unlabeled data. The glocal cluster structure is utilized in two complementary ways. First, GlocalMatch optimizes a Glocal Cluster Compacting (GCC) objective, that encourages feature representations of the same class, whether with in the same domain or across different domains, to become closer to each other. Second, GlocalMatch incorporates a Glocal Semantic Aggregation (GSA) strategy to produce more reliable pseudo-labels by aggregating predictions from neighboring clusters. Extensive experiments demonstrate that GlocalMatch outperforms the state-of-the-art SSL methods significantly, achieving superior performance for both in-domain and out-of-domain generalization. Zekun Li 0010, Lei Qi 0001, Yawen Li 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Exploring Flat Minima for Domain Generalization With Large Learning RatesabstractDomain Generalization (DG) aims to generalize to arbitrary unseen domains. A promising approach to improve model generalization in DG is the identification of flat minima. One typical method for this task is SWAD, which involves averaging weights along the training trajectory. However, the success of weight averaging depends on the diversity of weights, which is limited when training with a small learning rate. Instead, we observe that leveraging a large learning rate can simultaneously promote weight diversity and facilitate the identification of flat regions in the loss landscape. However, employing a large learning rate suffers from the convergence problem, which cannot be resolved by simply averaging the training weights. To address this issue, we introduce a training strategy called Lookahead which involves the weight interpolation, instead of average, between fast and slow weights. The fast weight explores the weight space with a large learning rate, which is not converged while the slow weight interpolates with it to ensure the convergence. Besides, weight interpolation also helps identify flat minima by implicitly optimizing the local entropy loss that measures flatness. To further prevent overfitting during training, we propose two variants to regularize the training weight with weighted averaged weight or with accumulated history weight. Taking advantage of this new perspective, our methods achieve state-of-the-art performance on both classification and semantic segmentation domain generalization benchmarks. The code is available athttps://github.com/koncle/DG-with-Large-LR. Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | MutexMatch: Semi-Supervised Learning With Mutex-Based Consistency RegularizationabstractThe core issue in semi-supervised learning (SSL) lies in how to effectively leverage unlabeled data, whereas most existing methods tend to put a great emphasis on the utilization of high-confidence samples yet seldom fully explore the usage of low-confidence samples. In this article, we aim to utilize low-confidence samples in a novel way with our proposed mutex-based consistency regularization, namely MutexMatch. Specifically, the high-confidence samples are required to exactly predict "what it is" by the conventional true-positive classifier (TPC), while low-confidence samples are employed to achieve a simpler goal-to predict with ease "what it is not" by the true-negative classifier (TNC). In this sense, we not only mitigate the pseudo-labeling errors but also make full use of the low-confidence unlabeled data by the consistency of dissimilarity degree. MutexMatch achieves superior performance on multiple benchmark datasets, i.e., Canadian Institute for Advanced Research (CIFAR)-10, CIFAR-100, street view house numbers (SVHN), self-taught learning 10 (STL-10), and mini-ImageNet. More importantly, our method further shows superiority when the amount of labeled data is scarce, e.g., 92.23% accuracy with only 20 labeled data on CIFAR-10. Code has been released at https://github.com/NJUyued/MutexMatch4SSL. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | PLACE Dropout: A Progressive Layer-wise and Channel-wise Dropout for Domain GeneralizationabstractDomain generalization (DG) aims to learn a generic model from multiple observed source domains that generalizes well to arbitrary unseen target domains without further training. The major challenge in DG is that the model inevitably faces a severe overfitting issue due to the domain gap between source and target domains. To mitigate this problem, some dropout-based methods have been proposed to resist overfitting by discarding part of the representation of the intermediate layers. However, we observe that most of these methods only conduct the dropout operation in some specific layers, leading to an insufficient regularization effect on the model. We argue that applying dropout at multiple layers can produce stronger regularization effects, which could alleviate the overfitting problem on source domains more adequately than previous layer-specific dropout methods. In this article, we develop a novel layer-wise and channel-wise dropout for DG, which randomly selects one layer and then randomly selects its channels to conduct dropout. Particularly, the proposed method can generate a variety of data variants to better deal with the overfitting issue. We also provide theoretical analysis for our dropout method and prove that it can effectively reduce the generalization error bound. Besides, we leverage the progressive scheme to increase the dropout ratio with the training progress, which can gradually boost the difficulty of training the model to enhance its robustness. Extensive experiments on three standard benchmark datasets have demonstrated that our method outperforms several state-of-the-art DG methods. Our code is available at https://github.com/lingeringlight/PLACEdropout . Jintao Guo, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | MultiMatch: Multi-task Learning for Semi-supervised Domain GeneralizationabstractDomain generalization (DG) aims at learning a model on source domains to well generalize on the unseen target domain. Although it has achieved great success, most of the existing methods require the label information for all training samples in source domains, which is time-consuming and expensive in the real-world application. In this article, we resort to solving the semi-supervised domain generalization (SSDG) task, where there are a few label information in each source domain. To address the task, we first analyze the theory of multi-domain learning, which highlights that (1) mitigating the impact of domain gap and (2) exploiting all samples to train the model can effectively reduce the generalization error in each source domain so as to improve the quality of pseudo-labels. According to the analysis, we propose MultiMatch, i.e., extending FixMatch to the multi-task learning framework, producing the high-quality pseudo-label for SSDG. To be specific, we consider each training domain as a single task (i.e., local task) and combine all training domains together (i.e., global task) to train an extra task for the unseen test domain. In the multi-task framework, we utilize the independent batch normalization and classifier for each task, which can effectively alleviate the interference from different domains during pseudo-labeling. Also, most of the parameters in the framework are shared, which can be trained by all training samples sufficiently. Moreover, to further boost the pseudo-label accuracy and the model’s generalization, we fuse the predictions from the global task and local task during training and testing, respectively. A series of experiments validate the effectiveness of the proposed method, and it outperforms the existing semi-supervised methods and the SSDG method on several benchmark DG datasets. Lei Qi 0001, Hongpeng Yang, Yinghuan Shi, Xin Geng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Orthogonal Annotation Benefits Barely-supervised Medical Image SegmentationabstractRecent trends in semi-supervised learning have significantly boosted the performance of 3D semi-supervised medical image segmentation. Compared with 2D images, 3D medical volumes involve information from different directions, e.g., transverse, sagittal, and coronal planes, so as to naturally provide complementary views. These complementary views and the intrinsic similarity among adjacent 3D slices inspire us to develop a novel annotation way and its corresponding semi-supervised model for effective segmentation. Specifically, we firstly propose the orthogonal annotation by only labeling two orthogonal slices in a labeled volume, which significantly relieves the burden of annotation. Then, we perform registration to obtain the initial pseudo labels for sparsely labeled volumes. Subsequently, by introducing unlabeled volumes, we propose a dual-network paradigm named Dense-Sparse Co-training (DeSCO) that exploits dense pseudo labels in early stage and sparse labels in later stage and meanwhile forces consistent output of two networks. Experimental results on three benchmark datasets validated our effectiveness in performance and efficiency in annotation. For example, with only 10 annotated slices, our method reaches a Dice up to 86.93% on KiTS19 dataset. Our code and models are available at https://github.com/HengCai-NJU/DeSCO. Heng Cai, Shumeng Li, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
CVPR | 5 |
| 2023 | ALOFT: A Lightweight MLP-Like Architecture with Dynamic Low-Frequency Transform for Domain GeneralizationabstractDomain generalization (DG) aims to learn a model that generalizes well to unseen target domains utilizing multiple source domains without re-training. Most existing DG works are based on convolutional neural networks (CNNs). However, the local operation of the convolution kernel makes the model focus too much on local representations (e.g., texture), which inherently causes the model more prone to overfit to the source domains and hampers its generalization ability. Recently, several MLP-based methods have achieved promising results in supervised learning tasks by learning global interactions among different patches of the image. Inspired by this, in this paper, we first analyze the difference between CNN and MLP methods in DG and find that MLP methods exhibit a better generalization ability because they can better capture the global representations (e.g., structure) than CNN methods. Then, based on a recent lightweight MLP method, we obtain a strong baseline that outperforms most state-of-the-art CNN-based methods. The baseline can learn global structure representations with a filter to suppress structureirrelevant information in the frequency space. Moreover, we propose a dynAmic LOw-Frequency spectrum Transform (ALOFT) that can perturb local texture features while preserving global structure features, thus enabling the filter to remove structure-irrelevant information sufficiently. Extensive experiments on four benchmarks have demonstrated that our method can achieve great performance improvement with a small number of parameters compared to SOTA CNN-based DG methods. Our code is available at https://github.com/lingeringlight/ALOFT/. Jintao Guo, Lei Qi 0001, Yinghuan Shi |
CVPR | 4 |
| 2023 | Revisiting Weak-to-Strong Consistency in Semi-Supervised Semantic SegmentationabstractIn this work, we revisit the weak-to-strong consistency framework, popularized by FixMatch from semi-supervised classification, where the prediction of a weakly perturbed image serves as supervision for its strongly perturbed version. Intriguingly, we observe that such a simple pipeline already achieves competitive results against recent advanced works, when transferred to our segmentation scenario. Its success heavily relies on the manual design of strong data augmentations, however, which may be limited and inadequate to explore a broader perturbation space. Motivated by this, we propose an auxiliary feature perturbation stream as a supplement, leading to an expanded perturbation space. On the other, to sufficiently probe original image-level augmentations, we present a dual-stream perturbation technique, enabling two strong views to be simultaneously guided by a common weak view. Consequently, our overall Unified Dual-Stream Perturbations approach (UniMatch) surpasses all existing methods significantly across all evaluation protocols on the Pascal, Cityscapes, and COCO benchmarks. Its superiority is also demonstrated in remote sensing interpretation and medical image analysis. We hope our reproduced FixMatch and our results can inspire more future works. Lihe Yang, Lei Qi 0001, Litong Feng, Wayne Zhang 0001, Yinghuan Shi |
CVPR | 5 |
| 2023 | DomainAdaptor: A Novel Approach to Test-time AdaptationabstractTo deal with the domain shift between training and test samples, current methods have primarily focused on learning generalizable features during training and ignore the specificity of unseen samples that are also critical during the test. In this paper, we investigate a more challenging task that aims to adapt a trained CNN model to unseen domains during the test. To maximumly mine the information in the test data, we propose a unified method called DomainAdaptor for the test-time adaptation, which consists of an AdaMixBN module and a Generalized Entropy Minimization (GEM) loss. Specifically, AdaMixBN addresses the domain shift by adaptively fusing training and test statistics in the normalization layer via a dynamic mixture co-efficient and a statistic transformation operation. To further enhance the adaptation ability of AdaMixBN, we design a GEM loss that extends the Entropy Minimization loss to better exploit the information in the test data. Extensive experiments show that DomainAdaptor consistently outperforms the state-of-the-art methods on four benchmarks. Furthermore, our method brings more remarkable improvement against existing methods on the few-data unseen domain. The code is available at https://github.com/koncle/DomainAdaptor. Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ICCV | 3 |
| 2023 | IOMatch: Simplifying Open-Set Semi-Supervised Learning with Joint Inliers and Outliers UtilizationabstractSemi-supervised learning (SSL) aims to leverage massive unlabeled data when labels are expensive to obtain. Unfortunately, in many real-world applications, the collected unlabeled data will inevitably contain unseen-class outliers not belonging to any of the labeled classes. To deal with the challenging open-set SSL task, the mainstream methods tend to first detect outliers and then filter them out. However, we observe a surprising fact that such approach could result in more severe performance degradation when labels are extremely scarce, as the unreliable outlier detector may wrongly exclude a considerable portion of valuable inliers. To tackle with this issue, we introduce a novel open-set SSL framework, IOMatch, which can jointly utilize inliers and outliers, even when it is difficult to distinguish exactly between them. Specifically, we propose to employ a multi-binary classifier in combination with the standard closed-set classifier for producing unified open-set classification targets, which regard all outliers as a single new class. By adopting these targets as open-set pseudo-labels, we optimize an open-set classifier with all unlabeled samples including both inliers and outliers. Extensive experiments have shown that IOMatch significantly outperforms the baseline methods across different benchmark datasets and different settings despite its remarkable simplicity. Our code and models are available at https://github.com/nukezil/IOMatch. Zekun Li 0010, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ICCV | 3 |
| 2023 | Towards Semi-supervised Learning with Non-random Missing LabelsabstractSemi-supervised learning (SSL) tackles the label missing problem by enabling the effective usage of unlabeled data. While existing SSL methods focus on the traditional setting, a practical and challenging scenario called label Missing Not At Random (MNAR) is usually ignored. In MNAR, the labeled and unlabeled data fall into different class distributions resulting in biased label imputation, which deteriorates the performance of SSL models. In this work, class transition tracking based Pseudo-Rectifying Guidance (PRG) is devised for MNAR. We explore the class-level guidance information obtained by the Markov random walk, which is modeled on a dynamically created graph built over the class tracking matrix. PRG unifies the historical information of class distribution and class transitions caused by the pseudo-rectifying procedure to maintain the model’s unbiased enthusiasm towards assigning pseudo-labels to all classes, so as the quality of pseudo-labels on both popular classes and rare classes in MNAR could be improved. Finally, we show the superior performance of PRG across a variety of MNAR scenarios, outperforming the latest SSL approaches combining bias removal solutions by a large margin. Code and model weights are available at https://github.com/NJUyued/PRG4SSL-MNAR. Yue Duan, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 6 |
| 2023 | Enhancing Sample Utilization through Sample Adaptive Augmentation in Semi-Supervised LearningabstractIn semi-supervised learning, unlabeled samples can be utilized through augmentation and consistency regularization. However, we observed certain samples, even undergoing strong augmentation, are still correctly classified with high confidence, resulting in a loss close to zero. It indicates that these samples have been already learned well and do not provide any additional optimization benefits to the model. We refer to these samples as "naive samples". Unfortunately, existing SSL models overlook the characteristics of naive samples, and they just apply the same learning strategy to all samples. To further optimize the SSL model, we emphasize the importance of giving attention to naive samples and augmenting them in a more diverse manner. Sample adaptive augmentation (SAA) is proposed for this stated purpose and consists of two modules: 1) sample selection module; 2) sample augmentation module. Specifically, the sample selection module picks out naive samples based on historical training information at each epoch, then the naive samples will be augmented in a more diverse manner in the sample augmentation module. Thanks to the extreme ease of implementation of the above modules, SAA is advantageous for being simple and lightweight. We add SAA on top of FixMatch and FlexMatch respectively, and experiments demonstrate SAA can significantly improve the models. For example, SAA helped improve the accuracy of FixMatch from 92.50% to 94.76% and that of FlexMatch from 95.01% to 95.31% on CIFAR-10 with 40 labels. The code is available at https://github.com/GuanGui-nju/SAA. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
ICCV | 6 |
| 2023 | DomainDrop: Suppressing Domain-Sensitive Channels for Domain GeneralizationabstractDeep Neural Networks have exhibited considerable success in various visual tasks. However, when applied to unseen test datasets, state-of-the-art models often suffer performance degradation due to domain shifts. In this paper, we introduce a novel approach for domain generalization from a novel perspective of enhancing the robustness of channels in feature maps to domain shifts. We observe that models trained on source domains contain a substantial number of channels that exhibit unstable activations across different domains, which are inclined to capture domain-specific features and behave abnormally when exposed to unseen target domains. To address the issue, we propose a DomainDrop framework to continuously enhance the channel robustness to domain shifts, where a domain discriminator is used to identify and drop unstable channels in feature maps of each network layer during forward propagation. We theoretically prove that our framework could effectively lower the generalization bound. Extensive experiments on several benchmarks indicate that our framework achieves state-of-the-art performance compared to other competing methods. Our code is available at https://github.com/lingeringlight/DomainDrop. Jintao Guo, Lei Qi 0001, Yinghuan Shi |
ICCV | 3 |
| 2023 | Generalizable Decision Boundaries: Dualistic Meta-Learning for Open Set Domain GeneralizationabstractDomain generalization (DG) is proposed to deal with the issue of domain shift, which occurs when statistical differences exist between source and target domains. However, most current methods do not account for a common realistic scenario where the source and target domains have different classes. To overcome this deficiency, open set domain generalization (OSDG) then emerges as a more practical setting to recognize unseen classes in unseen domains. An intuitive approach is to use multiple one-vs-all classifiers to define decision boundaries for each class and reject the outliers as unknown. However, the significant class imbalance between positive and negative samples often causes the boundaries biased towards positive ones, resulting in misclassification for known samples in the unseen target domain. In this paper, we propose a novel meta-learning-based framework called dualistic MEta-learning with joint DomaIn-Class matching (MEDIC), which considers gradient matching towards inter-domain and inter-class splits simultaneously to find a generalizable boundary balanced for all tasks. Experimental results demonstrate that MEDIC not only outperforms previous methods in open set scenarios, but also maintains competitive close set generalization ability at the same time. Our code is available at https://github.com/zzwdx/MEDIC. Xiran Wang, Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi |
ICCV | 4 |
| 2023 | Shrinking Class Space for Enhanced Certainty in Semi-Supervised LearningabstractSemi-supervised learning is attracting blooming attention, due to its success in combining unlabeled data. To mitigate potentially incorrect pseudo labels, recent frameworks mostly set a fixed confidence threshold to discard uncertain samples. This practice ensures high-quality pseudo labels, but incurs a relatively low utilization of the whole unlabeled set. In this work, our key insight is that these uncertain samples can be turned into certain ones, as long as the confusion classes for the top-1 class are detected and removed. Invoked by this, we propose a novel method dubbed ShrinkMatch to learn uncertain samples. For each uncertain sample, it adaptively seeks a shrunk class space, which merely contains the original top-1 class, as well as remaining less likely classes. Since the confusion ones are removed in this space, the re-calculated top-1 confidence can satisfy the pre-defined threshold. We then impose a consistency regularization between a pair of strongly and weakly augmented samples in the shrunk space to strive for discriminative representations. Furthermore, considering the varied reliability among uncertain samples and the gradually improved model during training, we correspondingly design two reweighting principles for our uncertain loss. Our method exhibits impressive performance on widely adopted benchmarks. Lihe Yang, Zhen Zhao 0001, Lei Qi 0001, Yu Qiao 0001, Yinghuan Shi, Hengshuang Zhao |
ICCV | 5 |
| 2023 | 3D Medical Image Segmentation with Sparse Annotation via Cross-Teaching Between 3D and 2D Networks
Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
MICCAI (3) | 4 |
| 2023 | FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsabstractSemantic segmentation has witnessed tremendous progress due to the proposal of various advanced network architectures. However, they are extremely hungry for delicate annotations to train, and the acquisition is laborious and unaffordable. Therefore, we present FreeMask in this work, which resorts to synthetic images from generative models to ease the burden of both data collection and annotation procedures. Concretely, we first synthesize abundant training images conditioned on the semantic masks provided by realistic datasets. This yields extra well-aligned image-mask training pairs for semantic segmentation models. We surprisingly observe that, solely trained with synthetic images, we already achieve comparable performance with real ones (e.g., 48.3 vs. 48.5 mIoU on ADE20K, and 49.3 vs. 50.5 on COCO-Stuff). Then, we investigate the role of synthetic images by joint training with real images, or pre-training for real images. Meantime, we design a robust filtering principle to suppress incorrectly synthesized regions. In addition, we propose to inequally treat different semantic masks to prioritize those harder ones and sample more corresponding synthetic images for them. As a result, either jointly trained or pre-trained with our filtered and re-sampled synthesized images, segmentation models can be greatly enhanced, e.g., from 48.7 to 52.0 on ADE20K. Lihe Yang, Xiaogang Xu 0002, Bingyi Kang, Yinghuan Shi, Hengshuang Zhao |
NeurIPS | 4 |
| 2023 | LibFewShot: A Comprehensive Library for Few-Shot LearningabstractFew-shot learning, especially few-shot image classification, has received increasing attention and witnessed significant advances in recent years. Some recent studies implicitly show that many generic techniques or "tricks", such as data augmentation, pre-training, knowledge distillation, and self-supervision, may greatly boost the performance of a few-shot learning method. Moreover, different works may employ different software platforms, backbone architectures and input image sizes, making fair comparisons difficult and practitioners struggle with reproducibility. To address these situations, we propose a comprehensive library for few-shot learning (LibFewShot) by re-implementing eighteen state-of-the-art few-shot learning methods in a unified framework with the same single codebase in PyTorch. Furthermore, based on LibFewShot, we provide comprehensive evaluations on multiple benchmarks with various backbone architectures to evaluate common pitfalls and effects of different training tricks. In addition, with respect to the recent doubts on the necessity of meta- or episodic-training mechanism, our evaluation results confirm that such a mechanism is still necessary especially when combined with pre-training. We hope our work can not only lower the barriers for beginners to enter the area of few-shot learning but also elucidate the effects of nontrivial tricks to facilitate intrinsic research on few-shot learning. Wenbin Li 0006, Xuesong Yang, Chuanqi Dong, Pinzhuo Tian, Tiexin Qin, Jing Huo, Yinghuan Shi, Lei Wang 0001, Yang Gao 0001, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Unsupervised generalizable multi-source person re-identification: A Domain-specific adaptive framework
Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Xin Geng 0001 |
Pattern Recognit. | 4 |
| 2023 | Better pseudo-label: Joint domain-aware label and dual-classifier for semi-supervised domain generalization
Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
Pattern Recognit. | 3 |
| 2023 | PLN: Parasitic-Like Network for Barely Supervised Medical Image SegmentationabstractIt is known that annotations for 3D medical image segmentation tasks are laborious, time-consuming and expensive. Considering the similarities existing in inter-slice and inter-volume, we believe that the delineation way and the model architecture should be tightly coupled. In this paper, by introducing an extremely sparse annotation way of labeling only one slice per 3D image, we investigate a novel barely-supervised segmentation setting with only a few sparsely-labeled images along with a large amount of unlabeled images. To achieve this goal, we present a new parasitic-like network including a registration module (as host) and a semi-supervised segmentation module (as parasite) to deal with inter-slice label propagation and inter-volume segmentation prediction, respectively. Specifically, our parasitism mechanism effectively achieves the collaboration of these two modules through three stages of infection, development and eclosion, providing accurate pseudo-labels for training. Extensive results demonstrate that our framework is capable of achieving high performance on extremely sparse annotation tasks, e.g., we achieve Dice of 84.83% on LA dataset with only 16 labeled slices. The code is available athttps://github.com/ShumengLI/PLN. Shumeng Li, Heng Cai, Lei Qi 0001, Qian Yu 0007, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | A Novel Mix-Normalization Method for Generalizable Multi-Source Person Re-IdentificationabstractPerson re-identification (Re-ID) has achieved great success in the supervised scenario. However, it is difficult to directly transfer the supervised model to arbitrary unseen domains due to the model overfitting to the seen source domains. In this paper, we aim to tackle the generalizable multi-source person Re-ID task (i.e., there are multiple available source domains, and the testing domain is unseen during training) from the data augmentation perspective, thus we put forward a novel method, termed MixNorm. It consists of domain-aware mix-normalization (DMN) and domain-aware center regularization (DCR). Different from the conventional data augmentation, the proposed domain-aware mix-normalization enhances the diversity of features during training from the normalization perspective of the neural network, which can effectively alleviate the model overfitting to the source domains, so as to boost the generalization capability of the model in the unseen domain. To further promote the efficacy of the proposed DMN, we exploit the domain-aware center regularization to better map the diversely generated features into the same space. Extensive experiments on multiple benchmark datasets validate the effectiveness of the proposed method and show that the proposed method can outperform the state-of-the-art methods. Besides, further analysis also reveals the superiority of the proposed method. Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | A Multilayer Framework for Online Metric LearningabstractOnline metric learning (OML) has been widely applied in classification and retrieval. It can automatically learn a suitable metric from data by restricting similar instances to be separated from dissimilar instances with a given margin. However, the existing OML algorithms have limited performance in real-world classifications, especially, when data distributions are complex. To this end, this article proposes a multilayer framework for OML to capture the nonlinear similarities among instances. Different from the traditional OML, which can only learn one metric space, the proposed multilayer OML (MLOML) takes an OML algorithm as a metric layer and learns multiple hierarchical metric spaces, where each metric layer follows a nonlinear layer for the complicated data distribution. Moreover, the forward propagation (FP) strategy and backward propagation (BP) strategy are employed to train the hierarchical metric layers. To build a metric layer of the proposed MLOML, a new Mahalanobis-based OML (MOML) algorithm is presented based on the passive-aggressive strategy and one-pass triplet construction strategy. Furthermore, in a progressively and nonlinearly learning way, MLOML has a stronger learning ability than traditional OML in the case of limited available training data. To make the learning process more explainable and theoretically guaranteed, theoretical analysis is provided. The proposed MLOML enjoys several nice properties, indeed learns a metric progressively, and performs better on the benchmark datasets. Extensive experiments with different settings have been conducted to verify these properties of the proposed MLOML. Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | LaSSL: Label-Guided Self-Training for Semi-supervised LearningabstractThe key to semi-supervised learning (SSL) is to explore adequate information to leverage the unlabeled data. Current dominant approaches aim to generate pseudo-labels on weakly augmented instances and train models on their corresponding strongly augmented variants with high-confidence results. However, such methods are limited in excluding samples with low-confidence pseudo-labels and under-utilization of the label information. In this paper, we emphasize the cruciality of the label information and propose a Label-guided Self-training approach to Semi-supervised Learning (LaSSL), which improves pseudo-label generations from two mutually boosted strategies. First, with the ground-truth labels and iteratively-polished pseudo-labels, we explore instance relations among all samples and then minimize a class-aware contrastive loss to learn discriminative feature representations that make same-class samples gathered and different-class samples scattered. Second, on top of improved feature representations, we propagate the label information to the unlabeled samples across the potential data manifold at the feature-embedding level, which can further improve the labelling of samples with reference to their neighbours. These two strategies are seamlessly integrated and mutually promoted across the whole training process. We evaluate LaSSL on several classification benchmarks under partially labeled settings and demonstrate its superiority over the state-of-the-art approaches. Zhen Zhao 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 4 |
| 2022 | ST++: Make Self-trainingWork Better for Semi-supervised Semantic SegmentationabstractSelf-training via pseudo labeling is a conventional, simple, and popular pipeline to leverage unlabeled data. In this work, we first construct a strong baseline of self-training (namely ST) for semi-supervised semantic segmentation via injecting strong data augmentations (SDA) on unlabeled images to alleviate overfitting noisy labels as well as decouple similar predictions between the teacher and student. With this simple mechanism, our ST outperforms all existing methods without any bells and whistles, e.g., iterative retraining. Inspired by the impressive results, we thoroughly investigate the SDA and provide some empirical analysis. Nevertheless, incorrect pseudo labels are still prone to accumulate and degrade the performance. To this end, we further propose an advanced self-training framework (namely ST++), that performs selective re-training via prioritizing reliable unlabeled images based on holistic prediction-level stability. Concretely, several model checkpoints are saved in the first stage supervised training, and the discrepancy of their predictions on the unlabeled image serves as a measurement for reliability. Our image-level selection offers holistic contextual information for learning. We demonstrate that it is more suitable for segmentation than common pixel-wise selection. As a result, ST+ further boosts the performance of our ST. Code is available at https://github.com/LiheYoung/ST-PlusPlus. Lihe Yang, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
CVPR | 4 |
| 2022 | DC-SSL: Addressing Mismatched Class Distribution in Semi-supervised LearningabstractConsistency-based Semi-supervised learning (SSL) has achieved promising performance recently. However, the success largely depends on the assumption that the labeled and unlabeled data share an identical class distribution, which is hard to meet in real practice. The distribution mismatch between the labeled and unlabeled sets can cause severe bias in the pseudo-labels of SSL, resulting in significant performance degradation. To bridge this gap, we put forward a new SSL learning framework, named Distribution Consistency SSL (DC-SSL), which rectifies the pseudolabels from a distribution perspective. The basic idea is to directly estimate a reference class distribution (RCD), which is regarded as a surrogate of the ground truth class distribution about the unlabeled data, and then improve the pseudo-labels by encouraging the predicted class distribution (PCD) of the unlabeled data to approach RCD gradually. To this end, this paper revisits the Exponentially Moving Average (EMA) model and utilizes it to estimate RCD in an iteratively improved manner, which is achieved with a momentum-update scheme throughout the training procedure. On top of this, two strategies are proposed for RCD to rectify the pseudo-label prediction, respectively. They correspond to an efficient training-free scheme and a training-based alternative that generates more accurate and reliable predictions. DC-SSL is evaluated on multiple SSL benchmarks and demonstrates remarkable performance improvement over competitive methods under matched- and mismatched-distribution scenarios. Zhen Zhao 0001, Luping Zhou, Yue Duan, Lei Wang 0001, Lei Qi 0001, Yinghuan Shi |
CVPR | 6 |
| 2022 | Generalizable Cross-modality Medical Image Segmentation via Style Augmentation and Dual NormalizationabstractFor medical image segmentation, imagine if a model was only trained using MR images in source domain, how about its performance to directly segment CT images in target domain? This setting, namely generalizable cross-modality segmentation, owning its clinical potential, is much more challenging than other related settings, e.g., domain adaptation. To achieve this goal, we in this paper propose a novel dual-normalization model by leveraging the augmented source-similar and source-dissimilar images during our generalizable segmentation. To be specific, given a single source domain, aiming to simulate the possible appearance change in unseen target domains, we first utilize a nonlinear transformation to augment source-similar and source-dissimilar images. Then, to sufficiently exploit these two types of augmentations, our proposed dualnormalization based model employs a shared backbone yet independent batch normalization layer for separate normalization. Afterward, we put forward a style-based selection scheme to automatically choose the appropriate path in the test stage. Extensive experiments on three publicly available datasets, i.e., BraTS, Cross-Modality Cardiac, and Abdominal Multi-Organ datasets, have demonstrated that our method outperforms other state-of-the-art domain generalization methods. Code is available at https://github.com/zzzqzhou/Dual-Normalization. Lei Qi 0001, Xin Yang 0009, Dong Ni 0001, Yinghuan Shi |
CVPR | 5 |
| 2022 | RDA: Reciprocal Distribution Alignment for Robust Semi-supervised Learning
Yue Duan, Lei Qi 0001, Lei Wang 0001, Luping Zhou, Yinghuan Shi |
ECCV (30) | 5 |
| 2022 | MVDG: A Unified Multi-view Framework for Domain Generalization
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ECCV (27) | 3 |
| 2022 | Generalizable Medical Image Segmentation via Random Amplitude Mixup and Domain-Specific Image Restoration
Lei Qi 0001, Yinghuan Shi |
ECCV (21) | 3 |
| 2022 | Improving Barely Supervised Learning by Discriminating Unlabeled Samples with Super-ClassabstractIn semi-supervised learning (SSL), a common practice is to learn consistent information from unlabeled data and discriminative information from labeled data to ensure both the immutability and the separability of the classification model. Existing SSL methods suffer from failures in barely-supervised learning (BSL), where only one or two labels per class are available, as the insufficient labels cause the discriminative information being difficult or even infeasible to learn. To bridge this gap, we investigate a simple yet effective way to leverage unlabeled samples for discriminative learning, and propose a novel discriminative information learning module to benefit model training. Specifically, we formulate the learning objective of discriminative information at the super-class level and dynamically assign different classes into different super-classes based on model performance improvement. On top of this on-the-fly process, we further propose a distribution-based loss to learn discriminative information by utilizing the similarity relationship between samples and super-classes. It encourages the unlabeled samples to stay closer to the distribution of their corresponding super-class than those of others. Such a constraint is softer than the direct assignment of pseudo labels, while the latter could be very noisy in BSL. We compare our method with state-of-the-art SSL and BSL methods through extensive experiments on standard SSL benchmarks. Our method can achieve superior results, \eg, an average accuracy of 76.76\% on CIFAR-10 with merely 1 label per class. Guan Gui 0002, Zhen Zhao 0001, Lei Qi 0001, Luping Zhou, Lei Wang 0001, Yinghuan Shi |
NeurIPS | 6 |
| 2022 | Generalizable model-agnostic semantic segmentation via target-specific normalization
Jian Zhang 0090, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
Pattern Recognit. | 3 |
| 2022 | Adversarial Camera Alignment Network for Unsupervised Cross-Camera Person Re-IdentificationabstractIn person re-identification (Re-ID), supervised methods usually need a large amount of expensive label information, while unsupervised ones are still unable to deliver satisfactory identification performance. In this paper, we introduce a novel person Re-ID task called unsupervised cross-camera person Re-ID, which only needs the within-camera (intra-camera) label information but not cross-camera (inter-camera) labels which are more expensive to obtain. In real-world applications, the intra-camera label information can be easily captured by tracking algorithms and few manual annotations. In this situation, the main challenge becomes the distribution discrepancy across different camera views, caused by the various body pose, occlusion, image resolution, illumination conditions, and background noises in different cameras. To address this situation, we propose a novel Adversarial Camera Alignment Network (ACAN) for unsupervised cross-camera person Re-ID. It consists of the camera-alignment task and the supervised within-camera learning task. To achieve the camera alignment, we develop a Multi-Camera Adversarial Learning (MCAL) to map images of different cameras into a shared subspace. Particularly, we investigate two different schemes, including the existing GRL (i.e., gradient reversal layer) scheme and the proposed scheme called “other camera equiprobability” (OCE), to conduct the multi-camera adversarial task. Based on this shared subspace, we then leverage the within-camera labels to train the network. Extensive experiments on five large-scale datasets demonstrate the superiority of ACAN over multiple state-of-the-art unsupervised methods that take advantage of labeled source domains and generated images by GAN-based models. In particular, we verify that the proposed multi-camera adversarial task does contribute to the significant improvement. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Xin Geng 0001, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Feature-Based Style Randomization for Domain GeneralizationabstractAs a recent noticeable topic, domain generalization (DG) aims to first learn a generic model on multiple source domains and then directly generalize to an arbitrary unseen target domain without any additional adaption. In previous DG models, by generating virtual data to supplement observed source domains, the data augmentation based methods have shown its effectiveness. To simulate the possible unseen domains, most of them enrich the diversity of original data via image-level style transformation. However, we argue that the potential styles are hard to be exhaustively illustrated and fully augmented due to the limited referred styles, leading the diversity could not be always guaranteed. Unlike image-level augmentation, we in this paper develop a simple yet effective feature-based style randomization module to achieve feature-level augmentation, which can produce random styles via integrating random noise into the original style. Compared with existing image-level augmentation, our feature-level augmentation favors a more goal-oriented and sample-diverse way. Furthermore, to sufficiently explore the efficacy of the proposed module, we design a novel progressive training strategy to enable all parameters of the network to be fully trained. Extensive experiments on three standard benchmark datasets,i.e., PACS, VLCS and Office-Home, highlight the superiority of our method compared to the state-of-the-art methods. Yue Wang 0076, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Label Distribution Learning for Generalizable Multisource Person Re-IdentificationabstractPerson re-identification (Re-ID) is a critical technique in the video surveillance system, which has achieved significant success in the supervised setting. However, it is difficult to directly apply the supervised model to arbitrary unseen domains due to the domain gap between the available source domains and unseen target domains. In this paper, we propose a novel label distribution learning (LDL) method to address generalizable multi-source person Re-ID task (i.e., there are multiple available source domains, and the testing domain is unseen during training), which aims to explore the relation of different classes and mitigate the domain-shift across different domains so as to improve the discrimination of the model and learn the domain-invariant feature, simultaneously. Specifically, during the training process, we produce the label distribution via the online manner to mine the relation information of different classes, thus it is beneficial for extracting the discriminative feature. Besides, for the label distribution of each class, we further revise it to give more and equal attention to the other domains that the class does not belong to, which can effectively reduce the domain gap across different domains and obtain the domain-invariant feature. Furthermore, we also give the theoretical analysis to demonstrate that the proposed method can effectively deal with the domain-shift issue. Extensive experiments on multiple benchmark datasets validate the effectiveness of the proposed method and show that the proposed method can outperform the state-of-the-art methods. Besides, further analysis also reveals the superiority of the proposed method. Lei Qi 0001, Jiaying Shen, Yinghuan Shi, Xin Geng 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Video Super-Resolution via a Spatio-Temporal Alignment NetworkabstractDeep convolutional neural network based video super-resolution (SR) models have achieved significant progress in recent years. Existing deep video SR methods usually impose optical flow to wrap the neighboring frames for temporal alignment. However, accurate estimation of optical flow is quite difficult, which tends to produce artifacts in the super-resolved results. To address this problem, we propose a novel end-to-end deep convolutional network that dynamically generates the spatially adaptive filters for the alignment, which are constituted by the local spatio-temporal channels of each pixel. Our method avoids generating explicit motion compensation and utilizes spatio-temporal adaptive filters to achieve the operation of alignment, which effectively fuses the multi-frame information and improves the temporal consistency of the video. Capitalizing on the proposed adaptive filter, we develop a reconstruction network and take the aligned frames as input to restore the high-resolution frames. In addition, we employ residual modules embedded with channel attention as the basic unit to extract more informative features for video SR. Both quantitative and qualitative evaluation results on three public video datasets demonstrate that the proposed method performs favorably against state-of-the-art super-resolution methods in terms of clearness and texture details. Weilei Wen, Wenqi Ren, Yinghuan Shi, Yunfeng Nie, Jingang Zhang, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2022 | Crosslink-Net: Double-Branch Encoder Network via Fusing Vertical and Horizontal Convolutions for Medical Image SegmentationabstractAccurate image segmentation plays a crucial role in medical image analysis, yet it faces great challenges caused by various shapes, diverse sizes, and blurry boundaries. To address these difficulties, square kernel-based encoder-decoder architectures have been proposed and widely used, but their performance remains unsatisfactory. To further address these challenges, we present a novel double-branch encoder architecture. Our architecture is inspired by two observations. (1) Since the discrimination of the features learned via square convolutional kernels needs to be further improved, we propose utilizing nonsquare vertical and horizontal convolutional kernels in a double-branch encoder so that the features learned by both branches can be expected to complement each other. (2) Considering that spatial attention can help models to better focus on the target region in a large-sized image, we develop an attention loss to further emphasize the segmentation of small-sized targets. With the above two schemes, we develop a novel double-branch encoder-based segmentation framework for medical image segmentation, namely, Crosslink-Net, and validate its effectiveness on five datasets with experiments. The code is released at https://github.com/Qianyu1226/Crosslink-Net. Qian Yu 0007, Lei Qi 0001, Yang Gao 0001, Wuzhang Wang, Yinghuan Shi |
IEEE Trans. Image Process. | 5 |
| 2022 | Deep Symmetric Adaptation Network for Cross-Modality Medical Image SegmentationabstractUnsupervised domain adaptation (UDA) methods have shown their promising performance in the cross-modality medical image segmentation tasks. These typical methods usually utilize a translation network to transform images from the source domain to target domain or train the pixel-level classifier merely using translated source images and original target images. However, when there exists a large domain shift between source and target domains, we argue that this asymmetric structure, to some extent, could not fully eliminate the domain gap. In this paper, we present a novel deep symmetric architecture of UDA for medical image segmentation, which consists of a segmentation sub-network, and two symmetric source and target domain translation sub-networks. To be specific, based on two translation sub-networks, we introduce a bidirectional alignment scheme via a shared encoder and two private decoders to simultaneously align features 1) from source to target domain and 2) from target to source domain, which is able to effectively mitigate the discrepancy between domains. Furthermore, for the segmentation sub-network, we train a pixel-level classifier using not only original target images and translated source images, but also original source images and translated target images, which could sufficiently leverage the semantic information from the images with different styles. Extensive experiments demonstrate that our method has remarkable advantages compared to the state-of-the-art methods in three segmentation tasks, such as cross-modality cardiac, BraTS, and abdominal multi-organ segmentation. Xiaoting Han, Lei Qi 0001, Qian Yu 0007, Yefeng Zheng 0001, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Inconsistency-Aware Uncertainty Estimation for Semi-Supervised Medical Image SegmentationabstractIn semi-supervised medical image segmentation, most previous works draw on the common assumption that higher entropy means higher uncertainty. In this paper, we investigate a novel method of estimating uncertainty. We observe that, when assigned different misclassification costs in a certain degree, if the segmentation result of a pixel becomes inconsistent, this pixel shows a relative uncertainty in its segmentation. Therefore, we present a new semi-supervised segmentation model, namely, conservative-radical network (CoraNet in short) based on our uncertainty estimation and separate self-training strategy. In particular, our CoraNet model consists of three major components: a conservative-radical module (CRM), a certain region segmentation network (C-SN), and an uncertain region segmentation network (UC-SN) that could be alternatively trained in an end-to-end manner. We have extensively evaluated our method on various segmentation tasks with publicly available benchmark datasets, including CT pancreas, MR endocardium, and MR multi-structures segmentation on the ACDC dataset. Compared with the current state of the art, our CoraNet has demonstrated superior performance. In addition, we have also analyzed its connection with and difference from conventional methods of uncertainty estimation in semi-supervised medical image segmentation. Yinghuan Shi, Jian Zhang 0090, Tong Ling, Jiwen Lu, Yefeng Zheng 0001, Qian Yu 0007, Lei Qi 0001, Yang Gao 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Manifold Alignment for Semantically Aligned Style TransferabstractMost existing style transfer methods follow the assumption that styles can be represented with global statistics (e.g., Gram matrices or covariance matrices), and thus address the problem by forcing the output and style images to have similar global statistics. An alternative is the assumption of local style patterns, where algorithms are designed to swap similar local features of content and style images. However, the limitation of these existing methods is that they neglect the semantic structure of the content image which may lead to corrupted content structure in the output. In this paper, we make a new assumption that image features from the same semantic region form a manifold and an image with multiple semantic regions follows a multi-manifold distribution. Based on this assumption, the style transfer problem is formulated as aligning two multi-manifold distributions and a Manifold Alignment based Style Transfer (MAST) framework is proposed. The proposed frame-work allows semantically similar regions between the output and the style image share similar style patterns. Moreover, the proposed manifold alignment method is flexible to allow user editing or using semantic segmentation maps as guidance for style transfer. To allow the method to be applicable to photorealistic style transfer, we propose a new adaptive weight skip connection network structure to preserve the content details. Extensive experiments verify the effectiveness of the proposed framework for both artistic and photorealistic style transfer. Code is available at https://github.com/NJUHuoJing/MAST. Jing Huo, Shiyin Jin, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yinghuan Shi, Yang Gao 0001 |
ICCV | 6 |
| 2021 | Mining Latent Classes for Few-shot SegmentationabstractFew-shot segmentation (FSS) aims to segment unseen classes given only a few annotated samples. Existing methods suffer the problem of feature undermining, i.e., potential novel classes are treated as background during training phase. Our method aims to alleviate this problem and enhance the feature embedding on latent novel classes. In our work, we propose a novel joint-training framework. Based on conventional episodic training on support-query pairs, we introduce an additional mining branch that exploits latent novel classes via transferable sub-clusters, and a new rectification technique on both background and fore-ground categories to enforce more stable prototypes. Over and above that, our transferable sub-cluster has the ability to leverage extra unlabeled data for further feature enhancement. Extensive experiments on two FSS benchmarks demonstrate that our method outperforms previous state-of-the-art by a large margin of 3.7% mIOU on PASCAL-5iand 7.0% mIOU on COCO-20iat the cost of 74% fewer parameters and 2.5x faster inference speed. The source code is available at https://github.com/LiheYoung/MiningFSS. Lihe Yang, Lei Qi 0001, Yinghuan Shi, Yang Gao 0001 |
ICCV | 4 |
| 2021 | Interactive medical image segmentation via a point-based interaction
Jian Zhang 0090, Yinghuan Shi, Jinquan Sun, Lei Wang 0001, Luping Zhou, Yang Gao 0001, Dinggang Shen |
Artif. Intell. Medicine | 2 |
| 2021 | A novel multiple instance learning framework for COVID-19 severity assessment via data augmentation and self-supervised learning
Zekun Li 0010, Wei Zhao 0040, Feng Shi 0001, Lei Qi 0001, Xingzhi Xie, Ying Wei 0009, Zhongxiang Ding, Yang Gao 0001, Shangjie Wu, Jun Liu 0075, Yinghuan Shi, Dinggang Shen |
Medical Image Anal. | 11 |
| 2021 | Synergistic learning of lung lobe segmentation and hierarchical multi-instance classification for automated severity assessment of COVID-19 in CT images
Kelei He, Wei Zhao 0040, Xingzhi Xie, Mingxia Liu 0001, Zhenyu Tang 0002, Yinghuan Shi, Feng Shi 0001, Yang Gao 0001, Jun Liu 0075, Dinggang Shen |
Pattern Recognit. | 7 |
| 2021 | Crossover-Net: Leveraging vertical-horizontal crossover relation for robust medical image segmentation
Qian Yu 0007, Yang Gao 0001, Yefeng Zheng 0001, Jianbing Zhu, Yakang Dai, Yinghuan Shi |
Pattern Recognit. | 6 |
| 2021 | Learning-Based Computer-Aided Prescription Model for Parkinson's Disease: A Data-Driven PerspectiveabstractIn this article, we study a novel problem: "automatic prescription recommendation for PD patients." To realize this goal, we first build a dataset by collecting 1) symptoms of PD patients, and 2) their prescription drug provided by neurologists. Then, we build a novel computer-aided prescription model by learning the relation between observed symptoms and prescription drug. Finally, for the new coming patients, we could recommend (predict) suitable prescription drug on their observed symptoms by our prescription model. From the methodology part, our proposed model, namely Prescription viA Learning lAtent Symptoms (PALAS), could recommend prescription using the multi-modality representation of the data. In PALAS, a latent symptom space is learned to better model the relationship between symptoms and prescription drug, as there is a large semantic gap between them. Moreover, we present an efficient alternating optimization method for PALAS. We evaluated our method using the data collected from 136 PD patients at Nanjing Brain Hospital, which can be regarded as a large dataset in PD research community. The experimental results demonstrate the effectiveness and clinical potential of our method in this recommendation task, if compared with other competing methods. Yinghuan Shi, Wanqi Yang, Kim-Han Thung, Hao Wang 0013, Yang Gao 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | GreyReID: A Novel Two-stream Deep Framework with RGB-grey Information for Person Re-identificationabstractIn this article, we observe that most false positive images (i.e., different identities with query images) in the top ranking list usually have the similar color information with the query image in person re-identification (Re-ID). Meanwhile, when we use the greyscale images generated from RGB images to conduct the person Re-ID task, some hard query images can obtain better performance compared with using RGB images. Therefore, RGB and greyscale images seem to be complementary to each other for person Re-ID. In this article, we aim to utilize both RGB and greyscale images to improve the person Re-ID performance. To this end, we propose a novel two-stream deep neural network with RGB-grey information, which can effectively fuse RGB and greyscale feature representations to enhance the generalization ability of Re-ID. First, we convert RGB images to greyscale images in each training batch. Based on these RGB and greyscale images, we train the RGB and greyscale branches, respectively. Second, to build up connections between RGB and greyscale branches, we merge the RGB and greyscale branches into a new joint branch. Finally, we concatenate the features of all three branches as the final feature representation for Re-ID. Moreover, in the training process, we adopt the joint learning scheme to simultaneously train each branch by the independent loss function, which can enhance the generalization ability of each branch. Besides, a global loss function is utilized to further fine-tune the final concatenated feature. The extensive experiments on multiple benchmark datasets fully show that the proposed method can outperform the state-of-the-art person Re-ID methods. Furthermore, using greyscale images can indeed improve the person Re-ID performance in the proposed deep framework. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Differentiable Meta-Learning Model for Few-Shot Semantic SegmentationabstractTo address the annotation scarcity issue in some cases of semantic segmentation, there have been a few attempts to develop the segmentation model in the few-shot learning paradigm. However, most existing methods only focus on the traditional 1-way segmentation setting (i.e., one image only contains a single object). This is far away from practical semantic segmentation tasks where the K-way setting (K > 1) is usually required by performing the accurate multi-object segmentation. To deal with this issue, we formulate the few-shot semantic segmentation task as a learning-based pixel classification problem, and propose a novel framework called MetaSegNet based on meta-learning. In MetaSegNet, an architecture of embedding module consisting of the global and local feature branches is developed to extract the appropriate meta-knowledge for the few-shot segmentation. Moreover, we incorporate a linear model into MetaSegNet as a base learner to directly predict the label of each pixel for the multi-object segmentation. Furthermore, our MetaSegNet can be trained by the episodic training mechanism in an end-to-end manner from scratch. Experiments on two popular semantic segmentation datasets, i.e., PASCAL VOC and COCO, reveal the effectiveness of the proposed MetaSegNet in the K-way few-shot semantic segmentation task. Pinzhuo Tian, Zhangkai Wu, Lei Qi 0001, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
AAAI | 5 |
| 2020 | Automatic Data Augmentation Via Deep Reinforcement Learning for Effective Kidney Tumor SegmentationabstractConventional data augmentation realized by performing simple pre-processing operations (e.g., rotation, crop, etc.) has been validated for its advantage in enhancing the performance for medical image segmentation. However, the data generated by these conventional augmentation methods are random and sometimes harmful to the subsequent segmentation. In this paper, we developed a novel automatic learning-based data augmentation method for medical image segmentation which models the augmentation task as a trial-and-error procedure using deep reinforcement learning (DRL). In our method, we innovatively combine the data augmentation module and the subsequent segmentation module in an end-to-end training manner with a consistent loss. Specifically, the best sequential combination of different basic operations is automatically learned by directly maximizing the performance improvement (i.e., Dice ratio) on the available validation set. We extensively evaluated our method on CT kidney tumor segmentation which validated the promising results of our method. Tiexin Qin, Kelei He, Yinghuan Shi, Yang Gao 0001, Dinggang Shen |
ICASSP | 4 |
| 2020 | Asymmetric Distribution Measure for Few-shot LearningabstractThe core idea of metric-based few-shot image classification is to directly measure the relations between query images and support classes to learn transferable feature embeddings. Previous work mainly focuses on image-level feature representations, which actually cannot effectively estimate a class's distribution due to the scarcity of samples. Some recent work shows that local descriptor based representations can achieve richer representations than image-level based representations. However, such works are still based on a less effective instance-level metric, especially a symmetric metric, to measure the relation between a query image and a support class. Given the natural asymmetric relation between a query image and a support class, we argue that an asymmetric measure is more suitable for metric-based few-shot learning. To that end, we propose a novel Asymmetric Distribution Measure (ADM) network for few-shot learning by calculating a joint local and global asymmetric measure between two multivariate local distributions of a query and a class. Moreover, a task-aware Contrastive Measure Strategy (CMS) is proposed to further enhance the measure function. On popular miniImageNet and tieredImageNet, ADM can achieve the state-of-the-art results, validating our innovative design of asymmetric distribution measures for few-shot learning. The source code can be downloaded from https://github.com/WenbinLee/ADM.git. Wenbin Li 0006, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001, Jiebo Luo 0001 |
IJCAI | 4 |
| 2020 | Consistent MetaReg: Alleviating Intra-task Discrepancy for Better Meta-knowledgeabstractIn the few-shot learning scenario, the data-distribution discrepancy between training data and test data in a task usually exists due to the limited data. However, most existing meta-learning approaches seldom consider this intra-task discrepancy in the meta-training phase which might deteriorate the performance. To overcome this limitation, we develop a new consistent meta-regularization method to reduce the intra-task data-distribution discrepancy. Moreover, the proposed meta-regularization method could be readily inserted into existing optimization-based meta-learning models to learn better meta-knowledge. Particularly, we provide the theoretical analysis to prove that using the proposed meta-regularization, the conventional gradient-based meta-learning method can reach the lower regret bound. The extensive experiments also demonstrate the effectiveness of our method, which indeed improves the performances of the state-of-the-art gradient-based meta-learning models in the few-shot classification task. Pinzhuo Tian, Lei Qi 0001, Shaokang Dong, Yinghuan Shi, Yang Gao 0001 |
IJCAI | 4 |
| 2020 | Progressive Cross-Camera Soft-Label Learning for Semi-Supervised Person Re-IdentificationabstractIn this paper, we focus on the semi-supervised person re-identification (Re-ID) case, which only has the intra-camera (within-camera) labels but not inter-camera (cross-camera) labels. In real-world applications, these intra-camera labels can be readily captured by tracking algorithms or few manual annotations, when compared with cross-camera labels. In this case, it is very difficult to explore the relationships between cross-camera persons in the training stage due to the lack of cross-camera label information. To deal with this issue, we propose a novel Progressive Cross-camera Soft-label Learning (PCSL) framework for the semi-supervised person Re-ID task, which can generate cross-camera soft-labels and utilize them to optimize the network. Concretely, we calculate an affinity matrix based on person-level features and adapt them to produce the similarities between cross-camera persons (i.e., cross-camera soft-labels). To exploit these soft-labels to train the network, we investigate the weighted cross-entropy loss and the weighted triplet loss from the classification and discrimination perspectives, respectively. Particularly, the proposed framework alternately generates progressive cross-camera soft-labels and gradually improves feature representations in the whole learning course. Extensive experiments on five large-scale benchmark datasets show that PCSL significantly outperforms the state-of-the-art unsupervised methods that employ labeled source domains or the images generated by the GANs-based models. Furthermore, the proposed method even has a competitive performance with respect to deep supervised Re-ID methods. Lei Qi 0001, Lei Wang 0001, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | An Effective MR-Guided CT Network Training for Segmenting Prostate in CT ImagesabstractSegmentation of prostate in medical imaging data (e.g., CT, MRI, TRUS) is often considered as a critical yet challenging task for radiotherapy treatment. It is relatively easier to segment prostate from MR images than from CT images, due to better soft tissue contrast of the MR images. For segmenting prostate from CT images, most previous methods mainly used CT alone, and thus their performances are often limited by low tissue contrast in the CT images. In this article, we explore the possibility of using indirect guidance from MR images for improving prostate segmentation in the CT images. In particular, we propose a novel deep transfer learning approach, i.e., MR-guided CT network training (namely MICS-NET), which can employ MR images to help better learning of features in CT images for prostate segmentation. In MICS-NET, the guidance from MRI consists of two steps: (1) learning informative and transferable features from MRI and then transferring them to CT images in a cascade manner, and (2) adaptively transferring the prostate likelihood of MRI model (i.e., well-trained convnet by purely using MR images) with a view consistency constraint. To illustrate the effectiveness of our approach, we evaluate MICS-NET on a real CT prostate image set, with the manual delineations available as the ground truth for evaluation. Our methods generate promising segmentation results which achieve (1) six percentages higher Dice Ratio than the CT model purely using CT images and (2) comparable performance with the MRI model purely using MR images. Wanqi Yang, Yinghuan Shi, Sanghyun Park 0004, Ming Yang 0014, Yang Gao 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Sample-Adaptive GANs: Linking Global and Local Mappings for Cross-Modality MR Image SynthesisabstractGenerative adversarial network (GAN) has been widely explored for cross-modality medical image synthesis. The existing GAN models usually adversarially learn a global sample space mapping from the source-modality to the target-modality and then indiscriminately apply this mapping to all samples in the whole space for prediction. However, due to the scarcity of training samples in contrast to the complicated nature of medical image synthesis, learning a single global sample space mapping that is "optimal" to all samples is very challenging, if not intractable. To address this issue, this paper proposes sample-adaptive GAN models, which not only cater for the global sample space mapping between the source- and the target-modalities but also explore the local space around each given sample to extract its unique characteristic. Specifically, the proposed sample-adaptive GANs decompose the entire learning model into two cooperative paths. The baseline path learns a common GAN model by fitting all the training samples as usual for the global sample space mapping. The new sample-adaptive path additionally models each sample by learning its relationship with its neighboring training samples and using the target-modality features of these training samples as auxiliary information for synthesis. Enhanced by this sample-adaptive path, the proposed sample-adaptive GANs are able to flexibly adjust themselves to different samples, and therefore optimize the synthesis performance. Our models have been verified on three cross-modality MR image synthesis tasks from two public datasets, and they significantly outperform the state-of-the-art methods in comparison. Moreover, the experiment also indicates that our sample-adaptive strategy could be utilized to improve various backbone GAN models. It complements the existing GANs models and can be readily integrated when needed. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Leveraging Coupled Interaction for Multimodal Alzheimer's Disease DiagnosisabstractAs the population becomes older worldwide, accurate computer-aided diagnosis for Alzheimer's disease (AD) in the early stage has been regarded as a crucial step for neurodegeneration care in recent years. Since it extracts the low-level features from the neuroimaging data, previous methods regarded this computer-aided diagnosis as a classification problem that ignored latent featurewise relation. However, it is known that multiple brain regions in the human brain are anatomically and functionally interlinked according to the current neuroscience perspective. Thus, it is reasonable to assume that the extracted features from different brain regions are related to each other to some extent. Also, the complementary information between different neuroimaging modalities could benefit multimodal fusion. To this end, we consider leveraging the coupled interactions in the feature level and modality level for diagnosis in this paper. First, we propose capturing the feature-level coupled interaction using a coupled feature representation. Then, to model the modality-level coupled interaction, we present two novel methods: 1) the coupled boosting (CB) that models the correlation of pairwise coupled-diversity on both inconsistently and incorrectly classified samples between different modalities and 2) the coupled metric ensemble (CME) that learns an informative feature projection from different modalities by integrating the intrarelation and interrelation of training samples. We systematically evaluated our methods with the AD neuroimaging initiative data set. By comparison with the baseline learning-based methods and the state-of-the-art methods that are specially developed for AD/MCI (mild cognitive impairment) diagnosis, our methods achieved the best performance with accuracy of 95.0% and 80.7% (CB), 94.9% and 79.9% (CME) for AD/NC (normal control), and MCI/NC identification, respectively. Yinghuan Shi, Heung-Il Suk, Yang Gao 0001, Seong-Whan Lee, Dinggang Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | A Novel Unsupervised Camera-Aware Domain Adaptation Framework for Person Re-IdentificationabstractUnsupervised cross-domain person re-identification (Re-ID) faces two key issues. One is the data distribution discrepancy between source and target domains, and the other is the lack of discriminative information in target domain. From the perspective of representation learning, this paper proposes a novel end-to-end deep domain adaptation framework to address them. For the first issue, we highlight the presence of camera-level sub-domains as a unique characteristic in person Re-ID, and develop a “camera-aware” domain adaptation method via adversarial learning. With this method, the learned representation reduces distribution discrepancy not only between source and target domains but also across all cameras. For the second issue, we exploit the temporal continuity in each camera of target domain to create discriminative information. This is implemented by dynamically generating online triplets within each batch, in order to maximally take advantage of the steadily improved representation in training process. Together, the above two methods give rise to a new unsupervised domain adaptation framework for person Re-ID. Extensive experiments and ablation studies conducted on benchmark datasets demonstrate its superiority and interesting properties. Lei Qi 0001, Lei Wang 0001, Jing Huo, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
ICCV | 5 |
| 2019 | A Mask Based Deep Ranking Neural Network for Person RetrievalabstractPerson retrieval faces many challenges including cluttered background, appearance variations (e.g., illumination, pose, occlusion) among different camera views and the similarity among different person's images. To address these issues, we put forward a novel mask based deep ranking neural network with a skipped fusing layer. Firstly, to alleviate the problem of cluttered background, masked images with only the foreground regions are incorporated as input in the proposed neural network. Secondly, to reduce the impact of the appearance variations, the multi-layer fusion scheme is developed to obtain more discriminative fine-grained information. Lastly, considering person retrieval is a special image retrieval task, we propose a novel ranking loss to optimize the whole network. The proposed ranking loss can further mitigate the interference problem of similar negative samples when producing ranking results. The extensive experiments validate the superiority of the proposed method compared with the state-of-the-art methods on many benchmark datasets. Lei Qi 0001, Jing Huo, Lei Wang 0001, Yinghuan Shi, Yang Gao 0001 |
ICME | 4 |
| 2019 | A Novel Deep Multi-Modal Feature Fusion Method for Celebrity Video IdentificationabstractIn this paper, we develop a novel multi-modal feature fusion method for the 2019 iQIYI Celebrity Video Identification Challenge, which is held in conjunction with ACM MM 2019. The purpose of this challenge is to retrieve all the video clips of a given identity in the testing set. In this challenge, the multi-modal features of a celebrity are encouraged to be combined for a promising performance, such as face features, head features, body features, and audio features. As we know, the features from different modalities usually have their own influences on the results. To achieve better results, a novel weighted multi-modal feature fusion method is designed to obtain the final feature representation. After many experimental verification, we found that different feature fusion weights for training and testing make the method robust to multi-modal person identification. Experiments on the iQIYI-VID-2019 dataset show that our multi-modal feature fusion strategy effectively improves the accuracy of person identification. Specifically, for competition, we use a single model to get the result of 0.8952 in mAP, which ranks TOP-5 among all the competitive results. Jianrong Chen, Jing Huo, Yinghuan Shi, Yang Gao 0001 |
ACM Multimedia | 5 |
| 2019 | MIDCN: A Multiple Instance Deep Convolutional Network for Image Classification
Kelei He, Jing Huo, Yinghuan Shi, Yang Gao 0001, Dinggang Shen |
PRICAI (1) | 3 |
| 2019 | Crossbar-Net: A Novel Convolutional Neural Network for Kidney Tumor Segmentation in CT ImagesabstractDue to the unpredictable location, fuzzy texture and diverse shape, accurate segmentation of the kidney tumor in CT images is an important yet challenging task. To this end, we in this paper present a cascaded trainable segmentation model termed as Crossbar-Net. Our method combines two novel schemes: (1) we originally proposed the crossbar patches, which consists of two orthogonal non-squared patches (i.e., the vertical patch and horizontal patch). The crossbar patches are able to capture both the global and local appearance information of the kidney tumors from both the vertical and horizontal directions simultaneously. (2) With the obtained crossbar patches, we iteratively train two sub-models (i.e., horizontal sub-model and vertical sub-model) in a cascaded training manner. During the training, the trained sub-models are encouraged to become more focus on the difficult parts of the tumor automatically (i.e., mis-segmented regions). Specifically, the vertical (horizontal) sub-model is required to help segment the mis-segmented regions for the horizontal (vertical) sub-model. Thus, the two sub-models could complement each other to achieve the self-improvement until convergence. In the experiment, we evaluate our method on a real CT kidney tumor dataset which is collected from 94 different patients including 3,500 CT slices. Compared with the state-of-the-art segmentation methods, the results demonstrate the superior performance of our method on the Dice similarity coefficient, true positive fraction, centroid distance and Hausdorff distance. Moreover, to exploit the generalization to other segmentation tasks, we also extend our Crossbar-Net to two related segmentation tasks: (1) cardiac segmentation in MR images and (2) breast mass segmentation in X-ray images, showing the promising results for these two tasks. Our implementation is released at https: //github.com/Qianyu1226/Crossbar-Net. Qian Yu 0007, Yinghuan Shi, Jinquan Sun, Yang Gao 0001, Jianbing Zhu, Yakang Dai |
IEEE Trans. Image Process. | 2 |
| 2019 | Machine Learning in Medical ImagingabstractThe papers in this special issue focus on machine learning for use in medical image processing applications. The use of machine learning in this area has become indispensable in diagnosis and treatment of many diseases. With advances in new imaging techniques, the need to take full advantage of abundant images draws more and more attention. Machine learning, including deep learning particularly, provides us a new paradigm to learn and to utilize the overwhelming volume of big imaging data smartly. Nowadays, machine learning in medical imaging has become one of the most promising and growing fields of research. The main aim of this special issue is to help advance the scientific research within the broad field of machine learning in medical imaging. The special issue was planned in conjunction with the International Workshop on Machine Learning in Medical Imaging (MLMI) 2017. Qian Wang 0001, Yinghuan Shi, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Pelvic Organ Segmentation Using Distinctive Curve Guided Fully Convolutional NetworksabstractAccurate segmentation of pelvic organs (i.e., prostate, bladder, and rectum) from CT image is crucial for effective prostate cancer radiotherapy. However, it is a challenging task due to: 1) low soft tissue contrast in CT images and 2) large shape and appearance variations of pelvic organs. In this paper, we employ a two-stage deep learning-based method, with a novel distinctive curve-guided fully convolutional network (FCN), to solve the aforementioned challenges. Specifically, the first stage is for fast and robust organ detection in the raw CT images. It is designed as a coarse segmentation network to provide region proposals for three pelvic organs. The second stage is for fine segmentation of each organ, based on the region proposal results. To better identify those indistinguishable pelvic organ boundaries, a novel morphological representation, namely, distinctive curve, is also introduced to help better conduct the precise segmentation. To implement this, in this second stage, a multi-task FCN is initially utilized to learn the distinctive curve and the segmentation map separately and then combine these two tasks to produce accurate segmentation map. The final segmentation results of all three pelvic organs are generated by a weighted max-voting strategy. We have conducted exhaustive experiments on a large and diverse pelvic CT data set for evaluating our proposed method. The experimental results demonstrate that our proposed method is accurate and robust for this challenging segmentation task, by also outperforming the state-of-the-art segmentation methods. Kelei He, Xiaohuan Cao, Yinghuan Shi, Dong Nie, Yang Gao 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Ea-GANs: Edge-Aware Generative Adversarial Networks for Cross-Modality MR Image SynthesisabstractMagnetic resonance (MR) imaging is a widely used medical imaging protocol that can be configured to provide different contrasts between the tissues in human body. By setting different scanning parameters, each MR imaging modality reflects the unique visual characteristic of scanned body part, benefiting the subsequent analysis from multiple perspectives. To utilize the complementary information from multiple imaging modalities, cross-modality MR image synthesis has aroused increasing research interest recently. However, most existing methods only focus on minimizing pixel/voxel-wise intensity difference but ignore the textural details of image content structure, which affects the quality of synthesized images. In this paper, we propose edge-aware generative adversarial networks (Ea-GANs) for cross-modality MR image synthesis. Specifically, we integrate edge information, which reflects the textural structure of image content and depicts the boundaries of different objects in images, to reduce this gap. Corresponding to different learning strategies, two frameworks are proposed, i.e., a generator-induced Ea-GAN (gEa-GAN) and a discriminator-induced Ea-GAN (dEa-GAN). The gEa-GAN incorporates the edge information via its generator, while the dEa-GAN further does this from both the generator and the discriminator so that the edge similarity is also adversarially learned. In addition, the proposed Ea-GANs are 3D-based and utilize hierarchical features to capture contextual information. The experimental results demonstrate that the proposed Ea-GANs, especially the dEa-GAN, outperform multiple state-of-the-art methods for cross-modality MR image synthesis in both qualitative and quantitative measures. Moreover, the dEa-GAN also shows excellent generality to generic image synthesis tasks on benchmark datasets about facades, maps, and cityscapes. Biting Yu, Luping Zhou, Lei Wang 0001, Yinghuan Shi, Jurgen Fripp, Pierrick Bourgeat |
IEEE Trans. Medical Imaging | 4 |
| 2018 | A Joint Local and Global Deep Metric Learning Method for Caricature Recognition
Wenbin Li 0006, Jing Huo, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Jiebo Luo 0001 |
ACCV (4) | 3 |
| 2018 | WebCaricature: a benchmark for caricature recognition
Jing Huo, Wenbin Li 0006, Yinghuan Shi, Yang Gao 0001, Hujun Yin |
BMVC | 3 |
| 2018 | Modelling Diffusion Process by Deep Neural Networks for Image Retrieval
Yan Zhao 0019, Lei Wang 0001, Luping Zhou, Yinghuan Shi, Yang Gao 0001 |
BMVC | 4 |
| 2018 | A Novel Image-Specific Transfer Approach for Prostate Segmentation in MR ImagesabstractProstate segmentation in Magnetic Resonance (MR) Images is a significant yet challenging task for prostate cancer treatment. Most of the existing works attempted to design a global classifier for all MR images, which neglect the discrepancy of images across different patients. To this end, we propose a novel transfer approach for prostate segmentation in MR images. Firstly, an image-specific classifier is built for each training image. Secondly, a pair of dictionaries and a mapping matrix are jointly obtained by a novel Semi-Coupled Dictionary Transfer Learning (SCDTL). Finally, the classifiers on the source domain could be selectively transferred to the target domain (i.e. testing images) by the dictionaries and the mapping matrix. The evaluation demonstrates that our approach has a competitive performance compared with the state-of-the-art transfer learning methods. Moreover, the proposed transfer approach outperforms the conventional deep neural network based method. Pinzhuo Tian, Lei Qi 0001, Yinghuan Shi, Luping Zhou, Yang Gao 0001, Dinggang Sheri |
ICASSP | 3 |
| 2018 | A Novel Two-Stage Deep Method for Mitosis Detection in Breast Cancer Histology ImagesabstractThe accurate detection and counting of mitosis in breast cancer histology images is very important for computer-aided diagnosis, which is manually completed by the pathologist according to her or his clinic experience. However, this procedure is extremely time consuming and tedious. Moreover, it always results in low agreement among different pathologists. Although several computer-aided detection methods have been developed recently, they suffer from high FN (false negative) and FP (false positive) with simply treating the detection task as a binary classification problem. In this paper, we present a novel two-stage detection method with multi-scale and similarity learning convnets (MSSN). Firstly, large amount of possible candidates will be generated in the first stage in order to reduce FN (i.e., prevent treating mitosis as non-mitosis), by using the different square and non-square filters, to capture the spatial relation from different scales. Secondly, a similarity prediction model is subsequently performed on the obtained candidates for the final detection to reduce FP, which is realized by imposing a large margin constraint. On both 2014 and 2012 ICPR MITOSIS datasets, our MSSN achieved a promising result with a highest Recall (outperforming other methods by a large margin) and a comparable F-score. Minglin Ma, Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001 |
ICPR | 2 |
| 2018 | Online multi-view subspace learning via group structure analysis for visual object tracking
Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Ming Yang 0014 |
Distributed Parallel Databases | 2 |
| 2018 | OPML: A one-pass closed-form solution for online metric learning
Wenbin Li 0006, Yang Gao 0001, Lei Wang 0001, Luping Zhou, Jing Huo, Yinghuan Shi |
Pattern Recognit. | 6 |
| 2018 | Heterogeneous Face Recognition by Margin-Based Cross-Modality Metric LearningabstractHeterogeneous face recognition deals with matching face images from different modalities or sources. The main challenge lies in cross-modal differences and variations and the goal is to make cross-modality separation among subjects. A margin-based cross-modality metric learning (MCM2L) method is proposed to address the problem. A cross-modality metric is defined in a common subspace where samples of two different modalities are mapped and measured. The objective is to learn such metrics that satisfy the following two constraints. The first minimizes pairwise, intrapersonal cross-modality distances. The second forces a margin between subject specific intrapersonal and interpersonal cross-modality distances. This is achieved by defining a hinge loss on triplet-based distance constraints for efficient optimization. It allows the proposed method to focus more on optimizing distances of those subjects whose intrapersonal and interpersonal distances are hard to separate. The proposed method is further extended to a kernelized MCM2L (KMCM2L). Both methods have been evaluated on an ID card face dataset and two other cross-modality benchmark datasets. Various feature extraction methods have also been incorporated in the study, including recent deep learned features. In extensive experiments and comparisons with the state-of-the-art methods, the MCM2L and KMCM2L methods achieved marked improvements in most cases. Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin |
IEEE Trans. Cybern. | 3 |
| 2018 | Cross-Modal Metric Learning for AUC OptimizationabstractCross-modal metric learning (CML) deals with learning distance functions for cross-modal data matching. The existing methods mostly focus on minimizing a loss defined on sample pairs. However, the numbers of intraclass and interclass sample pairs can be highly imbalanced in many applications, and this can lead to deteriorating or unsatisfactory performances. The area under the receiver operating characteristic curve (AUC) is a more meaningful performance measure for the imbalanced distribution problem. To tackle the problem as well as to make samples from different modalities directly comparable, a CML method is presented by directly maximizing AUC. The method can be further extended to focus on optimizing partial AUC (pAUC), which is the AUC between two specific false positive rates (FPRs). This is particularly useful in certain applications where only the performances assessed within predefined false positive ranges are critical. The proposed method is formulated as a log-determinant regularized semidefinite optimization problem. For efficient optimization, a minibatch proximal point algorithm is developed. The algorithm is experimentally verified stable with the size of sampled pairs that form a minibatch at each iteration. Several data sets have been used in evaluation, including three cross-modal data sets on face recognition under various scenarios and a single modal data set, the Labeled Faces in the Wild. Results demonstrate the effectiveness of the proposed methods and marked improvements over the existing methods. Specifically, pAUC-optimized CML proves to be more competitive for performance measures such as Rank-1 and verification rate at FPR = 0.1%. Jing Huo, Yang Gao 0001, Yinghuan Shi, Hujun Yin |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Incomplete-Data Oriented Multiview Dimension Reduction via Sparse Low-Rank RepresentationabstractFor dimension reduction on multiview data, most of the previous studies implicitly take an assumption that all samples are completed in all views. Nevertheless, this assumption could often be violated in real applications due to the presence of noise, limited access to data, equipment malfunction, and so on. Most of the previous methods will cease to work when missing values in one or multiple views occur, thus an incomplete-data oriented dimension reduction becomes an important issue. To this end, we mathematically formulate the above-mentioned issue as sparse low-rank representation through multiview subspace (SRRS) learning to impute missing values, by jointly measuring intraview relations (via sparse low-rank representation) and interview relations (through common subspace representation). Moreover, by exploiting various subspace priors in the proposed SRRS formulation, we develop three novel dimension reduction methods for incomplete multiview data: 1) multiview subspace learning via graph embedding; 2) multiview subspace learning via structured sparsity; and 3) sparse multiview feature selection via rank minimization. For each of them, the objective function and the algorithm to solve the resulting optimization problem are elaborated, respectively. We perform extensive experiments to investigate their performance on three types of tasks including data recovery, clustering, and classification. Both two toy examples (i.e., Swiss roll and -curve) and four real-world data sets (i.e., face images, multisource news, multicamera activity, and multimodality neuroimaging data) are systematically tested. As demonstrated, our methods achieve the performance superior to that of the state-of-the-art comparable methods. Also, the results clearly show the advantage of integrating the sparsity and low-rankness over using each of them separately. Wanqi Yang, Yinghuan Shi, Yang Gao 0001, Lei Wang 0001, Ming Yang 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Beyond IID: Learning to Combine Non-IID Metrics for Vision TasksabstractMetric learning has been widely employed, especially in various computer vision tasks, with the fundamental assumption that all samples (e.g., regions/superpixels in images/videos) are independent and identically distributed (IID). However, since the samples are usually spatially-connected or temporally-correlated with their physically-connected neighbours, they are not IID (non-IID for short), which cannot be directly handled by existing methods. Thus, we propose to learn and integrate non-IID metrics (NIME). To incorporate the non-IID spatial/temporal relations, instead of directly using non-IID features and metric learning as previous methods, NIME first builds several non-IID representations on original (non-IID) features by various graph kernel functions, and then automatically learns the metric under the best combination of various non-IID representations. NIME is applied to solve two typical computer vision tasks: interactive image segmentation and histology image identification. The results show that learning and integrating non-IID metrics improves the performance, compared to the IID methods. Moreover, our method achieves results comparable or better than that of the state-of-the-arts. Yinghuan Shi, Wenbin Li 0006, Yang Gao 0001, Longbing Cao, Dinggang Shen |
AAAI | 1 |
| 2017 | Revisiting Metric Learning for SPD Matrix Based Visual RepresentationabstractThe success of many visual recognition tasks largely depends on a good similarity measure, and distance metric learning plays an important role in this regard. Meanwhile, Symmetric Positive Definite (SPD) matrix is receiving increased attention for feature representation in multiple computer vision applications. However, distance metric learning on SPD matrices has not been sufficiently researched. A few existing works approached this by learning either d2× p or d × k transformation matrix for d× d SPD matrices. Different from these methods, this paper proposes a new member to the family of distance metric learning for SPD matrices. It learns only d parameters to adjust the eigenvalues of the SPD matrices through an efficient optimisation scheme. Also, it is shown that the proposed method can be interpreted as learning a sample-specific transformation matrix, instead of the fixed transformation matrix learned for all the samples in the existing works. The optimised d parameters can be used to massage the SPD matrices for better discrimination while still keeping them in the original space. From this perspective, the proposed method complements, rather than competes with, the existing linear-transformation-based methods, as the latter can always be applied to the output of the former to perform distance metric learning in further. The proposed method has been tested on multiple SPD-based visual representation data sets used in the literature, and the results demonstrate its interesting properties and attractive performance. Luping Zhou, Lei Wang 0001, Jianjia Zhang, Yinghuan Shi, Yang Gao 0001 |
CVPR | 4 |
| 2017 | Cost-Sensitive Alternating Direction Method of Multipliers for Large-Scale Classification
Yinghuan Shi, Xingguo Chen, Yang Gao 0001 |
IDEAL | 2 |
| 2017 | Does Manual Delineation only Provide the Side Information in CT Prostate Segmentation?
Yinghuan Shi, Wanqi Yang, Yang Gao 0001, Dinggang Shen |
MICCAI (3) | 1 |
| 2017 | Group-Based Alternating Direction Method of Multipliers for Distributed Linear ClassificationabstractThe alternating direction method of multipliers (ADMM) algorithm has been widely employed for distributed machine learning tasks. However, it suffers from several limitations, e.g., a relative low convergence speed, and an expensive time cost. To this end, in this paper, a novel method, namely the group-based ADMM (GADMM), is proposed for distributed linear classification. In particular, to accelerate the convergence speed and improve global consensus, a group layer is first utilized in GADMM to divide all the slave nodes into several groups. Then, all the local variables (from the slave nodes) are gathered in the group layer to generate different group variables. Finally, by using a weighted average method, the group variables are coordinated to update the global variable (from the master node) until the solution of the global problem is reached. According to the theoretical analysis, we found that: 1) GADMM can mathematically converge at the rate , where is the number of outer iterations and 2) by using the grouping methods, GADMM can improve the convergence speed compared with the distributed ADMM framework without grouping methods. Moreover, we systematically evaluate GADMM on four publicly available LIBSVM datasets. Compared with disADMM and stochastic dual coordinate ascent with alternating direction method of multipliers-ADMM, for distributed classification, GADMM is able to reduce the number of outer iterations, which leads to faster convergence speed and better global consensus. In particular, the statistical significance test has been experimentally conducted and the results validate that GADMM can significantly save up to 30% of the total time cost (with less than 0.6% accuracy loss) compared with disADMM on large-scale datasets, e.g., webspam and epsilon. Yang Gao 0001, Yinghuan Shi, Ruili Wang 0001 |
IEEE Trans. Cybern. | 3 |
| 2016 | A Fast Distributed Classification Algorithm for Large-Scale Imbalanced DataabstractThe Alternating Direction Method of Multipliers (ADMM) has been developed recently for distributed classification. Nevertheless, the widely-existing class imbalance problem has not been well investigated. Furthermore, previous imbalanced classification methods lack of efforts in studying the complex imbalance problem in a distributed environment. In this paper, we consider the imbalance problem as distributed data imbalance which includes three imbalance issues: (i) within-node class imbalance, (ii)between-node class imbalance, and (iii) between-node structure imbalance. In order to adequately deal with imbalanced data as well as improve time efficiency, a novel distributed Cost-Sensitive classification algorithm via Group-based ADMM (CS-GADMM) is proposed. Briefly, CS-GADMM derives the classification problem as a series of sub-problems with within-node class imbalance. To alleviate the time delay caused by between-node class imbalance, we propose a extension of dual coordinate descent method for the sub-problem optimization. Meanwhile, for between-node structure imbalance, we discreetly study the relationship between local functions, and combine the resulting local variables intra-group to update the global variables for prediction. The experimental results on various imbalanced datasets validate that CS-GADMM could be a efficient algorithm for imbalanced classification. Yang Gao 0001, Yinghuan Shi, Hao Wang 0013 |
ICDM | 3 |
| 2016 | Grouping Parallel Bayesian Network Structure Learning Algorithm Based on Variable Ordering
Xiaolong Qi, Yinghuan Shi, Hao Wang 0013, Yang Gao 0001 |
IDEAL | 2 |
| 2016 | Multi-view Subspace Clustering via a Global Low-Rank Affinity Matrix
Lei Qi 0001, Yinghuan Shi, Wanqi Yang, Yang Gao 0001 |
IDEAL | 2 |
| 2016 | Ensemble of Sparse Cross-Modal Metrics for Heterogeneous Face RecognitionabstractHeterogeneous face recognition aims to identify or verify person identity by matching facial images of different modalities. In practice, it is known that its performance is highly influenced by modality inconsistency, appearance occlusions, illumination variations and expressions. In this paper, a new method named as ensemble of sparse cross-modal metrics is proposed for tackling these challenging issues. In particular, a weak sparse cross-modal metric learning method is firstly developed to measure distances between samples of two modalities. It learns to adjust rank-one cross-modal metrics to satisfy two sets of triplet based cross-modal distance constraints in a compact form. Meanwhile, a group based feature selection is performed to enforce that features in the same position of two modalities are selected simultaneously. By neglecting features that attribute to "noise" in the face regions (eye glasses, expressions and so on), the performance of learned weak metrics can be markedly improved. Finally, an ensemble framework is incorporated to combine the results of differently learned sparse metrics into a strong one. Extensive experiments on various face datasets demonstrate the benefit of such feature selection especially when heavy occlusions exist. The proposed ensemble metric learning has been shown superiority over several state-of-the-art methods in heterogeneous face recognition. Jing Huo, Yang Gao 0001, Yinghuan Shi, Wanqi Yang, Hujun Yin |
ACM Multimedia | 3 |
| 2016 | A learning-based CT prostate segmentation method via joint transductive feature selection and regression
Yinghuan Shi, Yaozong Gao, Shu Liao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen |
Neurocomputing | 1 |
| 2015 | Interactive image segmentation via cascaded metric learningabstractIn this paper, we propose an interactive image segmentation method from a novel perspective of cascaded metric learning. Given an image with user-marked scribbles that are essentially uncertain and noisy, our method completes the segmentation task by solving a binary classification problem. Starting from the initial training samples with known class labels (i.e., regions of the image that are believed with high confidence to be foreground or background), we first find an optimal metric that can best describe the classification of these samples. After that, we classify the unlabeled samples using the learnt metric. Samples classified with high confidence are used as new training samples to refine the metric. This cycle of metric learning and classification repeats until the accomplishment of the image segmentation task. The proposed method is extensively evaluated on the MSRC image set. Experiment results show that our method outperforms the state-of-the-art methods. Wenbin Li 0006, Yinghuan Shi, Wanqi Yang, Hao Wang 0013, Yang Gao 0001 |
ICIP | 2 |
| 2015 | Semi-Automatic Segmentation of Prostate in CT Images via Coupled Feature Representation and Spatial-Constrained Transductive LassoabstractConventional learning-based methods for segmenting prostate in CT images ignore the relations among the low-level features by assuming all these features are independent. Also, their feature selection steps usually neglect the image appearance changes in different local regions of CT images. To this end, we present a novel semi-automatic learning-based prostate segmentation method in this article. For segmenting the prostate in a certain treatment image, the radiation oncologist will be first asked to take a few seconds to manually specify the first and last slices of the prostate. Then, prostate is segmented with the following two steps: (i) Estimation of 3D prostate-likelihood map to predict the likelihood of each voxel being prostate by employing the coupled feature representation, and the proposed Spatial-COnstrained Transductive LassO (SCOTO); (ii) Multi-atlases based label fusion to generate the final segmentation result by using the prostate shape information obtained from both planning and previous treatment images. The major contribution of the proposed method mainly includes: (i) incorporating radiation oncologist's manual specification to aid segmentation, (ii) adopting coupled features to relax previous assumption of feature independency for voxel representation, and (iii) developing SCOTO for joint feature selection across different local regions. The experimental result shows that the proposed method outperforms the state-of-the-art methods in a real-world prostate CT dataset, consisting of 24 patients with totally 330 images, all of which were manually delineated by the radiation oncologist for performance evaluation. Moreover, our method is also clinically feasible, since the segmentation performance can be improved by just requiring the radiation oncologist to spend only a few seconds for manual specification of ending slices in the current treatment CT image. Yinghuan Shi, Yaozong Gao, Shu Liao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | MRM-Lasso: A Sparse Multiview Feature Selection Method via Low-Rank AnalysisabstractLearning about multiview data involves many applications, such as video understanding, image classification, and social media. However, when the data dimension increases dramatically, it is important but very challenging to remove redundant features in multiview feature selection. In this paper, we propose a novel feature selection algorithm, multiview rank minimization-based Lasso (MRM-Lasso), which jointly utilizes Lasso for sparse feature selection and rank minimization for learning relevant patterns across views. Instead of simply integrating multiple Lasso from view level, we focus on the performance of sample-level (sample significance) and introduce pattern-specific weights into MRM-Lasso. The weights are utilized to measure the contribution of each sample to the labels in the current view. In addition, the latent correlation across different views is successfully captured by learning a low-rank matrix consisting of pattern-specific weights. The alternating direction method of multipliers is applied to optimize the proposed MRM-Lasso. Experiments on four real-life data sets show that features selected by MRM-Lasso have better multiview classification performance than the baselines. Moreover, pattern-specific weights are demonstrated to be significant for learning about multiview data, compared with view-specific weights. Wanqi Yang, Yang Gao 0001, Yinghuan Shi, Longbing Cao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | Joint Coupled-Feature Representation and Coupled Boosting for AD DiagnosisabstractRecently, there has been a great interest in computer-aided Alzheimer's Disease (AD) and Mild Cognitive Impairment (MCI) diagnosis. Previous learning based methods defined the diagnosis process as a classification task and directly used the low-level features extracted from neuroimaging data without considering relations among them. However, from a neuroscience point of view, it's well known that a human brain is a complex system that multiple brain regions are anatomically connected and functionally interact with each other. Therefore, it is natural to hypothesize that the low-level features extracted from neuroimaging data are related to each other in some ways. To this end, in this paper, we first devise a coupled feature representation by utilizing intra-coupled and inter-coupled interaction relationship. Regarding multi-modal data fusion, we propose a novel coupled boosting algorithm that analyzes the pairwise coupled-diversity correlation between modalities. Specifically, we formulate a new weight updating function, which considers both incorrectly and inconsistently classified samples. In our experiments on the ADNI dataset, the proposed method presented the best performance with accuracies of 94.7% and 80.1% for AD vs. Normal Control (NC) and MCI vs. NC classifications, respectively, outperforming the competing methods and the state-of-the-art methods. Yinghuan Shi, Heung-Il Suk, Yang Gao 0001, Dinggang Shen |
CVPR | 1 |
| 2014 | A Novel Ego-Centered Academic Community Detection Approach via Factor Graph Model
Yusheng Jia, Yang Gao 0001, Wanqi Yang, Jing Huo, Yinghuan Shi |
IDEAL | 5 |
| 2014 | mPadal: a joint local-and-global multi-view feature selection method for activity recognition
Wanqi Yang, Yang Gao 0001, Longbing Cao, Ming Yang 0014, Yinghuan Shi |
Appl. Intell. | 5 |
| 2013 | Prostate Segmentation in CT Images via Spatial-Constrained Transductive LassoabstractAccurate prostate segmentation in CT images is a significant yet challenging task for image guided radiotherapy. In this paper, a novel semi-automated prostate segmentation method is presented. Specifically, to segment the prostate in the current treatment image, the physician first takes a few seconds to manually specify the first and last slices of the prostate in the image space. Then, the prostate is segmented automatically by the proposed two steps: (i) The first step of prostate-likelihood estimation to predict the prostate likelihood for each voxel in the current treatment image, aiming to generate the 3-D prostate-likelihood map by the proposed Spatial-COnstrained Transductive LassO (SCOTO), (ii) The second step of multi-atlases based label fusion to generate the final segmentation result by using the prostate shape information obtained from the planning and previous treatment images. The experimental result shows that the proposed method outperforms several state-of-the-art methods on prostate segmentation in a real prostate CT dataset, consisting of 24 patients with 330 images. Moreover, it is also clinically feasible since our method just requires the physician to spend a few seconds on manual specification of the first and last slices of the prostate. Yinghuan Shi, Shu Liao, Yaozong Gao, Daoqiang Zhang, Yang Gao 0001, Dinggang Shen |
CVPR | 1 |
| 2013 | Image Super Resolution via Visual Prior Based Digital Image Characteristics
Yusheng Jia, Wanqi Yang, Yang Gao 0001, Hujun Yin, Yinghuan Shi |
IDEAL | 5 |
| 2013 | Voting-XCSc: A Consensus Clustering Method via Learning Classifier System
Liqiang Qian, Yinghuan Shi, Yang Gao 0001, Hujun Yin |
IDEAL | 2 |
| 2013 | Transductive cost-sensitive lung cancer image classification
Yinghuan Shi, Yang Gao 0001, Ruili Wang 0001, Dong Wang 0038 |
Appl. Intell. | 1 |
| 2011 | P2LSA and P2LSA+: Two Paralleled Probabilistic Latent Semantic Analysis Algorithms Based on the MapReduce Model
Yang Gao 0001, Yinghuan Shi, Lin Shang 0001, Ruili Wang 0001 |
IDEAL | 3 |
| 2011 | Xcsc: a Novel Approach to Clustering with Extended Classifier SystemabstractIn this paper, we propose a novel approach to clustering noisy and complex data sets based on the eXtend Classifier Systems (XCS). The proposed approach, termed XCSc, has three main processes: (a) a learning process to evolve the rule population, (b) a rule compacting process to remove redundant rules after the learning process, and (c) a rule merging process to deal with the overlapping rules that commonly occur between the clusters. In the first process, we have modified the clustering mechanisms of the current available XCS and developed a new accelerate learning method to improve the quality of the evolved rule population. In the second process, an effective rule compacting algorithm is utilized. The rule merging process is based on our newly proposed agglomerative hierarchical rule merging algorithm, which comprises the following steps: (i) all the generated rules are modeled by a graph, with each rule representing a node; (ii) the vertices in the graph are merged to form a number of sub-graphs (i.e. rule clusters) under some pre-defined criteria, which generates the final rule set to represent the clusters; (iii) each data is re-checked and assigned to a cluster that it belongs to, guided by the final rule set. In our experiments, we compared the proposed XCSc with CHAMELEON, a benchmark algorithm well known for its excellent performance, on a number of challenging data sets. The results show that the proposed approach outperforms CHAMELEON in the successful rate, and also demonstrates good stability. Liangdong Shi, Yinghuan Shi, Yang Gao 0001, Lin Shang 0001 |
Int. J. Neural Syst. | 2 |
| 2010 | Real-Time Abnormal Event Detection in Complicated ScenesabstractIn this paper, we proposed a novel real-time abnormal event detection framework that requires a short training period and has a fast processing speed. Our approach is based on phase correlation and our newly developed spatial-temporal co-occurrence Gaussian mixture models (STCOG)with the following steps: (i) a frame is divided into non-overlapping local regions; (ii) phase correlation is used to estimate the motion vectors between successive two frames for all corresponding local regions, and (iii) STCOG is used to model normal events and detect abnormal events if any deviation from the trained STCOG is found. Our proposed approach is also able to update the parameters incrementally and can be applied in complicated scenes. The proposed approach outperforms previous ones in terms of shorter training periods and lower computational complexity. Yinghuan Shi, Yang Gao 0001, Ruili Wang 0001 |
ICPR | 1 |
| 2009 | Apply ant colony optimization to TetrisabstractTetris is a falling block game where the player's objective is to arrange a sequence of different shaped tetrominoes smoothly in order to survive. In the intelligence games, agent imitates the real player and chooses the best move based on a linear value function. In this paper, we apply Ant Colony Optimization (ACO) method to learn the weights of the function, trying to search an optimal weight-path in the weight graph. We use dynamic heuristic to prevent premature convergence to local optima. Our experimental result is better than most of traditional reinforcement learning methods. Xingguo Chen, Hao Wang 0013, Weiwei Wang 0002, Yinghuan Shi, Yang Gao 0001 |
GECCO | 4 |
| 2009 | Clustering with XCS and Agglomerative Rule Merging
Liangdong Shi, Yinghuan Shi, Yang Gao 0001 |
IDEAL | 2 |