VLDB 2026 Research / reviewers in the wild / expert
Xiaoshuang Shi
dblp:87/10627
· DBLP profile ↗
76ranked-venue papers
19as first author
49since 2021 · last 2026
0000-0003-4934-0850ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 12 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 8 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk GuaranteesabstractUncertainty quantification (UQ) in foundation models is crucial for identifying and mitigating hallucinations in automatically generated text. However, heuristic UQ approaches lack statistical guarantees for key metrics such as the false discovery rate (FDR) in selective prediction tasks. Previous research adopts the split conformal prediction (SCP) framework to ensure desired coverage of admissible answers by constructing data-driven prediction sets, yet these sets typically contain incorrect candidates, undermining their practical effectiveness. To address this, we introduce COIN, an uncertainty-guarding selection framework that calibrates statistically valid uncertainty thresholds to filter a single generated answer per question under user-specified FDR constraints. COIN estimates the empirical error rate on the calibration set and applies confidence interval methods such as Clopper–Pearson to establish a high-probability upper bound on the true error rate (i.e., FDR). This enables the selection of the largest threshold that ensures FDR control on test data while significantly increasing sample retention. We demonstrate COIN's robustness in risk control, strong test-time power in retaining admissible answers, and predictive efficiency under limited calibration data across both general and multimodal text generation tasks. Furthermore, we show that employing alternative UQ and upper bound construction strategies can further boost COIN's power performance, which underscores its extensibility and adaptability to diverse application scenarios. Zhiyuan Wang 0007, Jinhao Duan, Qingni Wang, Xiaofeng Zhu 0001, Tianlong Chen 0001, Xiaoshuang Shi, Kaidi Xu |
AAAI | 6 |
| 2026 | DFPA: Dual-level Feature Perturbation Augmentation for Whole Slide Image Analysis
Chaojun Zhang, Zhenhua Guo 0001, Xiaoshuang Shi |
ICIC (7) | 4 |
| 2026 | Co-assistant networks by pathology foundation model and convolutional neural network for gigapixel whole slide image analysis
Meilian Xu, Xiaofeng Zhu 0001, Xiaoshuang Shi |
Medical Image Anal. | 6 |
| 2026 | Attention-guided knowledge distillation based deep multiple instance learning for gigapixel whole slide image analysis
Weiheng Fu, Yike Zhou, Meilian Xu, Tianfu Wen, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
Pattern Recognit. | 5 |
| 2026 | Feature-interpretable disease prediction from tabular data via dynamic GCN and LLMs
Shuting Pang, Junren Wang, He Lyu, Huan Song, Xiaoshuang Shi |
Pattern Recognit. | 9 |
| 2025 | SConU: Selective Conformal Uncertainty in Large Language ModelsabstractZhiyuan Wang, Qingni Wang, Yue Zhang, Tianlong Chen, Xiaofeng Zhu, Xiaoshuang Shi, Kaidi Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhiyuan Wang 0007, Qingni Wang, Yue Zhang 0025, Tianlong Chen 0001, Xiaofeng Zhu 0001, Xiaoshuang Shi, Kaidi Xu |
ACL (1) | 6 |
| 2025 | Interpretable Multi-View Fusion Network for Alzheimer's Disease Diagnosis with Large-Scale Pre-Trained Vision-Language ModelabstractRecent advances demonstrate that large-scale Vision-Language Models (VLMs) achieve impressive benchmarks on natural images via prompt learning. However, their applications to medical images are hampered by failing to consider domain-specific characteristics inherent to pathological regions, such as structural magnetic resonance imaging (sMRI), a key modality for Alzheimer's disease (AD) diagnosis. Additionally, existing single-view methods usually fail to capture sMRI's spatial hierarchy effectively while neglecting local pathological variations, and existing multi-view methods often disregard the strong inter-view consistency inherent in sMRI. To address these limitations, we propose a novel interpretable multi-view fusion framework, namely LaCA, by integrating loss-attentionbased prompt (LaP) learning and consistency-augmented multiview learning (CAML) with a large-scale pre-trained VLM backbone. Specifically, LaCA first applies a 3D CNN to the raw sMRI data to lower multi-view training costs. Then, the 3D CNN features are decomposed into multiple views, each of which is fed into the proposed LaP module to enhance lesion recognition capability while mitigating the domain gap between pre-training and downstream datasets. Next, the VLM leverages its robust feature extraction capabilities to derive deep representations of each view. After that, the designed CAML module is used to enhance inter-view consistency, and perform multi-view fusion by integrating complementary multiview information to improve AD diagnostic performance. Finally, experimental results demonstrate the superior classification and interpretation performance of the proposed framework over recent state-of-the-art AD diagnostic methods. The source codes are available at https://github.com/MakimaSasha/LaCA. Jinghao Xu, Xiaoshuang Shi |
BIBM | 6 |
| 2025 | TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination via Latent Truthful-Guided Pre-InterventionabstractObject Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such as hidden states, encode the "overall truthfulness" of generated responses. However, it remains under-explored how internal states in LVLMs function and whether they could serve as "per-token" hallucination indicators, which is essential for mitigating OH. In this paper, we first conduct an in-depth exploration of LVLM internal states with OH issues and discover that (1) LVLM internal states are high-specificity per-token indicators of hallucination behaviors. Moreover, (2) different LVLMs encode universal patterns of hallucinations in common latent subspaces, indicating that there exist "generic truthful directions" shared by various LVLMs. Based on these discoveries, we propose Truthful-Guided Pre-Intervention (TruthPrInt) that first learns the truthful direction of LVLM decoding and then applies truthful-guided inference-time intervention during LVLM decoding. We further propose TruthPrInt to enhance both cross-LVLM and cross-data hallucination detection transferability by constructing and aligning hallucination latent subspaces. We evaluate TruthPrInt in extensive experimental settings, including in-domain and out-of-domain scenarios, over popular LVLMs and OH benchmarks. Experimental results indicate that TruthPrInt significantly outperforms state-of-the-art methods. Codes will be available at https://github.com/jinhaoduan/TruthPrInt. Jinhao Duan, Fei Kong, Hao Cheng 0015, James Diffenderfer, Bhavya Kailkhura, Lichao Sun 0001, Xiaofeng Zhu 0001, Xiaoshuang Shi, Kaidi Xu |
ICCV | 8 |
| 2025 | Enhancing the Influence of Labels on Unlabeled Nodes in Graph Convolutional NetworksabstractThe message-passing mechanism of graph convolutional networks (i.e., GCNs) enables label information to reach more unlabeled neighbors, thereby increasing the utilization of labels. However, the additional label information does not always contribute positively to the GCN. To address this issue, we propose a new two-step framework called ELU-GCN. In the first stage, ELU-GCN conducts graph learning to learn a new graph structure (i.e., ELU-graph), which allows the additional label information to positively influence the predictions of GCN. In the second stage, we design a new graph contrastive learning on the GCN framework for representation learning by exploring the consistency and mutually exclusive information between the learned ELU graph and the original graph. Moreover, we theoretically demonstrate that the proposed method can ensure the generalization ability of GCNs. Extensive experiments validate the superiority of our method. Jincheng Huang 0005, Yujie Mo, Xiaoshuang Shi, Lei Feng 0006, Xiaofeng Zhu 0001 |
ICML | 3 |
| 2025 | Meta Label Correction with Generalization RegularizerabstractDeep neural networks can easily lead to the over-fitting issue due to the influence of noisy labels. However, previous label correction methods for dealing with noisy labels often need expensive computation cost to achieve effectiveness and ignore the generalization ability of the model. To address these issues, in this paper, we propose a new meta-based self-correction method to achieve accurate filtering of noisy labels and to enhance the generalization ability of the label correction model. Specifically, we first investigate a new gradient score method to filter noisy labels with less computation cost, and then theoretically design a new generalization regularizer into the meta-learner and the base learner, for correcting noisy labels as well as achieving the generalization ability. Experimental results on real datasets verify the effectiveness of our proposed method in terms of different classification tasks. Tao Tong, Yujie Mo, Yucheng Xie, Songyue Cai, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
IJCAI | 5 |
| 2025 | Word-Sequence Entropy: Towards uncertainty estimation in free-form medical question answering applications and beyond
Zhiyuan Wang 0007, Jinhao Duan, Chenxi Yuan, Qingyu Chen 0001, Tianlong Chen 0001, Yue Zhang 0025, Ren Wang 0008, Xiaoshuang Shi, Kaidi Xu |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Dynamic graph based weakly supervised deep hashing for whole slide image classification and retrieval
Haochen Jin, Xiaoshuang Shi, Kang Li 0004, Xiaofeng Zhu 0001 |
Medical Image Anal. | 4 |
| 2025 | Interpretable 2.5D network by hierarchical attention and consistency learning for 3D MRI classification
Shuting Pang, Xiaoshuang Shi, Rui Wang 0108, Mingzhe Dai, Xiaofeng Zhu 0001, Bin Song 0002, Kang Li 0004 |
Pattern Recognit. | 3 |
| 2025 | Exploring Unbiased Activation Maps for Weakly Supervised Tissue Segmentation of Histopathological ImagesabstractTissue segmentation in histopathological images plays a crucial role in computational pathology, owing to its significant potential to indicate the prognosis of cancer patients. Presently, numerous Weakly Supervised Semantic Segmentation (WSSS) methods strive to utilize image-level labels to achieve pixel-level segmentation, aiming to minimize the need for detailed annotations. Most of these methods rely on Class Activation Maps (CAM) extracted from classification models, frequently leading to poor coverage of objects. The major cause is attributed to the strong inductive bias of the classification model, focusing primarily on discriminative feature of objects, rather than non-discriminative features. Inspired by this, we propose a simple yet effective method that introduces a self-supervised task by exploiting both the discriminative and non-discriminative features, and generate Unbiased Activation Maps (UAM) to encompass the whole object. Specifically, our method entails clustering all spatial features of an object class to derive semantic centers. Each center then works as a spatial filter that amplifies similar feature and suppresses dissimilar feature, and extract high-quality pseudo-labels (some noise at object boundaries). Moreover, we further propose a Noise-Reduced (NR) Learning method to train the segmentation network towards credible signals and lessen the impact of false predictions. Comprehensive experimental results on two public histopathology image datasets demonstrate the superior performance of our method over the state-of-the-art weakly supervised segmentation methods. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Xiao Zhang 0028, Yaqiong Xing, Yuting Wen, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Feature Noise Boosts DNN Generalization Under Label NoiseabstractThe presence of label noise in the training data has a profound impact on the generalization of deep neural networks (DNNs). In this study, we introduce and theoretically demonstrate a simple feature noise (FN) method, which directly adds noise to the features of training data and can enhance the generalization of DNNs under label noise. Specifically, we conduct theoretical analyses to reveal that label noise leads to weakened DNN generalization by loosening the generalization bound, and FN results in better DNN generalization by imposing an upper bound on the mutual information between the model weights and the features, which constrains the generalization bound. Furthermore, we conduct a qualitative analysis to discuss the ideal type of FN that obtains good label noise generalization. Finally, extensive experimental results on several popular datasets demonstrate that the FN method can significantly enhance the label noise generalization of state-of-the-art methods. The source codes of the FN method are available on https://github.com/zlzenglu/FN. Xiaoshuang Shi, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Patch Target Guided Dual-Branch Deep Multiple Instance Learning for 3D MRI AnalysisabstractDeep multiple instance learning (MIL) has attracted considerable attention in medical image analysis, since it only requires image-level labels for model training without using fine-grained (or patch) annotations. Unfortunately, MIL-based methods might lose some significant patch features. Although pseudo-label-based methods, which assign a pre-defined label to each patch, can explore more patch-level features, they might bring label noise and make the patch-level features lose diversity, thereby possibly restricting the model performance. To overcome this issue, we propose a novel gradient-based patch target generation (PTG) module to dynamically produce a feature vector for each patch as its target. Additionally, based on the PTG module, we propose a patch-target guided dual-branch deep MIL framework for 3D MRI data analysis, where both the two branches consist of a CNN model to extract patch-level features, an attention module to interpret the significance of patches, and a bag-level classifier, while the second branch also contains the PTG module to generate patch targets of patches. Moreover, the two branches are alternatively updated in our framework, resulting in a bi-level optimization problem, and thus we design a bi-level optimization algorithm to solve our proposed objective function. Extensive experiments demonstrate the superior classification and interpretation performance of the proposed framework over recent state-of-the-art methods. Codes are available at https://github.com/daimz1213/mil2024. Mingzhe Dai, Xiaoshuang Shi, Xiaofeng Zhu 0001, Tingrui Pan, Kang Li 0004 |
BIBM | 2 |
| 2024 | ACT-Diffusion: Efficient Adversarial Consistency Training for One-Step Diffusion ModelsabstractThough diffusion models excel in image generation, their step-by-step denoising leads to slow generation speeds. Consistency training addresses this issue with single-step sampling but often produces lower-quality generations and requires high training costs. In this paper, we show that optimizing consistency training loss minimizes the Wasserstein distance between target and generated distributions. As timestep increases, the upper bound accumulates previous consistency training losses. Therefore, larger batch sizes are needed to reduce both current and accumulated losses. We propose Adversarial Consistency Training (ACT), which directly minimizes the Jensen-Shannon (JS) divergence between distributions at each timestep using a discriminator. Theoretically, ACT enhances generation quality, and convergence. By incorporating a discriminator into the consistency training framework, our method achieves improved FID scores on CIFAR10 and ImageNet 64×64 and LSUN Cat 256 ×256 datasets, retains zero-shot image inpainting capabilities, and uses less than 1/6 of the original batch size and fewer than 1/2 of the model parameters and training steps compared to the baseline method, this leads to a substantial reduction in resource consumption. Our code is available: https://github.com/kong13661/ACT Fei Kong, Jinhao Duan, Lichao Sun 0001, Hao Cheng 0015, Renjing Xu, Heng Tao Shen, Xiaofeng Zhu 0001, Xiaoshuang Shi, Kaidi Xu |
CVPR | 8 |
| 2024 | An Efficient Membership Inference Attack for the Diffusion Model by Proximal InitializationabstractRecently, diffusion models have achieved remarkable success in generating tasks, including image and audio generation. However, like other generative models, diffusion models are prone to privacy issues. In this paper, we propose an efficient query-based membership inference attack (MIA), namely Proximal Initialization Attack (PIA), which utilizes groundtruth trajectory obtained by $\epsilon$ initialized in $t=0$ and predicted point to infer memberships. Experimental results indicate that the proposed method can achieve competitive performance with only two queries that achieve at least 6$\times$ efficiency than the previous SOTA baseline on both discrete-time and continuous-time diffusion models. Moreover, previous works on the privacy of diffusion models have focused on vision tasks without considering audio tasks. Therefore, we also explore the robustness of diffusion models to MIA in the text-to-speech (TTS) task, which is an audio generation task. To the best of our knowledge, this work is the first to study the robustness of diffusion models to MIA in the TTS task. Experimental results indicate that models with mel-spectrogram (image-like) output are vulnerable to MIA, while models with audio output are relatively robust to MIA. Code is available at https://github.com/kong13661/PIA. Fei Kong, Jinhao Duan, Ruipeng Ma, Heng Tao Shen, Xiaoshuang Shi, Xiaofeng Zhu 0001, Kaidi Xu |
ICLR | 5 |
| 2024 | On Which Nodes Does GCN Fail? Enhancing GCN From the Node PerspectiveabstractThe label smoothness assumption is at the core of Graph Convolutional Networks (GCNs): nodes in a local region have similar labels. Thus, GCN performs local feature smoothing operation to adhere to this assumption. However, there exist some nodes whose labels obtained by feature smoothing conflict with the label smoothness assumption. We find that the label smoothness assumption and the process of feature smoothing are both problematic on these nodes, and call these nodes out of GCN's control (OOC nodes). In this paper, first, we design the corresponding algorithm to locate the OOC nodes, then we summarize the characteristics of OOC nodes that affect their representation learning, and based on their characteristics, we present DaGCN, an efficient framework that can facilitate the OOC nodes. Extensive experiments verify the superiority of the proposed method and demonstrate that current advanced GCNs are improvements specifically on OOC nodes; the remaining nodes under GCN's control (UC nodes) are already optimally represented by vanilla GCN on most datasets. Jincheng Huang 0005, Jialie Shen 0001, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
ICML | 3 |
| 2024 | Exploring the Role of Node Diversity in Directed Graph Representation Learning
Jincheng Huang 0005, Yujie Mo, Ping Hu 0001, Xiaoshuang Shi, Shangbo Yuan, Xiaofeng Zhu 0001 |
IJCAI | 4 |
| 2024 | Caterpillar: A Pure-MLP Architecture with Shifted-Pillars-ConcatenationabstractModeling in Computer Vision has evolved to MLPs. Vision MLPs naturally lack local modeling capability, to which the simplest treatment is combined with convolutional layers. Convolution, famous for its sliding window scheme, also suffers from this scheme of redundancy and lower parallel computation. In this paper, we seek to dispense with the windowing scheme and introduce a more elaborate and parallelizable method to exploit locality. To this end, we propose a new MLP module, namely Shifted-Pillars-Concatenation (SPC), that consists of two steps of processes: (1) Pillars-Shift, which generates four neighboring maps by shifting the input image along four directions, and (2) Pillars-Concatenation, which applies linear transformations and concatenation on the maps to aggregate local features. SPC module offers superior local modeling power and performance gains, making it a promising alternative to the convolutional layer. Then, we build a pure-MLP architecture called Caterpillar by replacing the convolutional layer with the SPC module in a hybrid model of sMLPNet. Extensive experiments show Caterpillar's excellent performance on both small-scale and ImageNet-1k classification benchmarks, with remarkable scalability and transfer capability possessed as well. The code is available at https://github.com/sunjin19126/Caterpillar. Jin Sun 0005, Xiaoshuang Shi, Zhiyuan Wang 0007, Kaidi Xu, Heng Tao Shen, Xiaofeng Zhu 0001 |
ACM Multimedia | 2 |
| 2024 | Interpretable medical deep framework by logits-constraint attention guiding graph-based multi-scale fusion for Alzheimer's disease analysisabstractDeep learning using structural MRI has been widely applied to early diagnosis study of Alzheimer’s disease. Among existing methods, attention-based 3D subject-level methods can not only provide diagnosis results but also interpret the significant brain regions, thereby attracting considerable attention. However, the performance of previous attention-based methods might be still restricted by: (i) the gap between attention scores and semantic significant regions; (ii) using only single-scale features or simply fusing multi-scale information by addition or concatenation for classification decision-making. To overcome these two issues, we propose an innovative dual-branch model called LA-GMF, which consists of two major modules: logits-constraint attention (LA) and graph-based multi-scale fusion (GMF). The LA module is designed to guide the model to focus on key areas to enhance the diagnostic performance of local lesions, by reducing the inconsistency between attention scores and class prediction probabilities. Meanwhile, by combining the graph neural network and the self-attention mechanism, the GMF module not only introduces the interaction between patches, but also explores the correlation and complementarity between features at different scales, thereby extracting feature representations more comprehensively. Experiments on the popular ADNI and AIBL datasets validate the potential of our model in boosting early AD diagnosis accuracy. Additionally, our interpretation experiments demonstrate the superior interpretability performance of the proposed method over recent state-of-the-art attention-based methods. Our source codes are released at: https://github.com/nollexu/LA-GMF . Jinghao Xu, Chenxi Yuan, Huifang Shang, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
Pattern Recognit. | 5 |
| 2024 | Reverse Graph Learning for Graph Neural NetworkabstractGraph neural networks (GNNs) conduct feature learning by taking into account the local structure preservation of the data to produce discriminative features, but need to address the following issues, i.e., 1) the initial graph containing faulty and missing edges often affect feature learning and 2) most GNN methods suffer from the issue of out-of-example since their training processes do not directly generate a prediction model to predict unseen data points. In this work, we propose a reverse GNN model to learn the graph from the intrinsic space of the original data points as well as to investigate a new out-of-sample extension method. As a result, the proposed method can output a high-quality graph to improve the quality of feature learning, while the new method of out-of-sample extension makes our reverse GNN method available for conducting supervised learning and semi-supervised learning. Experimental results on real-world datasets show that our method outputs competitive classification performance, compared to state-of-the-art methods, in terms of semi-supervised node classification, out-of-sample extension, random edge attack, link prediction, and image retrieval. Rongyao Hu, Fei Kong, Jiangzhang Gan, Yujie Mo, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | GRLC: Graph Representation Learning With ConstraintsabstractContrastive learning has been successfully applied in unsupervised representation learning. However, the generalization ability of representation learning is limited by the fact that the loss of downstream tasks (e.g., classification) is rarely taken into account while designing contrastive methods. In this article, we propose a new contrastive-based unsupervised graph representation learning (UGRL) framework by 1) maximizing the mutual information (MI) between the semantic information and the structural information of the data and 2) designing three constraints to simultaneously consider the downstream tasks and the representation learning. As a result, our proposed method outputs robust low-dimensional representations. Experimental results on 11 public datasets demonstrate that our proposed method is superior over recent state-of-the-art methods in terms of different downstream tasks. Our code is available at https://github.com/LarryUESTC/GRLC. Yujie Mo, Jie Xu 0044, Jialie Shen 0001, Xiaoshuang Shi, Xiaoxiao Li 0001, Heng Tao Shen, Xiaofeng Zhu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Multiplex Graph Representation Learning via Common and Private Information MiningabstractSelf-supervised multiplex graph representation learning (SMGRL) has attracted increasing interest, but previous SMGRL methods still suffer from the following issues: (i) they focus on the common information only (but ignore the private information in graph structures) to lose some essential characteristics related to downstream tasks, and (ii) they ignore the redundant information in node representations of each graph. To solve these issues, this paper proposes a new SMGRL method by jointly mining the common information and the private information in the multiplex graph while minimizing the redundant information within node representations. Specifically, the proposed method investigates the decorrelation losses to extract the common information and minimize the redundant information, while investigating the reconstruction losses to maintain the private information. Comprehensive experimental results verify the superiority of the proposed method, on four public benchmark datasets. Yujie Mo, Zongqian Wu, Yuhuan Chen, Xiaoshuang Shi, Heng Tao Shen, Xiaofeng Zhu 0001 |
AAAI | 4 |
| 2023 | Are Diffusion Models Vulnerable to Membership Inference Attacks?abstractDiffusion-based generative models have shown great potential for image synthesis, but there is a lack of research on the security and privacy risks they may pose. In this paper, we investigate the vulnerability of diffusion models to Membership Inference Attacks (MIAs), a common privacy concern. Our results indicate that existing MIAs designed for GANs or VAE are largely ineffective on diffusion models, either due to inapplicable scenarios (e.g., requiring the discriminator of GANs) or inappropriate assumptions (e.g., closer distances between synthetic samples and member samples). To address this gap, we propose Step-wise Error Comparing Membership Inference (SecMI), a query-based MIA that infers memberships by assessing the matching of forward process posterior estimation at each timestep. SecMI follows the common overfitting assumption in MIA where member samples normally have smaller estimation errors, compared with hold-out samples. We consider both the standard diffusion models, e.g., DDPM, and the text-to-image diffusion models, e.g., Latent Diffusion Models and Stable Diffusion. Experimental results demonstrate that our methods precisely infer the membership with high confidence on both of the two scenarios across multiple different datasets. Code is available at https://github.com/jinhaoduan/SecMI. Jinhao Duan, Fei Kong, Shiqi Wang 0002, Xiaoshuang Shi, Kaidi Xu |
ICML | 4 |
| 2023 | Disentangled Multiplex Graph Representation LearningabstractUnsupervised multiplex graph representation learning (UMGRL) has received increasing interest, but few works simultaneously focused on the common and private information extraction. In this paper, we argue that it is essential for conducting effective and robust UMGRL to extract complete and clean common information, as well as more-complementarity and less-noise private information. To achieve this, we first investigate disentangled representation learning for the multiplex graph to capture complete and clean common information, as well as design a contrastive constraint to preserve the complementarity and remove the noise in the private information. Moreover, we theoretically analyze that the common and private representations learned by our method are provably disentangled and contain more task-relevant and less task-irrelevant information to benefit downstream tasks. Extensive experiments verify the superiority of the proposed method in terms of different downstream tasks. Yujie Mo, Yajie Lei, Jialie Shen 0001, Xiaoshuang Shi, Heng Tao Shen, Xiaofeng Zhu 0001 |
ICML | 4 |
| 2023 | Improve Video Representation with Temporal Adversarial AugmentationabstractRecent works reveal that adversarial augmentation benefits the generalization of neural networks (NNs) if used in an appropriate manner. In this paper, we introduce Temporal Adversarial Augmentation (TA), a novel video augmentation technique that utilizes temporal attention. Unlike conventional adversarial augmentation, TA is specifically designed to shift the attention distributions of neural networks with respect to video clips by maximizing a temporal-related loss function. We demonstrate that TA will obtain diverse temporal views, which significantly affect the focus of neural networks. Training with these examples remedies the flaw of unbalanced temporal information perception and enhances the ability to defend against temporal shifts, ultimately leading to better generalization. To leverage TA, we propose Temporal Video Adversarial Fine-tuning (TAF) framework for improving video representations. TAF is a model-agnostic, generic, and interpretability-friendly training strategy. We evaluate TAF with four powerful models (TSM, GST, TAM, and TPN) over three challenging temporal-related benchmarks (Something-something V1&V2 and diving48). Experimental results demonstrate that TAF effectively improves the test accuracy of these models with notable margins without introducing additional parameters or computational costs. As a byproduct, TAF also improves the robustness under out-of-distribution (OOD) settings. Code is available at https://github.com/jinhaoduan/TAF. Jinhao Duan, Quanfu Fan, Hao Cheng 0015, Xiaoshuang Shi, Kaidi Xu |
IJCAI | 4 |
| 2023 | Co-assistant Networks for Label Correction
Weiheng Fu, Xiaoshuang Shi, Heng Tao Shen, Xiaofeng Zhu 0001 |
MICCAI (3) | 4 |
| 2023 | Segment Membranes and Nuclei from Histopathological Images via Nuclei Point-Level Supervision
Hansheng Li, Xiaoshuang Shi, Yuxin Kang, Qirong Bu, Hong Lv, Mingzhen Lin, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (6) | 4 |
| 2023 | Self-Weighted Contrastive Learning among Multiple Views for Mitigating Representation DegenerationabstractRecently, numerous studies have demonstrated the effectiveness of contrastive learning (CL), which learns feature representations by pulling in positive samples while pushing away negative samples. Many successes of CL lie in that there exists semantic consistency between data augmentations of the same instance. In multi-view scenarios, however, CL might cause representation degeneration when the collected multiple views inherently have inconsistent semantic information or their representations subsequently do not capture sufficient discriminative information. To address this issue, we propose a novel framework called SEM: SElf-weighted Multi-view contrastive learning with reconstruction regularization. Specifically, SEM is a general framework where we propose to first measure the discrepancy between pairwise representations and then minimize the corresponding self-weighted contrastive loss, and thus making SEM adaptively strengthen the useful pairwise views and also weaken the unreliable pairwise views. Meanwhile, we impose a self-supervised reconstruction term to regularize the hidden features of encoders, to assist CL in accessing sufficient discriminative information of data. Experiments on public multi-view datasets verified that SEM can mitigate representation degeneration in existing CL methods and help them achieve significant performance improvements. Ablation studies also demonstrated the effectiveness of SEM with different options of weighting strategies and reconstruction terms. Jie Xu 0044, Shuo Chen 0003, Yazhou Ren 0001, Xiaoshuang Shi, Heng Tao Shen, Gang Niu 0001, Xiaofeng Zhu 0001 |
NeurIPS | 4 |
| 2023 | IGCNN-FC: Boosting interpretability and generalization of convolutional neural networks for few chest X-rays analysis
Mengmeng Zhan, Xiaoshuang Shi, Rongyao Hu |
Inf. Process. Manag. | 2 |
| 2023 | Multi-scale representation attention based deep multiple instance learning for gigapixel whole slide image analysis
Hangchen Xiang, Qingguo Yan, Meilian Xu, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
Medical Image Anal. | 5 |
| 2023 | Self-paced resistance learning against overfitting on noisy labels
Xiaoshuang Shi, Zhenhua Guo 0001, Kang Li 0004, Yun Liang 0012, Xiaofeng Zhu 0001 |
Pattern Recognit. | 1 |
| 2023 | Adaptive Feature Projection With Distribution Alignment for Deep Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering (IMVC) analysis, where some views of multi-view data usually have missing data, has attracted increasing attention. However, existing IMVC methods still have two issues: 1) they pay much attention to imputing or recovering the missing data, without considering the fact that the imputed values might be inaccurate due to the unknown label information, 2) the common features of multiple views are always learned from the complete data, while ignoring the feature distribution discrepancy between the complete and incomplete data. To address these issues, we propose an imputation-free deep IMVC method and consider distribution alignment in feature learning. Concretely, the proposed method learns the features for each view by autoencoders and utilizes an adaptive feature projection to avoid the imputation for missing data. All available data are projected into a common feature space, where the common cluster information is explored by maximizing mutual information and the distribution alignment is achieved by minimizing mean discrepancy. Additionally, we design a new mean discrepancy loss for incomplete multi-view learning and make it applicable in mini-batch optimization. Extensive experiments demonstrate that our method achieves the comparable or superior performance compared with state-of-the-art methods. Jie Xu 0044, Chao Li 0034, Yazhou Ren 0001, Xiaoshuang Shi, Heng Tao Shen, Xiaofeng Zhu 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Multiplex Graph Representation Learning Via Dual Correlation ReductionabstractRecently, with the superior capacity for analyzing the multiplex graph data, self-supervised multiplex graph representation learning (SMGRL) has received much interest. However, existing SMGRL methods are still limited by the following issues: (i) they generally ignore the noisy information within each graph and the common information among different graphs, thus weakening the effectiveness of SMGRL, and (ii) they conduct negative sample encoding and complex pretext tasks for contrastive learning, thus weakening the efficiency of SMGRL. To solve these issues, in this work, we propose a new framework to conduct effective and efficient SMGRL. Specifically, the proposed method investigates the intra-graph and inter-graph decorrelation losses, respectively, for reducing the impact of noisy information within each graph and capturing the common information among different graphs, to achieve the effectiveness. Moreover, the proposed method does not need negative samples for the SMGRL and designs a simple pretext task, to achieve the efficiency. We further theoretically justify that our method achieves the maximal mutual information instead of directly conducting contrastive learning and theoretically justify that our method actually minimizes the multiplex graph information bottleneck, which guarantees the effectiveness. In addition, an extension for semi-supervised scenarios is proposed to fit the case that a few labels are provided in reality. Extensive experimental results verify the effectiveness and efficiency of the proposed method with respect to various downstream tasks. Yujie Mo, Yuhuan Chen, Yajie Lei, Xiaoshuang Shi, Chang-an Yuan 0001, Xiaofeng Zhu 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Why does batch normalization induce the model vulnerability on adversarial images?
Fei Kong, Kaidi Xu, Xiaoshuang Shi |
World Wide Web (WWW) | 4 |
| 2022 | Simple Unsupervised Graph Representation LearningabstractIn this paper, we propose a simple unsupervised graph representation learning method to conduct effective and efficient contrastive learning. Specifically, the proposed multiplet loss explores the complementary information between the structural information and neighbor information to enlarge the inter-class variation, as well as adds an upper bound loss to achieve the finite distance between positive embeddings and anchor embeddings for reducing the intra-class variation. As a result, both enlarging inter-class variation and reducing intra-class variation result in small generalization error, thereby obtaining an effective model. Furthermore, our method removes widely used data augmentation and discriminator from previous graph contrastive learning methods, meanwhile available to output low-dimensional embeddings, leading to an efficient model. Experimental results on various real-world datasets demonstrate the effectiveness and efficiency of our method, compared to state-of-the-art methods. The source codes are released at https://github.com/YujieMo/SUGRL. Yujie Mo, Jie Xu 0044, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
AAAI | 4 |
| 2022 | Deep Incomplete Multi-View Clustering via Mining Cluster ComplementarityabstractIncomplete multi-view clustering (IMVC) is an important unsupervised approach to group the multi-view data containing missing data in some views. Previous IMVC methods suffer from the following issues: (1) the inaccurate imputation or padding for missing data negatively affects the clustering performance, (2) the quality of features after fusion might be interfered by the low-quality views, especially the inaccurate imputed views. To avoid these issues, this work presents an imputation-free and fusion-free deep IMVC framework. First, the proposed method builds a deep embedding feature learning and clustering model for each view individually. Our method then nonlinearly maps the embedding features of complete data into a high-dimensional space to discover linear separability. Concretely, this paper provides an implementation of the high-dimensional mapping as well as shows the mechanism to mine the multi-view cluster complementarity. This complementary information is then transformed to the supervised information with high confidence, aiming to achieve the multi-view clustering consistency for the complete data and incomplete data. Furthermore, we design an EM-like optimization strategy to alternately promote feature learning and clustering. Extensive experiments on real-world multi-view datasets demonstrate that our method achieves superior clustering performance over state-of-the-art methods. Jie Xu 0044, Chao Li 0034, Yazhou Ren 0001, Yujie Mo, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
AAAI | 6 |
| 2022 | Invariant Content Synergistic Learning for Domain Generalization on Medical Image SegmentationabstractAlthough deep convolution neural networks (DC-NNs) can achieve remarkable success on medical image segmentation, their performance might significantly deteriorate when confronting testing data with the new distribution. Recent studies suggest that one major cause of this issue is the strong inductive bias of DCNNs, which towards image styles (e.g., superficial texture) that are sensitive to change, instead of the invariant content (e.g., object shapes). Inspired by this, we propose a novel method, named Invariant Content Synergistic Learning (ICSL), to improve the generalization ability of DCNNs on unseen data by controlling the inductive bias. Specifically, ICSL first mixes the style of training instances to perturb the training distribution, so that more diverse domains or styles would be made available for training DCNNs. Then, based on the perturbed distribution, we carefully design a dual-branches invariant content synergistic learning strategy to prevent style-biased predictions and maintain the invariant content. Extensive experimental results demonstrate the superior performance of the proposed method over state-of-the-art domain generalization methods on two typical medical segmentation tasks. Yuxin Kang, Hansheng Li, Xiaoshuang Shi, Feihong Liu, Qingguo Yan, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
BIBM | 4 |
| 2022 | Dual-Graph Learning Convolutional Networks for Interpretable Alzheimer's Disease Diagnosis
Tingsong Xiao, Xiaoshuang Shi, Xiaofeng Zhu 0001, Guorong Wu 0001 |
MICCAI (8) | 3 |
| 2022 | Complementary Graph Representation Learning for Functional Neuroimaging IdentificationabstractThe functional connectomics study on resting state functional magnetic resonance imaging (rs-fMRI) data has become a popular way for early disease diagnosis. However, previous methods did not jointly consider the global patterns, the local patterns, and the temporal information of the blood-oxygen-level-dependent (BOLD) signals, thereby restricting the model effectiveness for early disease diagnosis. In this paper, we propose a new graph convolutional network (GCN) method to capture local and global patterns for conducting dynamically functional connectivity analysis. Specifically, we first employ the sliding window method to partition the original BOLD signals into multiple segments, aiming at achieving the dynamically functional connectivity analysis, and then design a multi-view node classification and a temporal graph classification to output two kinds of representations, which capture the temporally global patterns and the temporally local patterns, respectively. We further fuse these two kinds of representation by the weighted concatenation method whose effectiveness is experimentally proved as well. Experimental results on real datasets demonstrate the effectiveness of our method, compared to comparison methods on different classification tasks. Rongyao Hu, Jiangzhang Gan, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
ACM Multimedia | 4 |
| 2022 | Simple Self-supervised Multiplex Graph Representation LearningabstractSelf-supervised multiplex graph representation learning (SMGRL) aims to capture the information from the multiplex graph, and generates discriminative embedding without labels. However, previous SMGRL methods still suffer from the issues of efficiency and effectiveness due to the processes, e.g., data augmentation, negative sample encoding, complex pretext tasks, etc. In this paper, we propose a simple method to achieve efficient and effective SMGRL. Specifically, the proposed method removes the processes (i.e., data augmentation and negative sample encoding) for the SMGRL and designs a simple pretext task, for achieving the efficiency. Moreover, the proposed method also designs an intra-graph decorrelation loss and an inter-graph decorrelation loss, respectively, to capture the common information within individual graphs and the common information across graphs, for achieving the effectiveness. Extensive experimental results verify the efficiency and effectiveness of our method, compared to 11 comparison methods on 4 public benchmark datasets, on the node classification task. Yujie Mo, Yuhuan Chen, Xiaoshuang Shi, Xiaofeng Zhu 0001 |
ACM Multimedia | 4 |
| 2022 | Multi-task multi-modality SVM for early COVID-19 Diagnosis using chest CT data
Rongyao Hu, Jiangzhang Gan, Xiaofeng Zhu 0001, Tong Liu 0016, Xiaoshuang Shi |
Inf. Process. Manag. | 5 |
| 2022 | Robust convolutional neural networks against adversarial attacks on medical imagesabstractConvolutional neural networks (CNNs) have been widely applied to medical images. However, medical images are vulnerable to adversarial attacks by perturbations that are undetectable to human experts. This poses significant security risks and challenges to CNN-based applications in clinic practice. In this work, we quantify the scale of adversarial perturbation imperceptible to clinical practitioners and investigate the cause of the vulnerability in CNNs. Specifically, we discover that noise (i.e., irrelevant or corrupted discriminative information) in medical images might be a key contributor to performance deterioration of CNNs against adversarial perturbations, as noisy features are learned unconsciously by CNNs in feature representations and magnified by adversarial perturbations. In response, we propose a novel defense method by embedding sparsity denoising operators in CNNs for improved robustness. Tested with various state-of-the-art attacking methods on two distinct medical image modalities, we demonstrate that the proposed method can successfully defend against those unnoticeable adversarial attacks by retaining as much as over 90% of its original performance. We believe our findings are critical for improving and deploying CNN-based medical applications in real-world scenarios. Xiaoshuang Shi, Yifan Peng 0002, Qingyu Chen 0001, Tiarnan D. Keenan, Alisa T. Thavikulwat, Sungwon Lee 0003, Yuxing Tang, Emily Y. Chew, Ronald M. Summers, Zhiyong Lu |
Pattern Recognit. | 1 |
| 2021 | Automatic whole slide pathology image diagnosis framework via unit stochastic selection and attention fusion
Pingjun Chen, Yun Liang 0012, Xiaoshuang Shi, Lin Yang 0002, Paul D. Gader |
Neurocomputing | 3 |
| 2021 | Text-Guided Neural Network Training for Image Recognition in Natural Scenes and MedicineabstractConvolutional neural networks (CNNs) are widely recognized as the foundation for machine vision systems. The conventional rule of teaching CNNs to understand images requires training images with human annotated labels, without any additional instructions. In this article, we look into a new scope and explore the guidance from text for neural network training. We present two versions of attention mechanisms to facilitate interactions between visual and semantic information and encourage CNNs to effectively distill visual features by leveraging semantic features. In contrast to dedicated text-image joint embedding methods, our method realizes asynchronous training and inference behavior: a trained model can classify images, irrespective of the text availability. This characteristic substantially improves the model scalability to multiple (multimodal) vision tasks. We also apply the proposed method onto medical imaging, which learns from richer clinical knowledge and achieves attention-based interpretable decision-making. With comprehensive validation on two natural and two medical datasets, we demonstrate that our method can effectively make use of semantic knowledge to improve CNN performance. Our method performs substantial improvement on medical image datasets. Meanwhile, it achieves promising performance for multi-label image classification and caption-image retrieval as well as excellent performance for phrase-based and multi-object localization on public benchmarks. Zizhao Zhang 0002, Pingjun Chen, Xiaoshuang Shi, Lin Yang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Loss-Based Attention for Interpreting Image-Level Prediction of Convolutional Neural NetworksabstractAlthough deep neural networks have achieved great success on numerous large-scale tasks, poor interpretability is still a notorious obstacle for practical applications. In this paper, we propose a novel and general attention mechanism, loss-based attention, upon which we modify deep neural networks to mine significant image patches for explaining which parts determine the image decision-making. This is inspired by the fact that some patches contain significant objects or their parts for image-level decision. Unlike previous attention mechanisms that adopt different layers and parameters to learn weights and image prediction, the proposed loss-based attention mechanism mines significant patches by utilizing the same parameters to learn patch weights and logits (class vectors), and image prediction simultaneously, so as to connect the attention mechanism with the loss function for boosting the patch precision and recall. Additionally, different from previous popular networks that utilize max-pooling or stride operations in convolutional layers without considering the spatial relationship of features, the modified deep architectures first remove them to preserve the spatial relationship of image patches and greatly reduce their dependencies, and then add two convolutional or capsule layers to extract their features. With the learned patch weights, the image-level decision of the modified deep architectures is the weighted sum on patches. Extensive experiments on large-scale benchmark databases demonstrate that the proposed architectures can obtain better or competitive performance to state-of-the-art baseline networks with better interpretability. The source codes are available on: https://github.com/xsshi2015/Loss-based-Attention-for-Interpreting-Image-level-Prediction-of-Convolutional-Neural-Networks. Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Pingjun Chen, Yun Liang 0012, Zhiyong Lu, Zhenhua Guo 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | A Scalable Optimization Mechanism for Pairwise Based Discrete HashingabstractMaintaining the pairwise relationship among originally high-dimensional data into a low-dimensional binary space is a popular strategy to learn binary codes. One simple and intuitive method is to utilize two identical code matrices produced by hash functions to approximate a pairwise real label matrix. However, the resulting quartic problem in term of hash functions is difficult to directly solve due to the non-convex and non-smooth nature of the objective. In this paper, unlike previous optimization methods using various relaxation strategies, we aim to directly solve the original quartic problem using a novel alternative optimization mechanism to linearize the quartic problem by introducing a linear regression model. Additionally, we find that gradually learning each batch of binary codes in a sequential mode, i.e. batch by batch, is greatly beneficial to the convergence of binary code learning. Based on this significant discovery and the proposed strategy, we introduce a scalable symmetric discrete hashing algorithm that gradually and smoothly updates each batch of binary codes. To further improve the smoothness, we also propose a greedy symmetric discrete hashing algorithm to update each bit of batch binary codes. Moreover, we extend the proposed optimization mechanism to solve the non-convex optimization problems for binary code learning in many other pairwise based hashing algorithms. Extensive experiments on benchmark single-label and multi-label databases demonstrate the superior performance of the proposed mechanism over recent state-of-the-art methods on two kinds of retrieval tasks: similarity and ranking order. The source codes are available on https://github.com/xsshi2015/Scalable-Pairwise-based-Discrete-Hashing. Xiaoshuang Shi, Fuyong Xing, Zizhao Zhang 0002, Manish Sapkota, Zhenhua Guo 0001, Lin Yang 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | Loss-Based Attention for Deep Multiple Instance LearningabstractAlthough attention mechanisms have been widely used in deep learning for many tasks, they are rarely utilized to solve multiple instance learning (MIL) problems, where only a general category label is given for multiple instances contained in one bag. Additionally, previous deep MIL methods firstly utilize the attention mechanism to learn instance weights and then employ a fully connected layer to predict the bag label, so that the bag prediction is largely determined by the effectiveness of learned instance weights. To alleviate this issue, in this paper, we propose a novel loss based attention mechanism, which simultaneously learns instance weights and predictions, and bag predictions for deep multiple instance learning. Specifically, it calculates instance weights based on the loss function, e.g. softmax+cross-entropy, and shares the parameters with the fully connected layer, which is to predict instance and bag predictions. Additionally, a regularization term consisting of learned weights and cross-entropy functions is utilized to boost the recall of instances, and a consistency cost is used to smooth the training process of neural networks for boosting the model generalization performance. Extensive experiments on multiple types of benchmark databases demonstrate that the proposed attention mechanism is a general, effective and efficient framework, which can achieve superior bag and image classification performance over other state-of-the-art MIL methods, with obtaining higher instance precision and recall than previous attention mechanisms. Source codes are available on https://github.com/xsshi2015/Loss-Attention. Xiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Zizhao Zhang 0002, Lei Cui 0004, Lin Yang 0002 |
AAAI | 1 |
| 2020 | A Novel Loss Calibration Strategy for Object Detection Networks Training on Sparsely Annotated Pathological Datasets
Hansheng Li, Yuxin Kang, Xiaoshuang Shi, Mengdi Yan, Zixu Tong, Qirong Bu, Lei Cui 0004, Jun Feng 0003, Lin Yang 0002 |
MICCAI (5) | 4 |
| 2020 | Anchor-Based Self-Ensembling for Semi-Supervised Deep Pairwise Hashing
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Yun Liang 0012, Lin Yang 0002 |
Int. J. Comput. Vis. | 1 |
| 2020 | Graph temporal ensembling based semi-supervised convolutional neural network with noisy labels for histopathology image analysis
Xiaoshuang Shi, Hai Su, Fuyong Xing, Yun Liang 0012, Gang Qu 0002, Lin Yang 0002 |
Medical Image Anal. | 1 |
| 2019 | Local and Global Consistency Regularized Mean Teacher for Semi-supervised Nuclei Classification
Hai Su, Xiaoshuang Shi, Jinzheng Cai, Lin Yang 0002 |
MICCAI (1) | 2 |
| 2019 | Towards pixel-to-pixel deep nucleus detection in microscopy imagesabstractBACKGROUND: Nucleus is a fundamental task in microscopy image analysis and supports many other quantitative studies such as object counting, segmentation, tracking, etc. Deep neural networks are emerging as a powerful tool for biomedical image computing; in particular, convolutional neural networks have been widely applied to nucleus/cell detection in microscopy images. However, almost all models are tailored for specific datasets and their applicability to other microscopy image data remains unknown. Some existing studies casually learn and evaluate deep neural networks on multiple microscopy datasets, but there are still several critical, open questions to be addressed. RESULTS: We analyze the applicability of deep models specifically for nucleus detection across a wide variety of microscopy image data. More specifically, we present a fully convolutional network-based regression model and extensively evaluate it on large-scale digital pathology and microscopy image datasets, which consist of 23 organs (or cancer diseases) and come from multiple institutions. We demonstrate that for a specific target dataset, training with images from the same types of organs might be usually necessary for nucleus detection. Although the images can be visually similar due to the same staining technique and imaging protocol, deep models learned with images from different organs might not deliver desirable results and would require model fine-tuning to be on a par with those trained with target data. We also observe that training with a mixture of target and other/non-target data does not always mean a higher accuracy of nucleus detection, and it might require proper data manipulation during model training to achieve good performance. CONCLUSIONS: We conduct a systematic case study on deep models for nucleus detection in a wide variety of microscopy images, aiming to address several important but previously understudied questions. We present and extensively evaluate an end-to-end, pixel-to-pixel fully convolutional regression network and report a few significant findings, some of which might have not been reported in previous studies. The model performance analysis and observations would be helpful to nucleus detection in microscopy images. Fuyong Xing, Yuanpu Xie, Xiaoshuang Shi, Pingjun Chen, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 3 |
| 2019 | Correction to: Towards pixel-to-pixel deep nucleus detection in microscopy imagesabstractFollowing publication of the original article [1], we have been notified of a few errors in the html version. Fuyong Xing, Yuanpu Xie, Xiaoshuang Shi, Pingjun Chen, Zizhao Zhang 0002, Lin Yang 0002 |
BMC Bioinform. | 3 |
| 2019 | Structured orthogonal matching pursuit for feature selection
Xiaoshuang Shi, Fuyong Xing, Zhenhua Guo 0001, Hai Su, Fujun Liu, Lin Yang 0002 |
Neurocomputing | 1 |
| 2019 | Similarity mapping for robust face recognition via a single training sample per person
Qin Li 0001, Xiaoshuang Shi, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Deep Convolutional Hashing for Low-Dimensional Binary Embedding of Histopathological ImagesabstractCompact binary representations of histopa-thology images using hashing methods provide efficient approximate nearest neighbor search for direct visual query in large-scale databases. They can be utilized to measure the probability of the abnormality of the query image based on the retrieved similar cases, thereby providing support for medical diagnosis. They also allow for efficient managing of large-scale image databases because of a low storage requirement. However, the effectiveness of binary representations heavily relies on the visual descriptors that represent the semantic information in the histopathological images. Traditional approaches with hand-crafted visual descriptors might fail due to significant variations in image appearance. Recently, deep learning architectures provide promising solutions to address this problem using effective semantic representations. In this paper, we propose a deep convolutional hashing method that can be trained "point-wise" to simultaneously learn both semantic and binary representations of histopathological images. Specifically, we propose a convolutional neural network that introduces a latent binary encoding (LBE) layer for low-dimensional feature embedding to learn binary codes. We design a joint optimization objective function that encourages the network to learn discriminative representations from the label information, and reduce the gap between the real-valued low-dimensional embedded features and desired binary values. The binary encoding for new images can be obtained by forward propagating through the network and quantizing the output of the LBE layer. Experimental results on a large-scale histopathological image dataset demonstrate the effectiveness of the proposed method. Manish Sapkota, Xiaoshuang Shi, Fuyong Xing, Lin Yang 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Iterative Attention Mining for Weakly Supervised Thoracic Disease Pattern Localization in Chest X-Rays
Jinzheng Cai, Le Lu 0001, Adam P. Harrison, Xiaoshuang Shi, Pingjun Chen, Lin Yang 0002 |
MICCAI (2) | 4 |
| 2018 | Robust principal component analysis via optimal mean by joint ℓ2, 1 and Schatten p-norms minimization
Xiaoshuang Shi, Feiping Nie 0001, Zhihui Lai 0001, Zhenhua Guo 0001 |
Neurocomputing | 1 |
| 2018 | Efficient and robust cell detection: A structured regression approach
Yuanpu Xie, Fuyong Xing, Xiaoshuang Shi, Xiangfei Kong, Hai Su, Lin Yang 0002 |
Medical Image Anal. | 3 |
| 2018 | Self-learning for face clustering
Xiaoshuang Shi, Zhenhua Guo 0001, Fuyong Xing, Jinzheng Cai, Lin Yang 0002 |
Pattern Recognit. | 1 |
| 2018 | Pairwise based deep ranking hashing for histopathology image classification and retrieval
Xiaoshuang Shi, Manish Sapkota, Fuyong Xing, Fujun Liu, Lei Cui 0004, Lin Yang 0002 |
Pattern Recognit. | 1 |
| 2018 | Revisiting graph construction for fast image segmentation
Zizhao Zhang 0002, Fuyong Xing, Hanzi Wang, Yan Yan 0001, Xiaoshuang Shi, Lin Yang 0002 |
Pattern Recognit. | 6 |
| 2017 | Asymmetric Discrete Graph HashingabstractRecently, many graph based hashing methods have been emerged to tackle large-scale problems. However, there exists two major bottlenecks: (1) directly learning discrete hashing codes is an NP-hardoptimization problem; (2) the complexity of both storage and computational time to build a graph with n data points is O(n2). To address these two problems, in this paper, we propose a novel yetsimple supervised graph based hashing method, asymmetric discrete graph hashing, by preserving the asymmetric discrete constraint and building an asymmetric affinity matrix to learn compact binary codes.Specifically, we utilize two different instead of identical discrete matrices to better preserve the similarity of the graph with short binary codes. We generate the asymmetric affinity matrix using m (m << n) selected anchors to approximate the similarity among all training data so that computational time and storage requirement can be significantly improved. In addition, the proposed method jointly learns discrete binary codes and a low-dimensional projection matrix to further improve the retrieval accuracy. Extensive experiments on three benchmark large-scale databases demonstrate its superior performance over the recent state of the arts with lower training time costs. Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Manish Sapkota, Lin Yang 0002 |
AAAI | 1 |
| 2017 | Cell Encoding for Histopathology Image Classification
Xiaoshuang Shi, Fuyong Xing, Yuanpu Xie, Hai Su, Lin Yang 0002 |
MICCAI (2) | 1 |
| 2017 | Supervised graph hashing for histopathology image retrieval and classification
Xiaoshuang Shi, Fuyong Xing, Kaidi Xu, Yuanpu Xie, Hai Su, Lin Yang 0002 |
Medical Image Anal. | 1 |
| 2017 | Active learning via local structure reconstruction
Qin Li 0001, Xiaoshuang Shi, Linfei Zhou, Zhifeng Bao, Zhenhua Guo 0001 |
Pattern Recognit. Lett. | 2 |
| 2016 | SemiContour: A Semi-Supervised Learning Approach for Contour DetectionabstractSupervised contour detection methods usually require many labeled training images to obtain satisfactory performance. However, a large set of annotated data might be unavailable or extremely labor intensive. In this paper, we investigate the usage of semi-supervised learning (SSL) to obtain competitive detection accuracy with very limited training data (three labeled images). Specifically, we propose a semi-supervised structured ensemble learning approach for contour detection built on structured random forests (SRF). To allow SRF to be applicable to unlabeled data, we present an effective sparse representation approach to capture inherent structure in image patches by finding a compact and discriminative low-dimensional subspace representation in an unsupervised manner, enabling the incorporation of abundant unlabeled patches with their estimated structured labels to help SRF perform better node splitting. We re-examine the role of sparsity and propose a novel and fast sparse coding algorithm to boost the overall learning efficiency. To the best of our knowledge, this is the first attempt to apply SSL for contour detection. Extensive experiments on the BSDS500 segmentation dataset and the NYU Depth dataset demonstrate the superiority of the proposed method. Zizhao Zhang 0002, Fuyong Xing, Xiaoshuang Shi, Lin Yang 0002 |
CVPR | 3 |
| 2016 | Kernel-Based Supervised Discrete Hashing for Image Retrieval
Xiaoshuang Shi, Fuyong Xing, Jinzheng Cai, Zizhao Zhang 0002, Yuanpu Xie, Lin Yang 0002 |
ECCV (7) | 1 |
| 2016 | Transfer Shape Modeling Towards High-Throughput Microscopy Image Segmentation
Fuyong Xing, Xiaoshuang Shi, Zizhao Zhang 0002, Jinzheng Cai, Yuanpu Xie, Lin Yang 0002 |
MICCAI (3) | 2 |
| 2016 | Two-Dimensional Whitening Reconstruction for Enhancing Robustness of Principal Component AnalysisabstractPrincipal component analysis (PCA) is widely applied in various areas, one of the typical applications is in face. Many versions of PCA have been developed for face recognition. However, most of these approaches are sensitive to grossly corrupted entries in a 2D matrix representing a face image. In this paper, we try to reduce the influence of grosses like variations in lighting, facial expressions and occlusions to improve the robustness of PCA. In order to achieve this goal, we present a simple but effective unsupervised preprocessing method, two-dimensional whitening reconstruction (TWR), which includes two stages: 1) A whitening process on a 2D face image matrix rather than a concatenated 1D vector; 2) 2D face image matrix reconstruction. TWR reduces the pixel redundancy of the internal image, meanwhile maintains important intrinsic features. In this way, negative effects introduced by gross-like variations are greatly reduced. Furthermore, the face image with TWR preprocessing could be approximate to a Gaussian signal, on which PCA is more effective. Experiments on benchmark face databases demonstrate that the proposed method could significantly improve the robustness of PCA methods on classification and clustering, especially for the faces with severe illumination changes. Xiaoshuang Shi, Zhenhua Guo 0001, Feiping Nie 0001, Lin Yang 0002, Jane You, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Within-class penalty based multi-class support vector machineabstractSupport vector machine (SVM) is a widely used maximum margin classifier, but the classification performance is largely affected by outliers. In this paper, we propose a novel multi-class SVM method to reduce the influence of outliers on the classification performance. Our proposed method includes an efficient optimization model via considering the within-class scatter and an optimization way. Specifically, the method is based on one assumption that penalizing the within-class scatter can reduce the number of misclassified outliers near the decision boundary, because data points of each class could be compacted by the within-class penalty. Experiments on benchmark databases demonstrate the effectiveness of the assumption and the proposed method. Xiaoshuang Shi, Zhenhua Guo 0001, Yujiu Yang 0001, Lin Yang 0002 |
ICIP | 1 |
| 2015 | A Framework of Joint Graph Embedding and Sparse Regression for Dimensionality ReductionabstractOver the past few decades, a large number of algorithms have been developed for dimensionality reduction. Despite the different motivations of these algorithms, they can be interpreted by a common framework known as graph embedding. In order to explore the significant features of data, some sparse regression algorithms have been proposed based on graph embedding. However, the problem is that these algorithms include two separate steps: (1) embedding learning and (2) sparse regression. Thus their performance is largely determined by the effectiveness of the constructed graph. In this paper, we present a framework by combining the objective functions of graph embedding and sparse regression so that embedding learning and sparse regression can be jointly implemented and optimized, instead of simply using the graph spectral for sparse regression. By the proposed framework, supervised, semisupervised, and unsupervised learning algorithms could be unified. Furthermore, we analyze two situations of the optimization problem for the proposed framework. By adopting an ℓ2,1-norm regularization for the proposed framework, it can perform feature selection and subspace learning simultaneously. Experiments on seven standard databases demonstrate that joint graph embedding and sparse regression method can significantly improve the recognition performance and consistently outperform the sparse regression method. Xiaoshuang Shi, Zhenhua Guo 0001, Zhihui Lai 0001, Yujiu Yang 0001, Zhifeng Bao, David Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Face recognition by sparse discriminant analysis via joint L2, 1-norm minimization
Xiaoshuang Shi, Yujiu Yang 0001, Zhenhua Guo 0001, Zhihui Lai 0001 |
Pattern Recognit. | 1 |