Shihui Ying

dblp:52/2125 · DBLP profile ↗
← Back
103ranked-venue papers
8as first author
73since 2021 · last 2027
0000-0001-9423-0146ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 7 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 31 · 1 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2027 URG-MRI: Unpaired reference guided MRI reconstruction
Xuanmin Chen, Jingtian Gu, Shihui Ying, Liyan Ma
Expert Syst. Appl.4
2026 SCADA: Sparse cross attention for domain adaptive semantic segmentation
Qizhe Fan, Xiaoqin Shen, Yuanbo Chen, Shihui Ying, Jue Jiang, Shaoyi Du
Neural Networks4
2026 Knowledge-Embedded Hypergraph Neural Networks
abstract
Hypergraph Neural Networks (HGNNs) enhance graph-based modeling by representing complex relationships, with applications in brain network analysis, recommendation systems, and computer vision. However, conventional HGNNs often struggle with effective knowledge extraction and discriminative feature representation, leading to performance limitations. This paper presents Knowledge-Embedded Hypergraph Neural Networks (Knowledge HGNN), a framework that addresses these challenges with two complementary encoders and a multi-dimensional fusion strategy. The High-Order Incidence Encoder (HOI-Encoder) explicitly embeds structural knowledge by capturing permutation-invariant high-order incidence patterns that are typically overlooked by standard HGNNs. In contrast, the Task-Driven Rule Encoder (TDR-Encoder) focuses on feature-level knowledge, extracting task-related rules from vertex attributes through gradient boosted decision tree pre-training and encoding both rule content and positional importance. A Multi-Dimensional Knowledge Fusion module then integrates structural and rule-based embeddings, bridging semantic and dimensional gaps to form enriched vertex representations. The framework includes two implementations: Rule-Driven HGNN, which emphasizes rule-based knowledge, and Dual-Driven HGNN, which jointly leverages structural and rule-based knowledge for comprehensive feature extraction. Extensive experiments on ten datasets, together with ablation studies, demonstrate that Knowledge HGNN significantly improves performance, achieving a 7.3% gain on the Cora dataset and an average improvement of 2.5% across all datasets. These results highlight the effectiveness of explicitly differentiating and fusing structural and rule-based knowledge, setting a new standard for hypergraph applications in complex, data-driven scenarios.
Yifan Feng 0001, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 HGNN Shield: Defending Hypergraph Neural Networks Against High-Order Structure Attack
abstract
Hypergraph Neural Networks (HGNNs) are crucial in modeling complex high-order correlations in diverse domains, utilizing hyperedges that connect multiple vertices. However, their susceptibility to structural attacks and irrational connections can disrupt message propagation and degrade performance. To address these issues, we introduce the HGNN Shield, a defense framework incorporating two key modules: Hyperedge-Dependent Estimation (HDE) and High-Order Shield (HOS). The HDE module prioritizes vertex dependencies within hyperedges and adapts traditional connectivity measures to hypergraphs, facilitating precise structural modifications. This adaptation allows for a nuanced assessment of vertex relationships within hyperedges, contributing theoretically by extending classical graph-based connection dependency measures to hypergraphs. Following HDE, the HOS module, positioned before convolutional layers, consists of three submodules: Hyperpath Cut, Hyperpath Link, and Hyperpath Refine. These components collectively detect, disconnect, and refine adversarial connections, ensuring robust message propagation. The theoretical contribution of the HOS module lies in maintaining hyperpath integrity and learning trajectory under adversarial conditions, providing a certifiable defense mechanism against high-order structural attacks. Experiments on six hypergraph datasets indicate that HGNN Shield significantly enhances robustness and maintains data integrity against targeted attacks, outperforming existing methods (an average performance improvement of 9.33% over other methods). Our framework not only improves HGNN reliability but also advances security in hypergraph-based applications.
Yifan Feng 0001, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 HGNNv2: Stable Hypergraph Neural Networks
abstract
Hypergraph neural networks (HGNNs) are widely used models for analyzing higher-order relational data. HGNNs suffer from the rapid performance degradation with increasing layers. Hypergraph dynamic system (HDS) is a potential way to deal with this challenge. However, hypergraph dynamic system is confined to a time-continuous isotropic model, lacking positional information in the structural space of the hypergraph. In contrast, anisotropic diffusion can capture structural space differences among vertices, providing a more precise representation of the information propagation process in hypergraph structures than isotropic diffusion. In this paper, we introduce HGNNv2, a stable hypergraph neural network, which is built as a hypergraph dynamic system with partial differential equation (PDE). This model incorporates a position-aware anisotropic diffusion term and an external control term. We further present the vertex-rooted subtree method to determine anisotropic diffusion intensity. HGNNv2 has properties that vertices occupying equivalent positions in the structural space share equivalent structural labels and positional features. Experiments on 6 hypergraph datasets and 3 graph datasets reveal that HGNNv2 outperforms all 12 compared methods. HGNNv2 is capable of achieving stable final representations and task accuracy even under noisy conditions. HGNNv2 achieves stable performance with fewer layers than hypergraph dynamic systems employing isotropic diffusion. We provide feature visualizations to illustrate the evolution of representations.
Yue Gao 0002, Jielong Yan, Yifan Feng 0001, Xiangmin Han, Shihui Ying, Zongze Wu 0001, Han Hu 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Hypergraph-Based High-Order Correlation Analysis for Large-Scale Long-Tailed Data Classification
abstract
High-order correlations, which capture complex interactions among multiple entities, extend beyond traditional graph representations and support a wider range of applications. However, existing neural network models for high-order correlations encounter scalability issues on large datasets due to the substantial computational complexity involved in processing large-scale structures. In addition, long-tailed distributions, which are common in real-world data, result in underrepresented categories and hinder the model's ability to learn effective high-order interaction patterns for rare instances. To address these issues, we introduce a novel framework known as HyperGraph-based High-order Correlation analysis (HGHC) for large-scale long-tailed data classification. Firstly, to tackle the long-tailed distribution problem, HGHC generates synthetic vertices and computes their attributed high-order correlations using an oversampling module inspired by SMOTE, termed HSMOTE, to enhance the representation of tail categories. Secondly, for efficient computational scaling, we treat the data as having two modalities: the structural modality capturing high-order relationships and the feature modality representing individual attributes. We perform computations on both CPU and GPU separately and then fuse the results to achieve a lightweight vertex transformation and aggregation scheme for high-order correlation data. Additionally, we contribute the first benchmark for large-scale long-tailed datasets involving high-order correlations, known as Amazon-LT, which includes multiple datasets with varying imbalance ratios. Our experimental results demonstrate that HGHC achieves state-of-the-art performance in handling high-order correlation analysis issues for large-scale, long-tailed data.
Xiangmin Han, Yubo Zhang 0006, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Graph Quality Matters on Revealing the Semantics Behind the Data in Physical World
Jielong Yan, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Reinterpreting Hypergraph Kernels: Insights Through Homomorphism Analysis
abstract
Designing expressive hypergraph kernels that can effectively capture high-order structural information is a fundamental challenge in hypergraph learning. In this paper, we propose a novel comparison framework based on hypergraph homomorphisms to evaluate and compare the expressive ability of existing hypergraph kernels. We revisit classical kernels such as Hypergraph Weisfeiler-Lehman (HG WL) and Hypergraph Rooted kernels, providing theoretical conditions under which they fail to distinguish non-isomorphic hypergraphs. Motivated by these insights, we introduce the Hypergraph Subtree-Cycle Kernel, which augments subtree-based features with cycle-based structural patterns to enhance expressiveness. We propose two variants: HG SCKernelv1 and HG SCKernelv2. Extensive experiments on five graph and ten hypergraph classification benchmarks demonstrate the superior performance of our methods, confirming the effectiveness of integrating homomorphism-guided design into hypergraph kernels.
Shaoyi Du, Yifan Feng 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Multi-task dynamic graph learning for brain disorder identification with functional MRI
Yunling Ma, Chaojun Zhang, Han Zhang 0002, Shihui Ying
Pattern Recognit.5
2026 Intrinsic Semiparametric Mixed Effects Model for Longitudinal Manifold-Valued Data
abstract
Abstract. This paper introduces a novel intrinsic semiparametric mixed-effects model (ISMEM) to characterize the complex relationships between manifold-valued responses and Euclidean-valued covariates in the longitudinal setting. Manifold-valued data, characterized by their inherent nonlinearity, high dimensionality, and specific geometric structures, present significant challenges for longitudinal analysis. Most existing intrinsic models are developed for cross-sectional data and unsuited for longitudinal studies. To address this limitation, we propose ISMEM, an efficient framework designed for longitudinal manifold-valued datasets that capture such relationships at both group and individual levels. Key features of ISMEM include (i) the integration of fixed and random effects on Riemannian manifolds, enabling analysis at multiple levels; (ii) a semiparametric structure combining parametric and nonparametric components, enhancing both flexibility and interpretability; and (iii) the preservation of the geometric properties of manifold-valued observations, offering improved robustness and interpretability compared to Euclidean-based models. In addition, we develop a two-stage iterative estimation procedure and validate our approach through simulations on the symmetric positive definite (SPD) manifold. Finally, we apply ISMEM to longitudinal data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI), reconstructing and comparing continuous three-dimensional (3D) shape trajectories of the lateral ventricle in Kendall shape space across distinct groups.
Junfei Huang, Xinjian Xu, Shihui Ying
SIAM J. Imaging Sci.4
2026 MSAFed: Generalized Multi-Stage and Adaptive Federated Learning for Test-Time Medical Segmentation
abstract
Federated learning (FL) enables collaborative model training across multiple medical centers without sharing data, offering significant promise for privacy-preserving AI in healthcare. However, FL models often lack generalization across all participating clients (inside FL) and perform poorly when deployed to unseen clients (outside FL), particularly in heterogeneous domains. Current test-time adaptation methods for outside FL fail to address biases in personalized models toward source distributions, limiting their clinical applications. To tackle these challenges, we propose MSAFed, a generalized multi-stage adaptive FL framework that enhances both inside generalization and outside test-time adaptation. During pretraining, intra-client and inter-client contrastive learning with prototype-aware aggregation produces a generalized global model. An adaptive learning rate strategy further improves inside FL generalization. For unseen clients, source knowledge, including adaptive learning rates and prototypes, is leveraged to dynamically adapt the network architecture during test time. Experiments on three real-world multi-center medical datasets demonstrate the effectiveness of MSAFed, achieving superior performance on both inside and outside FL tasks.
Jiajie Jin, Xuanmin Chen, Liyan Ma, Shihui Ying, Guang Yang 0006, Tieyong Zeng
IEEE J. Biomed. Health Informatics4
2026 Re-Visible Dual-Domain Self-Supervised Deep Unfolding Network for MRI Reconstruction
abstract
Magnetic Resonance Imaging (MRI) is widely used in clinical practice, but suffers from prolonged acquisition time. Although deep learning methods have been proposed to accelerate acquisition and demonstrate promising performance, they rely on high-quality fully-sampled datasets for training in a supervised manner. However, such datasets are time-consuming and expensive-to-collect, which constrains their broader applications. On the other hand, self-supervised methods offer an alternative by enabling learning from under-sampled data alone, but most existing methods rely on further partitioned under-sampled k-space data as model's input for training, which causes an input distribution shift between the the training stage and the inference stage. Additionally, their models have not effectively incorporated comprehensive image priors, leading to degraded reconstruction performance. In this paper, we propose a novel re-visible dual-domain self-supervised deep unfolding network to address these issues when only under-sampled datasets are available. Specifically, by incorporating re-visible dual-domain loss, all under-sampled k-space data are utilized during training to mitigate the input distribution shift caused by further partitioning. This design enables the model to implicitly adapt to all under-sampled k-space data as input. Additionally, we design a Deep Unfolding Network based on Chambolle and Pock Proximal Point Algorithm (DUN-CP-PPA) to achieve end-to-end reconstruction. By employing a Spatial-Frequency Feature Extraction (SFFE) block to capture both global and local representations, the model effectively integrates imaging physics with comprehensive image priors to enhance reconstruction performance. Experiments on both single-coil and multi-coil datasets demonstrate that our method outperforms state-of-the-art approaches in terms of reconstruction performance and generalization capability.
Hao Zhang 0026, Qi Wang 0128, Jian Sun 0009, Zhijie Wen, Jun Shi 0004, Shihui Ying
IEEE J. Biomed. Health Informatics6
2025 Beyond Graphs: Can Large Language Models Comprehend Hypergraphs?
abstract
Existing benchmarks like NLGraph and GraphQA evaluate LLMs on graphs by focusing mainly on pairwise relationships, overlooking the high-order correlations found in real-world data. Hypergraphs, which can model complex beyond-pairwise relationships, offer a more robust framework but are still underexplored in the context of LLMs. To address this gap, we introduce LLM4Hypergraph, the first comprehensive benchmark comprising 21,500 problems across eight low-order, five high-order, and two isomorphism tasks, utilizing both synthetic and real-world hypergraphs from citation networks and protein structures. We evaluate six prominent LLMs, including GPT-4o, demonstrating our benchmark’s effectiveness in identifying model strengths and weaknesses. Our specialized prompt- ing framework incorporates seven hypergraph languages and introduces two novel techniques, Hyper-BAG and Hyper-COT, which enhance high-order reasoning and achieve an average 4% (up to 9%) performance improvement on structure classification tasks. This work establishes a foundational testbed for integrating hypergraph computational capabilities into LLMs, advancing their comprehension.
Yifan Feng 0001, Chengwu Yang, Xingliang Hou, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002
ICLR5
2025 Mitigating noisy labels in long-tailed image classification via multi-level collaborative learning
Xinyang Zhou, Zhijie Wen, Yuandi Zhao, Jun Shi 0004, Shihui Ying
Appl. Intell.5
2025 HCNM: Hierarchical cognitive neural model for small-sample image classification
Dequan Jin, Ruoge Li, Xuanlu Xiang, Shihui Ying
Expert Syst. Appl.6
2025 Multi-resolution based dual-channel UNet with cross clique for medical image dense prediction
Xueying Zhou, Ge Jin 0002, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Expert Syst. Appl.6
2025 Open set label noise learning with robust sample selection and margin-guided module
Yuandi Zhao, Qianxi Xia, Zhijie Wen, Liyan Ma, Shihui Ying
Knowl. Based Syst.6
2025 HGMSurvNet: A two-stage hypergraph learning network for multimodal cancer survival prediction
Saisai Ding, Linjin Li, Ge Jin 0002, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Medical Image Anal.5
2025 LMS-Net: A learned Mumford-Shah network for binary few-shot medical image segmentation
Shengdong Zhang, Hao Zhang 0026, Jun Shi 0004, Liyan Ma, Shihui Ying
Medical Image Anal.7
2025 Deep unfolding network with spatial alignment for multi-modal MRI reconstruction
Hao Zhang 0026, Qi Wang 0128, Jun Shi 0004, Shihui Ying, Zhijie Wen
Medical Image Anal.4
2025 Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation
abstract
We introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12% and 9% improvements.
Yifan Feng 0001, Jiangang Huang, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Guiguang Ding, Rongrong Ji, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Self-Supervised Hypergraph Training Framework via Structure-Aware Learning
abstract
Hypergraphs, with their ability to model complex, beyond pair-wise correlations, presents a significant advancement over traditional graphs for capturing intricate relational data across diverse domains. However, the integration of hypergraphs into self-supervised learning (SSL) frameworks has been hindered by the intricate nature of high-order structural variations. This paper introduces the Self-Supervised Hypergraph Training Framework via Structure-Aware Learning (SS-HT), designed to enhance the perception and measurement of these variations within hypergraphs. The SS-HT framework employs a "Masking and Re-Masking" strategy to bolster feature reconstruction in Hypergraph Neural Networks (HGNNs), addressing the limitations of traditional SSL methods. It also introduces a metric strategy for local high-order correlation changes, streamlining the computational efficiency of structural distance calculations. Extensive experiments on 11 datasets demonstrate SS-HT's superior performance over existing SSL methods for both low-order and high-order data. Notably, the framework significantly reduces data labeling dependency, achieving a 32% improvement over HGNN in the downstream task fine-tuning phase under the 1% labeled data setting in the Cora-CC dataset. Ablation studies further validate SS-HT's scalability and its capacity to augment the performance of various HGNN methods, underscoring its robustness and applicability in real-world scenarios.
Yifan Feng 0001, Shiquan Liu, Shihui Ying, Shaoyi Du, Zongze Wu 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Kernelized Hypergraph Neural Networks
abstract
Hypergraph Neural Networks (HGNNs) have attracted much attention for high-order structural data learning. Existing methods mainly focus on simple mean-based aggregation or manually combining multiple aggregations to capture multiple information on hypergraphs. However, those methods inherently lack continuous non-linear modeling ability and are sensitive to varied distributions. Although some kernel-based aggregations on GNNs and CNNs can capture non-linear patterns to some degree, those methods are restricted in the low-order correlation and may cause unstable computation in training. In this work, we introduce Kernelized Hypergraph Neural Networks (KHGNN) and its variant, Half-Kernelized Hypergraph Neural Networks (H-KHGNN), which synergize mean-based and max-based aggregation functions to enhance representation learning on hypergraphs. KHGNN's kernelized aggregation strategy adaptively captures both semantic and structural information via learnable parameters, offering a mathematically grounded blend of kernelized aggregation approaches for comprehensive feature extraction. H-KHGNN addresses the challenge of overfitting in less intricate hypergraphs by employing non-linear aggregation selectively in the vertex-to-hyperedge message-passing process, thus reducing model complexity. Our theoretical contributions reveal a bounded gradient for kernelized aggregation, ensuring stability during training and inference. Empirical results demonstrate that KHGNN and H-KHGNN outperform state-of-the-art models across 10 graph/hypergraph datasets, with ablation studies demonstrating the effectiveness and computational stability of our method.
Yifan Feng 0001, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Multi-View Spectral Clustering on the Grassmannian Manifold With Hypergraph Representation
abstract
Graph-based multi-view spectral clustering methods have achieved notable progress recently, yet they often fall short in either oversimplifying pairwise relationships or struggling with inefficient spectral decompositions in high-dimensional Euclidean spaces. In this paper, we introduce a novel approach that begins to generate hypergraphs by leveraging sparse representation learning from data points. Based on the generated hypergraph, we propose an optimization function with orthogonality constraints for multi-view hypergraph spectral clustering, which incorporates spectral clustering for each view and ensures consistency across different views. In Euclidean space, solving the orthogonality-constrained optimization problem may yield local maxima and approximation errors. Innovately, we transform this problem into an unconstrained form on the Grassmannian manifold. Finally, we devise an alternating iterative Riemannian optimization algorithm to solve the problem. To validate the effectiveness of the proposed algorithm, we test it on four real-world multi-view datasets and compare its performance with six state-of-the-art multi-view clustering algorithms. The experimental results demonstrate that our method outperforms the baselines in terms of clustering performance due to its superior low-dimensional and resilient feature representation.
Murong Yang, Shihui Ying, Xin-Jian Xu, Yue Gao 0002
IEEE Trans. Big Data2
2025 A Trustworthy Curriculum Learning Guided Multi-Target Domain Adaptation Network for Autism Spectrum Disorder Classification
abstract
Domain adaptation has demonstrated success in classification of multi-center autism spectrum disorder (ASD). However, current domain adaptation methods primarily focus on classifying data in a single target domain with the assistance of one or multiple source domains, lacking the capability to address the clinical scenario of identifying ASD in multiple target domains. In response to this limitation, we propose a Trustworthy Curriculum Learning Guided Multi-Target Domain Adaptation (TCL-MTDA) network for identifying ASD in multiple target domains. To effectively handle varying degrees of data shift in multiple target domains, we propose a trustworthy curriculum learning procedure based on the Dempster-Shafer (D-S) Theory of Evidence. Additionally, a domain-contrastive adaptation method is integrated into the TCL-MTDA process to align data distributions between source and target domains, facilitating the learning of domain-invariant features. The proposed TCL-MTDA method is evaluated on 437 subjects (including 220 ASD patients and 217 NCs) from the Autism Brain Imaging Data Exchange (ABIDE). Experimental results validate the effectiveness of our proposed method in multi-target ASD classification, achieving an average accuracy of 71.46% (95% CI: 68.85% - 74.06%) across four target domains, significantly outperforming most baseline methods (p<0.05).
Jiale Dun, Jun Wang 0024, Juncheng Li 0003, Qianhui Yang, Wenlong Hang, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics7
2025 Bidirectional Projection-Based Multi-Modal Fusion Transformer for Early Detection of Cerebral Palsy in Infants
abstract
Periventricular white matter injury (PWMI) is the most frequent magnetic resonance imaging (MRI) finding in infants with Cerebral Palsy (CP). We aim to detect CP and identify subtle, sparse PWMI lesions in infants under two years of age with immature brain structures. Based on the characteristic that the responsible lesions are located within five target regions, we first construct a multi-modal dataset including 243 cases with the mask annotations of five target regions for delineating anatomical structures on T1-Weighted Imaging (T1WI) images, masks for lesions on T2-Weighted Imaging (T2WI) images, and categories (CP or Non-CP). Furthermore, we develop a bidirectional projection-based multi-modal fusion transformer (BiP-MFT), incorporating a Bidirectional Projection Fusion Module (BPFM) for integrating the features between five target regions on T1WI images and lesions on T2WI images. Our BiP-MFT achieves subject-level classification accuracy of 0.90, specificity of 0.87, and sensitivity of 0.94. It surpasses the best results of nine comparative methods, with 0.10, 0.08, and 0.09 improvements in classification accuracy, specificity and sensitivity respectively. Our BPFM outperforms eight compared feature fusion strategies using Transformer and U-Net backbones on our dataset. Ablation studies on the dataset annotations and model components justify the effectiveness of our annotation method and the model rationality. The proposed dataset and codes are available at https://github.com/Kai-Qi/BiP-MFT.
Kai Qi, Yizhe Yang, Shihui Ying, Jian Sun 0009
IEEE Trans. Medical Imaging5
2025 Exploring Local and Global Consistent Correlation on Hypergraph for Rotation Invariant Point Cloud Analysis
abstract
Rotation invariant point cloud analysis is essential for many real-world applications where objects can appear in arbitrary orientations. Traditional local rotation-invariant methods rely on lossy region descriptors, limiting the global comprehension of 3D objects. Conversely, global features derived from pose alignment can capture complementary information. To leverage both local and global consistency for enhanced accuracy, we propose the Global-Local-Consistent Hypergraph Cross-Attention Network (GLC-HCAN). This framework includes the Global Consistent Feature (GCF) representation branch, the Local Consistent Feature (LCF) representation branch, and the Hypergraph Cross-Attention (HyperCA) network to model complex correlations through the global-local-consistent hypergraph representation learning. Specifically, the GCF branch employs a multi-pose grouping and aggregation strategy based on PCA for improved global comprehension. Simultaneously, the LCF branch uses local farthest reference point features to enhance local region descriptions. To capture high-order and complex global-local correlations, we construct hypergraphs that integrate both features, mutually enhancing and fusing the representations. The inductive HyperCA module leverages attention techniques to better utilize these high-order relations for comprehensive understanding. Consequently, GLC-HCAN offers an effective and robust rotation-invariant point cloud analysis network, suitable for object classification and shape retrieval tasks in SO(3). Experimental results on both synthetic and scanned point cloud datasets demonstrate that GLC-HCAN outperforms state-of-the-art methods.
Yue Dai 0003, Shihui Ying, Yue Gao 0002
IEEE Trans. Multim.2
2025 Mode Hypergraph Neural Network
abstract
The hypergraph neural network (HGNN) is an emerging powerful tool for modeling and learning complex, high-order correlations among entities upon hypergraph structures. While existing HGNN-based approaches excel in modeling high-order correlations among data using hyperedges, they often have difficulties in distinguishing diverse semantics (e.g., bioactivities between drug and target in biological networks) of different correlations, making it challenging to learn accurate final representations. The underlying reason is that the specific semantic information of each hyperedge cannot be captured and distinguished during the modeling and learning process. To address this, we propose a mode HGNN ( $\textsf {MHGNN}$ ) framework that extends the vanilla hypergraph structure by endowing hyperedges with mode information for encapsulating their semantics and then performs mode-aware high-order message passing upon mode hypergraph for achieving comprehensive node representations. Extensive evaluations on four real-world datasets under two representative tasks have demonstrated the outstanding performance of $\textsf {MHGNN}$ against the state of the arts.
Shuyi Ji, Yifan Feng 0001, Donglin Di, Shihui Ying, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.4
2025 Multiview Representation Learning With One-to-Many Dynamic Relationships
abstract
Integrating information from multiple views to obtain potential representations with stronger expressive ability has received significant attention in practical applications. Most existing algorithms usually focus on learning either the consistent or complementary representation of views and, subsequently, integrate one-to-one corresponding sample representations between views. Although these approaches yield effective results, they do not fully exploit the information available from multiple views, limiting the potential for further performance improvement. In this article, we propose an unsupervised multiview representation learning method based on sample relationships, which enables the one-to-many fusion of intraview and interview information. Due to the heterogeneity of views, we need mainly face the two following challenges: 1) the discrepancy in the dimensions of data across different views and 2) the characterization and utilization of sample relationships across these views. To address these two issues, we adopt two modules: the dimension consistency relationship enhancement module and the multiview graph learning module. Thereinto, the relationship enhancement module addresses the discrepancy in data dimensions across different views and dynamically selects data dimensions for each sample that bolsters intraview relationships. The multiview graph learning module devises a novel multiview adjacency matrix to capture both intraview and interview sample relationships. To achieve one-to-many fusion and obtain multiview representations, we employ the graph autoencoder structure. Furthermore, we extend the proposed architecture to the supervised case. We conduct extensive experiments on various real-world multiview datasets, focusing on clustering and multilabel classification tasks, to evaluate the effectiveness of our method. The results demonstrate that our approach significantly improves performance compared to existing methods, highlighting the potential of leveraging sample relationships for multiview representation learning. Our code is released at https://github.com/lilidan-orm/one-to-many-multiview on GitHub.
Haibao Wang, Shihui Ying
IEEE Trans. Neural Networks Learn. Syst.3
2024 LightHGNN: Distilling Hypergraph Neural Networks into MLPs for 100x Faster Inference
abstract
Hypergraph Neural Networks (HGNNs) have recently attracted much attention and exhibited satisfactory performance due to their superiority in high-order correlation modeling. However, it is noticed that the high-order modeling capability of hypergraph also brings increased computation complexity, which hinders its practical industrial deployment. In practice, we find that one key barrier to the efficient deployment of HGNNs is the high-order structural dependencies during inference. In this paper, we propose to bridge the gap between the HGNNs and inference-efficient Multi-Layer Perceptron (MLPs) to eliminate the hypergraph dependency of HGNNs and thus reduce computational complexity as well as improve inference speed. Specifically, we introduce LightHGNN and LightHGNN$^+$ for fast inference with low complexity. LightHGNN directly distills the knowledge from teacher HGNNs to student MLPs via soft labels, and LightHGNN$^+$ further explicitly injects reliable high-order correlations into the student MLPs to achieve topology-aware distillation and resistance to over-smoothing. Experiments on eight hypergraph datasets demonstrate that even without hypergraph dependency, the proposed LightHGNNs can still achieve competitive or even better performance than HGNNs and outperform vanilla MLPs by $16.3$ on average. Extensive experiments on three graph datasets further show the average best performance of our LightHGNNs compared with all other methods. Experiments on synthetic hypergraphs with 5.5w vertices indicate LightHGNNs can run $100\times$ faster than HGNNs, showcasing their ability for latency-sensitive deployments.
Yifan Feng 0001, Yihe Luo, Shihui Ying, Yue Gao 0002
ICLR3
2024 Hypergraph Dynamic System
abstract
Recently, hypergraph neural networks (HGNNs) exhibit the potential to tackle tasks with high-order correlations and have achieved success in many tasks. However, existing evolution on the hypergraph has poor controllability and lacks sufficient theoretical support (like dynamic systems), thus yielding sub-optimal performance. One typical scenario is that only one or two layers of HGNNs can achieve good results and more layers lead to degeneration of performance. Under such circumstances, it is important to increase the controllability of HGNNs. In this paper, we first introduce hypergraph dynamic systems (HDS), which bridge hypergraphs and dynamic systems and characterize the continuous dynamics of representations. We then propose a control-diffusion hypergraph dynamic system by an ordinary differential equation (ODE). We design a multi-layer HDS$^{ode}$ as a neural implementation, which contains control steps and diffusion steps. HDS$^{ode}$ has the properties of controllability and stabilization and is allowed to capture long-range correlations among vertices. Experiments on $9$ datasets demonstrate HDS$^{ode}$ beat all compared methods. HDS$^{ode}$ achieves stable performance with increased layers and solves the poor controllability of HGNNs. We also provide the feature visualization of the evolutionary process to demonstrate the controllability and stabilization of HDS$^{ode}$.
Jielong Yan, Yifan Feng 0001, Shihui Ying, Yue Gao 0002
ICLR3
2024 Cross-Evaluation and Re-weighting for Multi-Source-Free Domain Adaptation
abstract
In this paper, we investigate the multi-source-free domain adaptation (MSFDA), a specific case of unsupervised domain adaptation where multiple source models are adapted to the target domain without accessing the source datasets. We propose sample cross-evaluation and re-weighting for MSFDA. Specifically, we use a mixture model to model the loss distribution of each sample, dynamically obtaining pseudo-labels and their probabilities of being clean samples. In this process, we use each source model to perform sample cross-evaluation for the other source models, which facilitates source models to guide each other and avoids confirmation bias. We then utilize the clean probabilities for re-weighting to make clean samples to dominate the training, progressively adapting the models to the target domain. The experiments on four benchmark datasets demonstrate the superiority of our proposed method.
Ying Li 0028, Shihui Ying
ICME3
2024 Multi-View disentanglement-based bidirectional generalized distillation for diagnosis of liver cancers with ultrasound images
Lehang Guo, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Inf. Process. Manag.5
2024 Hypergraph Isomorphism Computation
abstract
The isomorphism problem is a fundamental problem in network analysis, which involves capturing both low-order and high-order structural information. In terms of extracting low-order structural information, graph isomorphism algorithms analyze the structural equivalence to reduce the solver space dimension, which demonstrates its power in many applications, such as protein design, chemical pathways, and community detection. For the more commonly occurring high-order relationships in real-life scenarios, the problem of hypergraph isomorphism, which effectively captures these high-order structural relationships, cannot be straightforwardly addressed using graph isomorphism methods. Besides, the existing hypergraph kernel methods may suffer from high memory consumption or inaccurate sub-structure identification, thus yielding sub-optimal performance. In this paper, to address the abovementioned problems, we first propose the hypergraph Weisfiler-Lehman test algorithm for the hypergraph isomorphism test problem by generalizing the Weisfiler-Lehman test algorithm from graphs to hypergraphs. Secondly, based on the presented algorithm, we propose a general hypergraph Weisfieler-Lehman kernel framework and implement two instances, which are Hypergraph Weisfeiler-Lehamn Subtree Kernel (Hypergraph WL Subtree Kernel) and Hypergraph Weisfeiler-Lehamn Hyperedge Kernel (Hypergraph WL Hyperedge Kernel). The Hypergraph WL Subtree Kernel counts different types of rooted subtrees and generates the final feature vector for a given hypergraph by comparing the number of different types of rooted subtrees. The Hypergraph WL Hyperedge Kernel is developed to process hypergraphs with more degrees of hyperedges, which counts the vertex labels that are connected by each hyperedge to generate the feature vector. Mathematically, we prove the proposed Hypergraph WL Subtree Kernel can degenerate into the typical Graph Weisfeiler-Lehman Subtree Kernel when dealing with low-order graph structures. In order to fulfill our research objectives, a comprehensive set of experiments was meticulously designed, including seven graph classification datasets and 12 hypergraph classification datasets. Results on graph classification datasets indicate that the Hypergraph WL Subtree Kernel can achieve the same performance compared with the classical Graph Weisfeiler-Lehman Subtree Kernel. Results on hypergraph classification datasets show significant improvements compared to other typical kernel-based methods, which demonstrates the effectiveness of the proposed methods. In our evaluation, we found that our proposed methods outperform the second-best method in terms of runtime, running over 80 times faster when handling complex hypergraph structures. This significant speed advantage highlights the great potential of our methods in real-world applications.
Yifan Feng 0001, Jiashu Han, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Self-adaptive subspace representation from a geometric intuition
Lipeng Cai, Jun Shi 0004, Shaoyi Du, Yue Gao 0002, Shihui Ying
Pattern Recognit.5
2024 Sliding at First-Order: Higher-Order Momentum Distributions for Discontinuous Image Registration
abstract
Abstract. In this paper, we propose a new approach to deformable image registration that captures sliding motions. The large deformation diffeomorphic metric mapping (LDDMM) registration method faces challenges in representing sliding motion since it per construction generates smooth warps. To address this issue, we extend LDDMM by incorporating both zeroth- and first-order momenta with a nondifferentiable kernel. This allows us to represent both discontinuous deformation at switching boundaries and diffeomorphic deformation in homogeneous regions. We provide a mathematical analysis of the proposed deformation model from the viewpoint of discontinuous systems. To evaluate our approach, we conduct experiments on both artificial images and the publicly available DIR-Lab 4DCT dataset. Results show the effectiveness of our approach in capturing plausible sliding motion.
Lili Bao, Shihui Ying, Stefan Sommer
SIAM J. Imaging Sci.3
2024 Strong Consistency of Spectral Clustering for the Sparse Degree-Corrected Hypergraph Stochastic Block Model
abstract
We prove strong consistency of spectral clustering under the degree-corrected hypergraph stochastic block model in the sparse regime where the maximum expected hyperdegree is as small as$\Omega (\log n)$with$n$denoting the number of nodes. We show that the basic spectral clustering without preprocessing or postprocessing is strongly consistent in an even wider range of the model parameters, in contrast to previous studies that either trim high-degree nodes or perform local refinement. At the heart of our analysis is the entry-wise eigenvector perturbation bound derived by the “leave-one-out”technique. To the best of our knowledge, this is the first entry-wise error bound for degree-corrected hypergraph models, resulting in the strong consistency for clustering non-uniform hypergraphs with heterogeneous hyperdegrees.
Chong Deng, Xin-Jian Xu, Shihui Ying
IEEE Trans. Inf. Theory3
2024 FEFA: Frequency Enhanced Multi-Modal MRI Reconstruction With Deep Feature Alignment
abstract
Integrating complementary information from multiple magnetic resonance imaging (MRI) modalities is often necessary to make accurate and reliable diagnostic decisions. However, the different acquisition speeds of these modalities mean that obtaining information can be time consuming and require significant effort. Reference-based MRI reconstruction aims to accelerate slower, under-sampled imaging modalities, such as T2-modality, by utilizing redundant information from faster, fully sampled modalities, such as T1-modality. Unfortunately, spatial misalignment between different modalities often negatively impacts the final results. To address this issue, we propose FEFA, which consists of cascading FEFA blocks. The FEFA block first aligns and fuses the two modalities at the feature level. The combined features are then filtered in the frequency domain to enhance the important features while simultaneously suppressing the less essential ones, thereby ensuring accurate reconstruction. Furthermore, we emphasize the advantages of combining the reconstruction results from multiple cascaded blocks, which also contributes to stabilizing the training process. Compared to existing registration-then-reconstruction and cross-attention-based approaches, our method is end-to-end trainable without requiring additional supervision, extensive parameters, or heavy computation. Experiments on the public fastMRI, IXI and in-house datasets demonstrate that our approach is effective across various under-sampling patterns and ratios.
Xuanmin Chen, Liyan Ma, Shihui Ying, Dinggang Shen, Tieyong Zeng
IEEE J. Biomed. Health Informatics3
2024 Correction to: Multi-View Feature Transformation Based SVM+ for Computer-Aided Diagnosis of Liver Cancers With Ultrasound Image
abstract
Presents corrections to the paper, Multi-View Feature Transformation Based SVM+ for Computer-Aided Diagnosis of Liver Cancers With Ultrasound Image.
Lehang Guo, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2024 OTCLDA: Optimal Transport and Contrastive Learning for Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptive (UDA) semantic segmentation aims to assign a predetermined semantic label to every single pixel of the unannotated target data by exploiting a model that is trained on the labeled source data. Numerous current methods only display concern for grouping similar features together but ignore dispersing those features across various classes, so that some feature representations can not be well-separated. Therefore, we propose to employ contrastive learning (CL) method to increase the similarity of pixel features, propelling similar features closer and dispelling different ones far away. Furthermore, due to the domain shift, the UDA model frequently has poor generalization on the target domain. Accordingly, we design an optimal transport (OT) module to enhance UDA by comparing and aligning sample distributions to minimize transport loss between them. By taking advantage of this, the domain shift can be efficaciously mitigated by bringing the target probability distribution closer to that of the source. Specially, due to its simplicity, our OT module can be integrated into various UDA methods. In light of the aforementioned viewpoints, we put forth an ingenious approach, named OTCLDA, which successfully combines OT and CL while enhancing the performance of the UDA model. Multitudinous experiments demonstrate the importance of our method involving OT and CL. It significantly gains mIoU of 75.1% on benchmark GTA$\rightarrow$Cityscapes, and 66.9% on SYNTHIA$\rightarrow$Cityscapes respectively, displaying a competitive performance compared with previous works. The source code of OTCLDA is publicly available at https://github.com/YYDSDD/OTCLDA.
Qizhe Fan, Xiaoqin Shen, Shihui Ying, Shaoyi Du
IEEE Trans. Intell. Transp. Syst.3
2024 Penalized Flow Hypergraph Local Clustering
abstract
In recent years, hypergraph analysis have attracted increasing attention due to their ability to model complex data correlation, with hypergraph clustering being one of the most important tasks. However, when the scale of hypergraph is large enough, clustering is difficult based on global consistency. Existing flow-based hypergraph local clustering methods have good theoretical cut improvements and runtime guarantees. However, these methods exhibit poor performance when the initial reference node set is small and are prone to causing the output set to shrink into a small subset, resulting in local minima. To address this issue, we propose the Penalized Flow Hypergraph Local Clustering(PFHLC) and provide new conductance guarantees and runtime analyses for our method. First, we use the random walk method to grow the initial seed set, and introduce the random walk information of nodes as penalized flow into the flow-based framework to optimize the output. Second, we propose a generalized objective function containing random walk information, which takes full advantage of the semi-supervised information of the target cluster to protect important nodes. This feature can avoid the local minima of previous flow-based methods. Importantly, our method is strongly-local and can run efficiently on large-scale hypergraphs. We contribute a real-world dataset and the experiments on real-world large-scale datasets show that PFHLC achieves the state-of-the-art significantly.
Yubo Zhang 0006, Chenggang Yan 0001, Zuxing Xuan, Ting Yu 0004, Ji Zhang 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.7
2024 Multimodal Co-Attention Fusion Network With Online Data Augmentation for Cancer Subtype Classification
abstract
It is an essential task to accurately diagnose cancer subtypes in computational pathology for personalized cancer treatment. Recent studies have indicated that the combination of multimodal data, such as whole slide images (WSIs) and multi-omics data, could achieve more accurate diagnosis. However, robust cancer diagnosis remains challenging due to the heterogeneity among multimodal data, as well as the performance degradation caused by insufficient multimodal patient data. In this work, we propose a novel multimodal co-attention fusion network (MCFN) with online data augmentation (ODA) for cancer subtype classification. Specifically, a multimodal mutual-guided co-attention (MMC) module is proposed to effectively perform dense multimodal interactions. It enables multimodal data to mutually guide and calibrate each other during the integration process to alleviate inter- and intra-modal heterogeneities. Subsequently, a self-normalizing network (SNN)-Mixer is developed to allow information communication among different omics data and alleviate the high-dimensional small-sample size problem in multi-omics data. Most importantly, to compensate for insufficient multimodal samples for model training, we propose an ODA module in MCFN. The ODA module leverages the multimodal knowledge to guide the data augmentations of WSIs and maximize the data diversity during model training. Extensive experiments are conducted on the public TCGA dataset. The experimental results demonstrate that the proposed MCFN outperforms all the compared algorithms, suggesting its effectiveness.
Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE Trans. Medical Imaging4
2024 Weakly Supervised Lesion Detection and Diagnosis for Breast Cancers With Partially Annotated Ultrasound Images
abstract
Deep learning (DL) has proven highly effective for ultrasound-based computer-aided diagnosis (CAD) of breast cancers. In an automatic CAD system, lesion detection is critical for the following diagnosis. However, existing DL-based methods generally require voluminous manually-annotated region of interest (ROI) labels and class labels to train both the lesion detection and diagnosis models. In clinical practice, the ROI labels, i.e. ground truths, may not always be optimal for the classification task due to individual experience of sonologists, resulting in the issue of coarse annotation to limit the diagnosis performance of a CAD model. To address this issue, a novel Two-Stage Detection and Diagnosis Network (TSDDNet) is proposed based on weakly supervised learning to improve diagnostic accuracy of the ultrasound-based CAD for breast cancers. In particular, all the initial ROI-level labels are considered as coarse annotations before model training. In the first training stage, a candidate selection mechanism is then designed to refine manual ROIs in the fully annotated images and generate accurate pseudo-ROIs for the partially annotated images under the guidance of class labels. The training set is updated with more accurate ROI labels for the second training stage. A fusion network is developed to integrate detection network and classification network into a unified end-to-end framework as the final CAD model in the second training stage. A self-distillation strategy is designed on this model for joint optimization to further improves its diagnosis performance. The proposed TSDDNet is evaluated on three B-mode ultrasound datasets, and the experimental results indicate that it achieves the best performance on both lesion detection and diagnosis tasks, suggesting promising application potential.
Jian Wang 0135, Shichong Zhou, Jun Wang 0024, Juncheng Li 0003, Shihui Ying, Cai Chang, Jun Shi 0004
IEEE Trans. Medical Imaging7
2024 Spatial and Modal Optimal Transport for Fast Cross-Modal MRI Reconstruction
abstract
Multi-modal magnetic resonance imaging (MRI) plays a crucial role in comprehensive disease diagnosis in clinical medicine. However, acquiring certain modalities, such as T2-weighted images (T2WIs), is time-consuming and prone to be with motion artifacts. It negatively impacts subsequent multi-modal image analysis. To address this issue, we propose an end-to-end deep learning framework that utilizes T1-weighted images (T1WIs) as auxiliary modalities to expedite T2WIs' acquisitions. While image pre-processing is capable of mitigating misalignment, improper parameter selection leads to adverse pre-processing effects, requiring iterative experimentation and adjustment. To overcome this shortage, we employ Optimal Transport (OT) to synthesize T2WIs by aligning T1WIs and performing cross-modal synthesis, effectively mitigating spatial misalignment effects. Furthermore, we adopt an alternating iteration framework between the reconstruction task and the cross-modal synthesis task to optimize the final results. Then, we prove that the reconstructed T2WIs and the synthetic T2WIs become closer on the T2 image manifold with iterations increasing, and further illustrate that the improved reconstruction result enhances the synthesis process, whereas the enhanced synthesis result improves the reconstruction process. Finally, experimental results from FastMRI and internal datasets confirm the effectiveness of our method, demonstrating significant improvements in image reconstruction quality even at low sampling rates.
Qi Wang 0128, Zhijie Wen, Jun Shi 0004, Qian Wang 0001, Dinggang Shen, Shihui Ying
IEEE Trans. Medical Imaging6
2024 Histopathology Image Classification With Noisy Labels via The Ranking Margins
abstract
Clinically, histopathology images always offer a golden standard for disease diagnosis. With the development of artificial intelligence, digital histopathology significantly improves the efficiency of diagnosis. Nevertheless, noisy labels are inevitable in histopathology images, which lead to poor algorithm efficiency. Curriculum learning is one of the typical methods to solve such problems. However, existing curriculum learning methods either fail to measure the training priority between difficult samples and noisy ones or need an extra clean dataset to establish a valid curriculum scheme. Therefore, a new curriculum learning paradigm is designed based on a proposed ranking function, which is named The Ranking Margins (TRM). The ranking function measures the 'distances' between samples and decision boundaries, which helps distinguish difficult samples and noisy ones. The proposed method includes three stages: the warm-up stage, the main training stage and the fine-tuning stage. In the warm-up stage, the margin of each sample is obtained through the ranking function. In the main training stage, samples are progressively fed into the networks for training, starting from those with larger margins to those with smaller ones. Label correction is also performed in this stage. In the fine-tuning stage, the networks are retrained on the samples with corrected labels. In addition, we provide theoretical analysis to guarantee the feasibility of TRM. The experiments on two representative histopathologies image datasets show that the proposed method achieves substantial improvements over the latest Label Noise Learning (LNL) methods.
Zhijie Wen, Haixia Wu, Shihui Ying
IEEE Trans. Medical Imaging3
2024 Pseudo-Data Based Self-Supervised Federated Learning for Classification of Histopathological Images
abstract
Computer-aided diagnosis (CAD) can help pathologists improve diagnostic accuracy together with consistency and repeatability for cancers. However, the CAD models trained with the histopathological images only from a single center (hospital) generally suffer from the generalization problem due to the straining inconsistencies among different centers. In this work, we propose a pseudo-data based self-supervised federated learning (FL) framework, named SSL-FT-BT, to improve both the diagnostic accuracy and generalization of CAD models. Specifically, the pseudo histopathological images are generated from each center, which contain both inherent and specific properties corresponding to the real images in this center, but do not include the privacy information. These pseudo images are then shared in the central server for self-supervised learning (SSL) to pre-train the backbone of global mode. A multi-task SSL is then designed to effectively learn both the center-specific information and common inherent representation according to the data characteristics. Moreover, a novel Barlow Twins based FL (FL-BT) algorithm is proposed to improve the local training for the CAD models in each center by conducting model contrastive learning, which benefits the optimization of the global model in the FL procedure. The experimental results on four public histopathological image datasets indicate the effectiveness of the proposed SSL-FL-BT on both diagnostic accuracy and generalization.
Xiangmin Han, Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE Trans. Medical Imaging7
2023 PGF-BIQA: Blind image quality assessment via probability multi-grained cascade forest
Hao Liu 0060, Ce Li 0001, Shangang Jin, Weizhe Gao, Fenghua Liu, Shaoyi Du, Shihui Ying
Comput. Vis. Image Underst.7
2023 GAME: GAussian Mixture Error-based meta-learning architecture
Jinhe Dong, Jun Shi 0004, Yue Gao 0002, Shihui Ying
Neural Comput. Appl.4
2023 JSMix: a holistic algorithm for learning with label noise
Zhijie Wen, Shihui Ying
Neural Comput. Appl.3
2023 B-mode ultrasound based CAD for liver cancers via multi-view privileged information learning
Xiangmin Han, Bangming Gong, Lehang Guo, Jun Wang 0024, Shihui Ying, Shuo Li 0001, Jun Shi 0004
Neural Networks5
2023 Attention-Guided Optimal Transport for Unsupervised Domain Adaptation with Class Structure Prior
Ying Li 0028, Shihui Ying
Neural Process. Lett.3
2023 Structure Evolution on Manifold for Graph Learning
abstract
Graph has been widely used in various applications, while how to optimize the graph is still an open question. In this paper, we propose a framework to optimize the graph structure via structure evolution on graph manifold. We first define the graph manifold and search the best graph structure on this manifold. Concretely, associated with the data features and the prediction results of a given task, we define a graph energy to measure how the graph fits the graph manifold from an initial graph structure. The graph structure then evolves by minimizing the graph energy. In this process, the graph structure can be evolved on the graph manifold corresponding to the update of the prediction results. Alternatively iterating these two processes, both the graph structure and the prediction results can be updated until converge. It achieves the suitable structure for graph learning without searching all hyperparameters. To evaluate the performance of the proposed method, we have conducted experiments on eight datasets and compared with the recent state-of-the-art methods. Experiment results demonstrate that our method outperforms the state-of-the-art methods in both transductive and inductive settings.
Hai Wan, Xinwei Zhang 0012, Yubo Zhang 0006, Xibin Zhao, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Graph Learning on Millions of Data in Seconds: Label Propagation Acceleration on Graph Using Data Distribution
abstract
Graph-based semi-supervised learning methods have been used in a wide range of real-world applications, e.g., from social relationship mining to multimedia classification and retrieval. However, existing methods are limited along with high computational complexity or not facilitating incremental learning, which may not be powerful to deal with large-scale data, whose scale may continuously increase, in real world. This paper proposes a new method called Data Distribution Based Graph Learning (DDGL) for semi-supervised learning on large-scale data. This method can achieve a fast and effective label propagation and supports incremental learning. The key motivation is to propagate the labels along smaller-scale data distribution model parameters, rather than directly dealing with the raw data as previous methods, which accelerate the data propagation significantly. It also improves the prediction accuracy since the loss of structure information can be alleviated in this way. To enable incremental learning, we propose an adaptive graph updating strategy which can update the model when there is distribution bias between new data and the already seen data. We have conducted comprehensive experiments on multiple datasets with sample sizes increasing from seven thousand to five million. Experimental results on the classification task on large-scale data demonstrate that our proposed DDGL method improves the classification accuracy by a large margin while consuming much less time compared to state-of-the-art methods.
Yubo Zhang 0006, Shuyi Ji, Changqing Zou, Xibin Zhao, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 ML-DSVM+: A meta-learning based deep SVM+ for computer-aided diagnosis
Xiangmin Han, Jun Wang 0024, Shihui Ying, Jun Shi 0004, Dinggang Shen
Pattern Recognit.3
2023 Multi-Scale Efficient Graph-Transformer for Whole Slide Image Classification
abstract
The multi-scale information among the whole slide images (WSIs) is essential for cancer diagnosis. Although the existing multi-scale vision Transformer has shown its effectiveness for learning multi-scale image representation, it still cannot work well on the gigapixel WSIs due to their extremely large image sizes. To this end, we propose a novel Multi-scale Efficient Graph-Transformer (MEGT) framework for WSI classification. The key idea of MEGT is to adopt two independent efficient Graph-based Transformer (EGT) branches to process the low-resolution and high-resolution patch embeddings (i.e., tokens in a Transformer) of WSIs, respectively, and then fuse these tokens via a multi-scale feature fusion module (MFFM). Specifically, we design an EGT to efficiently learn the local-global information of patch tokens, which integrates the graph representation into Transformer to capture spatial-related information of WSIs. Meanwhile, we propose a novel MFFM to alleviate the semantic gap among different resolution patches during feature fusion, which creates a non-patch token for each branch as an agent to exchange information with another branch by cross-attention mechanism. In addition, to expedite network training, a new token pruning module is developed in EGT to reduce the redundant tokens. Extensive experiments on both TCGA-RCC and CAMELYON16 datasets demonstrate the effectiveness of the proposed MEGT.
Saisai Ding, Juncheng Li 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2023 Reconstruction of Quantitative Susceptibility Mapping From Total Field Maps With Local Field Maps Guided UU-Net
abstract
Quantitative susceptibility mapping (QSM) is an emerging computational technique based on the magnetic resonance imaging (MRI) phase signal, which can provide magnetic susceptibility values of tissues. The existing deep learning-based models mainly reconstruct QSM from local field maps. However, the complicated inconsecutive reconstruction steps not only accumulate errors for inaccurate estimation, but also are inefficient in clinical practice. To this end, a novel local field maps guided UU-Net with Self- and Cross-Guided Transformer (LGUU-SCT-Net) is proposed to reconstruct QSM directly from the total field maps. Specifically, we propose to additionally generate the local field maps as the auxiliary supervision during the training stage. This strategy decomposes the more complicated mapping from total maps to QSM into two relatively easier ones, effectively alleviating the difficulty of direct mapping. Meanwhile, an improved U-Net model, named LGUU-SCT-Net, is further designed to promote the nonlinear mapping ability. The long-range connections are designed between two sequentially stacked U-Nets to bring more feature fusions and facilitate the information flow. The Self- and Cross-Guided Transformer integrated into these connections further captures multi-scale channel-wise correlations and guides the fusion of multi-scale transferred features, assisting in the more accurate reconstruction. The experimental results on an in-vivo dataset demonstrate the superior reconstruction results of our proposed algorithm.
Shihui Ying, Jun Wang 0024, Hongjian He, Jun Shi 0004
IEEE J. Biomed. Health Informatics2
2023 Two-Stage Self-Supervised Cycle-Consistency Transformer Network for Reducing Slice Gap in MR Images
abstract
Magnetic resonance (MR) images are usually acquired with large slice gap in clinical practice, i.e., low resolution (LR) along the through-plane direction. It is feasible to reduce the slice gap and reconstruct high-resolution (HR) images with the deep learning (DL) methods. To this end, the paired LR and HR images are generally required to train a DL model in a popular fully supervised manner. However, since the HR images are hardly acquired in clinical routine, it is difficult to get sufficient paired samples to train a robust model. Moreover, the widely used convolutional Neural Network (CNN) still cannot capture long-range image dependencies to combine useful information of similar contents, which are often spatially far away from each other across neighboring slices. To this end, a Two-stage Self-supervised Cycle-consistency Transformer Network (TSCTNet) is proposed to reduce the slice gap for MR images in this work. A novel self-supervised learning (SSL) strategy is designed with two stages respectively for robust network pre-training and specialized network refinement based on a cycle-consistency constraint. A hybrid Transformer and CNN structure is utilized to build an interpolation model, which explores both local and global slice representations. The experimental results on two public MR image datasets indicate that TSCTNet achieves superior performance over other compared SSL-based algorithms.
Zhiyang Lu, Jian Wang 0135, Shihui Ying, Jun Wang 0024, Jun Shi 0004, Dinggang Shen
IEEE J. Biomed. Health Informatics4
2023 Multi-View Feature Transformation Based SVM+ for Computer-Aided Diagnosis of Liver Cancers With Ultrasound Images
abstract
It is feasible to improve the performance of B-mode ultrasound (BUS) based computer-aided diagnosis (CAD) for liver cancers by transferring knowledge from contrast-enhanced ultrasound (CEUS) images. In this work, we propose a novel feature transformation based support vector machine plus (SVM+) algorithm for this transfer learning task by introducing feature transformation into the SVM+ framework (named FSVM+). Specifically, the transformation matrix in FSVM+ is learned to minimize the radius of the enclosing ball of all samples, while the SVM+ is used to maximize the margin between two classes. Moreover, to capture more transferable information from multiple CEUS phase images, a multi-view FSVM+ (MFSVM+) is further developed, which transfers knowledge from three CEUS images from three phases, i.e., arterial phase, portal venous phase, and delayed phase, to the BUS-based CAD model. MFSVM+ innovatively assigns appropriate weights for each CEUS image by calculating the maximum mean discrepancy between a pair of BUS and CEUS images, which can capture the relationship between source and target domains. The experimental results on a bi-modal ultrasound liver cancer dataset demonstrate that MFSVM+ achieves the best classification accuracy of 88.24±1.28%, sensitivity of 88.32±2.88%, specificity of 88.17±2.91%, suggesting its effectiveness in promoting the diagnostic accuracy of BUS-based CAD.
Lehang Guo, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2023 A Weighting Method for Feature Dimension by Semisupervised Learning With Entropy
abstract
In this article, a semisupervised weighting method for feature dimension based on entropy is proposed for classification, dimension reduction, and correlation analysis. For real-world data, different feature dimensions usually show different importance. Generally, data in the same class are supposed to be similar, so their entropy should be small; and those in different classes are supposed to be dissimilar, so their entropy should be large. According to this, we propose a way to construct the weights of feature dimensions with the whole entropy and the innerclass entropies. The weights indicate the contribution of their corresponding feature dimensions in classification. They can be used to improve the performance of classification by giving a weighted distance metric and can be applied to dimension reduction and correlation analysis as well. Some numerical experiments are given to test the proposed method by comparing it with some other representative methods. They demonstrate that the proposed method is feasible and efficient in classification, dimension reduction, and correlation analysis.
Dequan Jin, Murong Yang, Ziyan Qin, Shihui Ying
IEEE Trans. Neural Networks Learn. Syst.5
2022 Task-Driven Self-Supervised BI-Channel Networks Learning for Diagnosis of Breast Cancers with Mammography
abstract
Deep learning can promote mammography-based computer-aided diagnosis (CAD) for breast cancers, but it generally suffers from the small size sample problem. In this work, a task-driven self-supervised bi-channel networks learning (TSBNL) framework is proposed to improve the network performance with limited mammograms. In particular, a new gray-scale image mapping (GSIM) task for image restoration is designed as the pretext task to improve discriminative feature representation with label information of mammograms. TSBNL then innovatively integrates this image restoration network and the downstream classification network into a unified SSL framework, and transfers the knowledge from the pretext network to the classification network with improved diagnostic accuracy. The proposed algorithm is evaluated on a public INbreast mammogram dataset. The experimental results indicate that it outperforms the conventional SSL algorithms for the diagnosis of breast cancers with limited samples.
Ronglin Gong, Shihui Ying, Jun Shi 0004
ICIP2
2022 Intrinsic partial linear models for manifold-valued data
Shihui Ying, Hongtu Zhu
Inf. Process. Manag.2
2022 Multi-Class ASD Classification via Label Distribution Learning with Class-Shared and Class-Specific Decomposition
Jun Wang 0024, Fengyexin Zhang, Xiuyi Jia, Xin Wang 0084, Han Zhang 0002, Shihui Ying, Qian Wang 0001, Jun Shi 0004, Dinggang Shen
Medical Image Anal.6
2022 Multitask transfer learning with kernel representation
Shihui Ying, Zhijie Wen
Neural Comput. Appl.2
2022 Long Time Series Deep Forecasting with Multiscale Feature Extraction and Seq2seq Attention Mechanism
Xin Wang 0084, Yixian Luo, Zhijie Wen, Shihui Ying
Neural Process. Lett.5
2022 Kullback-Leibler Divergence Metric Learning
abstract
The Kullback-Leibler divergence (KLD), which is widely used to measure the similarity between two distributions, plays an important role in many applications. In this article, we address the KLD metric-learning task, which aims at learning the best KLD-type metric from the distributions of datasets. Concretely, first, we extend the conventional KLD by introducing a linear mapping and obtain the best KLD to well express the similarity of data distributions by optimizing such a linear mapping. It improves the expressivity of data distribution, which means it makes the distributions in the same class close and those in different classes far away. Then, the KLD metric learning is modeled by a minimization problem on the manifold of all positive-definite matrices. To deal with this optimization task, we develop an intrinsic steepest descent method, which preserves the manifold structure of the metric in the iteration. Finally, we apply the proposed method along with ten popular metric-learning approaches on the tasks of 3-D object classification and document classification. The experimental results illustrate that our proposed method outperforms all other methods.
Shuyi Ji, Zizhao Zhang 0003, Shihui Ying, Xibin Zhao, Yue Gao 0002
IEEE Trans. Cybern.3
2022 Rotation-Invariant Point Cloud Representation for 3-D Model Recognition
abstract
Three-dimensional (3-D) data have many applications in the field of computer vision and a point cloud is one of the most popular modalities. Therefore, how to establish a good representation for a point cloud is a core issue in computer vision, especially for 3-D object recognition tasks. Existing approaches mainly focus on the invariance of representation under the group of permutations. However, for point cloud data, it should also be rotation invariant. To address such invariance, in this article, we introduce a relation of equivalence under the action of rotation group, through which the representation of point cloud is located in a homogeneous space. That is, two point clouds are regarded as equivalent when they are only different from a rotation. Our network is flexibly incorporated into existing frameworks for point clouds, which guarantees the proposed approach to be rotation invariant. Besides, a sufficient analysis on how to parameterize the group SO(3) into a convolutional network, which captures a relation with all rotations in 3-D Euclidean space [Formula: see text]. We select the optimal rotation as the best representation of point cloud and propose a solution for minimizing the problem on the rotation group SO(3) by using its geometric structure. To validate the rotation invariance, we combine it with two existing deep models and evaluate them on ModelNet40 dataset and its subset ModelNet10. Experimental results indicate that the proposed strategy improves the performance of those existing deep models when the data involve arbitrary rotations.
Yan Wang 0076, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Cybern.3
2022 A Convolutional Neural Network and Graph Convolutional Network Based Framework for Classification of Breast Histopathological Images
abstract
The spatial correlation among different tissue components is an essential characteristic for diagnosis of breast cancers based on histopathological images. Graph convolutional network (GCN) can effectively capture this spatial feature representation, and has been successfully applied to the histopathological image based computer-aided diagnosis (CAD). However, the current GCN-based approaches need complicated image preprocessing for graph construction. In this work, we propose a novel CAD framework for classification of breast histopathological images, which integrates both convolutional neural network (CNN) and GCN (named CNN-GCN) into a unified framework, where CNN learns high-level features from histopathological images for further adaptive graph construction, and the generated graph is then fed to GCN to learn the spatial features of histopathological images for the classification task. In particular, a novel clique GCN (cGCN) is proposed to learn more effective graph representation, which can arrange both forward and backward connections between any two graph convolution layers. Moreover, a new group graph convolution is further developed to replace the classical graph convolution of each layer in cGCN, so as to reduce redundant information and implicitly select superior fused feature representation. The proposed clique group GCN (cgGCN) is then embedded in the CNN-GCN framework (named CNN-cgGCN) to promote the learned spatial representation for diagnosis of breast cancers. The experimental results on two public breast histopathological image datasets indicate the effectiveness of the proposed CNN-cgGCN with superior performance to all the compared algorithms.
Zhiyang Gao, Zhiyang Lu, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2022 Self-Supervised Bi-Channel Transformer Networks for Computer-Aided Diagnosis
abstract
Self-supervised learning (SSL) can alleviate the issue of small sample size, which has shown its effectiveness for the computer-aided diagnosis (CAD) models. However, since the conventional SSL methods share the identical backbone in both the pretext and downstream tasks, the pretext network generally cannot be well trained in the pre-training stage, if the pretext task is totally different from the downstream one. In this work, we propose a novel task-driven SSL method, namely Self-Supervised Bi-channel Transformer Networks (SSBTN), to improve the diagnostic accuracy of a CAD model by enhancing SSL flexibility. In SSBTN, we innovatively integrate two different networks for the pretext and downstream tasks, respectively, into a unified framework. Consequently, the pretext task can be flexibly designed based on the data characteristics, and the corresponding designed pretext network thus learns more effective feature representation to be transferred to the downstream network. Furthermore, a transformer-based transfer module is developed to efficiently enhance knowledge transfer by conducting feature alignment between two different networks. The proposed SSBTN is evaluated on two publicly available datasets, namely the full-field digital mammography INbreast dataset and the wireless video capsule CrohnIPI dataset. The experimental results indicate that the proposed SSBTN outperforms all the compared algorithms.
Ronglin Gong, Xiangmin Han, Jun Wang 0024, Shihui Ying, Jun Shi 0004
IEEE J. Biomed. Health Informatics4
2022 Doubly Supervised Transfer Classifier for Computer-Aided Diagnosis With Imbalanced Modalities
abstract
Transfer learning (TL) can effectively improve diagnosis accuracy of single-modal-imaging-based computer-aided diagnosis (CAD) by transferring knowledge from other related imaging modalities, which offers a way to alleviate the small-sample-size problem. However, medical imaging data generally have the following characteristics for the TL-based CAD: 1) The source domain generally has limited data, which increases the difficulty to explore transferable information for the target domain; 2) Samples in both domains often have been labeled for training the CAD model, but the existing TL methods cannot make full use of label information to improve knowledge transfer. In this work, we propose a novel doubly supervised transfer classifier (DSTC) algorithm. In particular, DSTC integrates the support vector machine plus (SVM+) classifier and the low-rank representation (LRR) into a unified framework. The former makes full use of the shared labels to guide the knowledge transfer between the paired data, while the latter adopts the block-diagonal low-rank (BLR) to perform supervised TL between the unpaired data. Furthermore, we introduce the Schatten-p norm for BLR to obtain a tighter approximation to the rank function. The proposed DSTC algorithm is evaluated on the Alzheimer's disease neuroimaging initiative (ADNI) dataset and the bimodal breast ultrasound image (BBUI) dataset. The experimental results verify the effectiveness of the proposed DSTC algorithm.
Xiangmin Han, Xiaoyan Fei, Jun Wang 0024, Tao Zhou 0002, Shihui Ying, Jun Shi 0004, Dinggang Shen
IEEE Trans. Medical Imaging5
2021 Lightweight adaptive weighted network for single image super-resolution
Chaofeng Wang 0003, Jun Wang 0024, Shihui Ying, Jun Shi 0004
Comput. Vis. Image Underst.4
2021 Doubly supervised parameter transfer classifier for diagnosis of breast cancer with imbalanced ultrasound imaging modalities
Xiaoyan Fei, Shichong Zhou, Xiangmin Han, Jun Wang 0024, Shihui Ying, Cai Chang, Jun Shi 0004
Pattern Recognit.5
2021 GCSBA-Net: Gabor-Based and Cascade Squeeze Bi-Attention Network for Gland Segmentation
abstract
Colorectal cancer is the second and the third most common cancer in women and men, respectively. Pathological diagnosis is the "gold standard" for tumor diagnosis. Accurate segmentation of glands from tissue images is a crucial step in assisting pathologists in their diagnosis. The typical methods for gland segmentation form a dense image representation, ignoring its texture and multi-scale attention information. Therefore, we utilize a Gabor-based module to extract texture information at different scales and directions in histopathology images. This paper also designs a Cascade Squeeze Bi-Attention (CSBA) module. Specifically, we add Atrous Cascade Spatial Pyramid (ACSP), Squeeze Position Attention (SPA) module and Squeeze Channel Attention module (SCA) to model semantic correlation and maintain the multi-level aggregation on the spatial pyramid with different dilations. Besides, to solve the imbalance of data distribution and boundary blur, we propose a hybrid loss function to response the object boudary better. The experimental results show that the proposed method achieves state-of-the-art performance on the GlaS challenge dataset and CRAG colorectal adenocarcinoma dataset, respectively.
Zhijie Wen, Ru Feng, Jingxin Liu 0005, Ying Li 0028, Shihui Ying
IEEE J. Biomed. Health Informatics5
2021 Multi-Source Transfer Learning Via Multi-Kernel Support Vector Machine Plus for B-Mode Ultrasound-Based Computer-Aided Diagnosis of Liver Cancers
abstract
B-mode ultrasound (BUS) imaging is a routine tool for diagnosis of liver cancers, while contrast-enhanced ultrasound (CEUS) provides additional information to BUS on the local tissue vascularization and perfusion to promote diagnostic accuracy. In this work, we propose to improve the BUS-based computer aided diagnosis for liver cancers by transferring knowledge from the multi-view CEUS images, including the arterial phase, portal venous phase, and delayed phase, respectively. To make full use of the shared labels of paired of BUS and CEUS images to guide knowledge transfer, support vector machine plus (SVM+), a specifically designed transfer learning (TL) classifier for paired data with shared labels, is adopted for this supervised TL. A nonparallel hyperplane based SVM+ (NHSVM+) is first proposed to improve the TL performance by transferring the per-class knowledge from source domain to the corresponding target domain. Moreover, to handle the issue of multi-source TL, a multi-kernel learning based NHSVM+ (MKL-NHSVM+) algorithm is further developed to effectively transfer multi-source knowledge from multi-view CEUS images. The experimental results indicate that the proposed MKL-NHSVM+ outperforms all the compared algorithms for diagnosis of liver cancers, whose mean classification accuracy, sensitivity, and specificity are 88.18 ± 3.16 %, 86.98 ± 4.77 %, and 89.42±3.77%, respectively.
Lehang Guo, Jun Wang 0024, Lili Bao, Shihui Ying, Huixiong Xu, Jun Shi 0004
IEEE J. Biomed. Health Informatics6
2020 CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence
abstract
In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective function through bidirectional correspondence. Globally, error metric of correntropy is introduced as noise model to resist outliers. Importantly, the close resemblance between normal-distributions transform (NDT) and correntropy is revealed. To ease the minimization step, an on-manifold parameterization of the special Euclidean group is proposed. Extensive experiments validate that CoBigICP outperforms several well-known and state-of-the-art methods.
Pengyu Yin, Di Wang 0028, Shaoyi Du, Shihui Ying, Yue Gao 0002, Nanning Zheng 0001
IROS4
2020 Deep Doubly Supervised Transfer Network for Diagnosis of Breast Cancer with Imbalanced Ultrasound Imaging Modalities
Xiangmin Han, Jun Wang 0024, Cai Chang, Shihui Ying, Jun Shi 0004
MICCAI (6)5
2020 IExpressNet: Facial Expression Recognition with Incremental Classes
abstract
Existing methods on facial expression recognition (FER) are mainly trained in the setting when all expression classes are fixed in advance. However, in real applications, expression classes are becoming increasingly fine-grained and incremental. To deal with sequential expression classes, we can fine-tune or re-train these models, but this often results in poor performance or large computing resources consumption. To address these problems, we develop an Incremental Facial Expression Recognition Network (IExpressNet), which can learn a competitive multi-class classifier at any time with a lower requirement of computing resources. Specifically, IExpressNet consists of two novel components. First, we construct an exemplar set by dynamically selecting representative samples from old expression classes. Then, the exemplar set and new expression classes samples constitute the training set. Second, we design a novel center-expression-distilled loss. As for facial expression in the wild, center-expression-distilled loss enhances the discriminative power of the deeply learned features and prevents catastrophic forgetting. Extensive experiments are conducted on two large-scale FER datasets in the wild, RAF-DB and AffectNet. The results demonstrate the superiority of the proposed method as compared to state-of-the-art incremental learning approaches.
Bingjun Luo, Sicheng Zhao, Shihui Ying, Xibin Zhao, Yue Gao 0002
ACM Multimedia4
2020 Multi-metric Joint Discrimination Network for Few-Shot Classification
Zhijie Wen, Liyan Ma, Shihui Ying
PRCV (3)4
2020 Projective parameter transfer based sparse multiple empirical kernel learning Machine for diagnosis of brain disease
Xiaoyan Fei, Jun Wang 0024, Shihui Ying, Zhongyi Hu 0001, Jun Shi 0004
Neurocomputing3
2019 Manifold Alignment and Distribution Adaptation for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation is a problem which exploits the knowledge learned from the resource-rich domain to obtain an accurate classifier for the resource-poor domain. Most of the existing methods lift performance by reducing the differences between distributions, such as the difference between marginal probability distributions, the difference between conditional probability distributions, or both. However, all these methods consider the two distributions to be equally important, which could lead to poor classification performance in practical applications. Therefore, a balanced factor is required to weigh the two distributions to compensate for the degraded performance. In this paper, we first introduce this balance factor to weigh the distribution importance. On this base, we utilize the marginal distribution, introduce the ideas of manifold regularization, and then preserve the neighboring structures of the data sets, with the dimension reduction as much as possible. By this way, we propose the manifold alignment and balanced distribution adaptation algorithm. A large number of experiments have also been conducted, showing that our algorithm behaves much better than the previous ones.
Ying Li 0028, Yaxin Peng, Zhijie Wen, Shihui Ying
ICME5
2019 Asymmetric Local Metric Learning with PSD Constraint for Person Re-identification
abstract
Person re-identification is one of the key issues in both machine learning and video monitor application. In particular, defining an appropriate distance metric between the person images is very important. Existing metric learning approaches used in person re-identification either learn a single measure, or ignore the positive semi-definite (PSD) of measurement matrix, at the same time, since the number of negative sample pairs largely exceeds the number of positive sample pairs, some metric learning methods are largely influenced by the sample imbalance. Considering the above issues, we propose a new adaptive local metric learning method with positive semi-definite (PSD) constraint. Unlike existing metric learning methods which learn a single distance metric, we use an approximation error bound of a smooth metric matrix function over the data manifold to learn local metrics as linear combinations of basis metrics defined on anchor points over different regions of the instance space. Besides, we develop an efficient two stage algorithm that first learns the anchor points and the linear combinations of each instance, then learns the metric matrices of the anchor points. We employ the fast iterative shrinkage-thresholding algorithm which is a fast first-order optimization algorithm in the learning process of the linear combinations as well as the basis metrics of the anchor points. Our metric learning method has excellent performance. We firstly apply the proposed method on 5 UCI databases, which are widely used in machine learning, to test and evaluate the effectiveness of the proposed method. Then the proposed approach is applied for person re-identification, achieving better performance on three challenging databases (GRID, VIPeR, CUHK01) than the existing methods. The experimental results show that the proposed method can prvide the theoretical and practical support for the person re-identification problem.
Zhijie Wen, Ying Li 0028, Shihui Ying, Yaxin Peng
ICRA4
2019 Geometric Understanding for Unsupervised Subspace Learning
abstract
In this paper, we address the unsupervised subspace learning from a geometric viewpoint. First, we formulate the subspace learning as an inverse problem on Grassmannian manifold by considering all subspaces as points on it. Then, to make the model computable, we parameterize the Grassmannian manifold by using an orbit of rotation group action on all standard subspaces, which are spanned by the orthonormal basis. Further, to improve the robustness, we introduce a low-rank regularizer which makes the dimension of subspace as low as possible. Thus, the subspace learning problem is transferred to a minimization problem with variables of rotation and dimension. Then, we adopt the alternately iterative strategy to optimize the variables, where a structure-preserving method, based on the geodesic structure of the rotation group, is designed to update the rotation. Finally, we compare the proposed approach with six state-of-the-art methods on three different kinds of real datasets. The experimental results validate that our proposed method outperforms all compared methods.
Shihui Ying, Lipeng Cai, Changzhou He, Yaxin Peng
IJCAI1
2019 Quaternion Grassmann average network for learning representation of histopathological image
Jun Shi 0004, Jinjie Wu, Bangming Gong, Qi Zhang 0003, Shihui Ying
Pattern Recognit.6
2019 MR Image Super-Resolution via Wide Residual Networks With Fixed Skip Connection
abstract
Spatial resolution is a critical imaging parameter in magnetic resonance imaging. The image super-resolution (SR) is an effective and cost efficient alternative technique to improve the spatial resolution of MR images. Over the past several years, the convolutional neural networks (CNN)-based SR methods have achieved state-of-the-art performance. However, CNNs with very deep network structures usually suffer from the problems of degradation and diminishing feature reuse, which add difficulty to network training and degenerate the transmission capability of details for SR. To address these problems, in this work, a progressive wide residual network with a fixed skip connection (named FSCWRN) based SR algorithm is proposed to reconstruct MR images, which combines the global residual learning and the shallow network based local residual learning. The strategy of progressive wide networks is adopted to replace deeper networks, which can partially relax the above-mentioned problems, while a fixed skip connection helps provide rich local details at high frequencies from a fixed shallow layer network to subsequent networks. The experimental results on one simulated MR image database and three real MR image databases show the effectiveness of the proposed FSCWRN SR algorithm, which achieves improved reconstruction performance compared with other algorithms.
Jun Shi 0004, Shihui Ying, Chaofeng Wang 0003, Qingping Liu, Qi Zhang 0003, Pingkun Yan
IEEE J. Biomed. Health Informatics3
2018 Global Nonlinear Metric Learning by Gluing Local Linear Metrics
abstract
We address the nonlinear metric learning by constructing a smooth nonlinear metric from the data. First, we locally define an initial linear metric on each cluster by principal component analysis. Second, we glue such local linear metrics to form a smooth nonlinear metric by a partition of unity on the sample space, and further learn the global nonlinear metric. Third, we conduct the intrinsic steepest descent algorithm on matrix manifolds for implementation. Finally, we compare our approach with several state-of-the-art methods on a variety of datasets. The results validate that the robustness and accuracy of classification are both improved under our nonlinear metric. The novelty of our global smooth nonlinear metric learning model lies in that it has completely overcome drawbacks of local metric learning methods: the partition coefficients obtained by the partition of unity is smooth, while the metric at any point on the manifold can be directly defined.
Yaxin Peng, Lingfang Hu, Shihui Ying, Chaomin Shen 0001
SDM3
2018 Nonlinear Semi-Supervised Metric Learning Via Multiple Kernels and Local Topology
abstract
Changing the metric on the data may change the data distribution, hence a good distance metric can promote the performance of learning algorithm. In this paper, we address the semi-supervised distance metric learning (ML) problem to obtain the best nonlinear metric for the data. First, we describe the nonlinear metric by the multiple kernel representation. By this approach, we project the data into a high dimensional space, where the data can be well represented by linear ML. Then, we reformulate the linear ML by a minimization problem on the positive definite matrix group. Finally, we develop a two-step algorithm for solving this model and design an intrinsic steepest descent algorithm to learn the positive definite metric matrix. Experimental results validate that our proposed method is effective and outperforms several state-of-the-art ML methods.
Yanqin Bai, Yaxin Peng, Shaoyi Du, Shihui Ying
Int. J. Neural Syst.5
2018 Neuroimaging-based diagnosis of Parkinson's disease with deep neural mapping large margin distribution machine
Bangming Gong, Jun Shi 0004, Shihui Ying, Yakang Dai, Qi Zhang 0003, Hedi An, Yingchun Zhang
Neurocomputing3
2018 A P-ADMM for sparse quadratic kernel-free least squares semi-supervised support vector machine
Yaru Zhan, Yanqin Bai, Wei Zhang 0172, Shihui Ying
Neurocomputing4
2018 Multimodal Neuroimaging Feature Learning With Multimodal Stacked Deep Polynomial Networks for Diagnosis of Alzheimer's Disease
abstract
The accurate diagnosis of Alzheimer's disease (AD) and its early stage, i.e., mild cognitive impairment, is essential for timely treatment and possible delay of AD. Fusion of multimodal neuroimaging data, such as magnetic resonance imaging (MRI) and positron emission tomography (PET), has shown its effectiveness for AD diagnosis. The deep polynomial networks (DPN) is a recently proposed deep learning algorithm, which performs well on both large-scale and small-size datasets. In this study, a multimodal stacked DPN (MM-SDPN) algorithm, which MM-SDPN consists of two-stage SDPNs, is proposed to fuse and learn feature representation from multimodal neuroimaging data for AD diagnosis. Specifically speaking, two SDPNs are first used to learn high-level features of MRI and PET, respectively, which are then fed to another SDPN to fuse multimodal neuroimaging information. The proposed MM-SDPN algorithm is applied to the ADNI dataset to conduct both binary classification and multiclass classification tasks. Experimental results indicate that MM-SDPN is superior over the state-of-the-art multimodal feature-learning-based algorithms for AD diagnosis.
Jun Shi 0004, Yan Li 0066, Qi Zhang 0003, Shihui Ying
IEEE J. Biomed. Health Informatics5
2018 Manifold Preserving: An Intrinsic Approach for Semisupervised Distance Metric Learning
abstract
In this paper, we address the semisupervised distance metric learning problem and its applications in classification and image retrieval. First, we formulate a semisupervised distance metric learning model by considering the metric information of inner classes and interclasses. In this model, an adaptive parameter is designed to balance the inner metrics and intermetrics by using data structure. Second, we convert the model to a minimization problem whose variable is symmetric positive-definite matrix. Third, in implementation, we deduce an intrinsic steepest descent method, which assures that the metric matrix is strictly symmetric positive-definite at each iteration, with the manifold structure of the symmetric positive-definite matrix manifold. Finally, we test the proposed algorithm on conventional data sets, and compare it with other four representative methods. The numerical results validate that the proposed method significantly improves the classification with the same computational efficiency.
Shihui Ying, Zhijie Wen, Jun Shi 0004, Yaxin Peng, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.1
2017 Histopathological Image Classification With Color Pattern Random Binary Hashing-Based PCANet and Matrix-Form Classifier
abstract
The computer-aided diagnosis for histopathological images has attracted considerable attention. Principal component analysis network (PCANet) is a novel deep learning algorithm for feature learning with the simple network architecture and parameters. In this study, a color pattern random binary hashing-based PCANet (C-RBH-PCANet) algorithm is proposed to learn an effective feature representation from color histopathological images. The color norm pattern and angular pattern are extracted from the principal component images of R, G, and B color channels after cascaded PCA networks. The random binary encoding is then performed on both color norm pattern images and angular pattern images to generate multiple binary images. Moreover, we rearrange the pooled local histogram features by spatial pyramid pooling to a matrix-form for reducing the dimension of feature and preserving spatial information. Therefore, a C-RBH-PCANet and matrix-form classifier-based feature learning and classification framework is proposed for diagnosis of color histopathological images. The experimental results on three color histopathological image datasets show that the proposed C-RBH-PCANet algorithm is superior to the original PCANet and other conventional unsupervised deep learning algorithms, while the best performance is achieved by the proposed feature learning and classification framework that combines C-RBH-PCANet and matrix-form classifier.
Jun Shi 0004, Jinjie Wu, Yan Li 0066, Qi Zhang 0003, Shihui Ying
IEEE J. Biomed. Health Informatics5
2016 Compute Karcher means on SO(n) by the geometric conjugate gradient method
Shihui Ying, Han Qin, Yaxin Peng, Zhijie Wen
Neurocomputing1
2016 Nonlinear 2D shape registration via thin-plate spline and Lie group representation
Shihui Ying, Yuanwei Wang, Zhijie Wen, Yuping Lin
Neurocomputing1
2016 Virus image classification using multi-scale completed local binary pattern features extracted from filtered images by multi-scale principal component analysis
Zhijie Wen, Zhuojun Li, Yaxin Peng, Shihui Ying
Pattern Recognit. Lett.4
2014 LieTrICP: An improvement of trimmed iterative closest point algorithm
Yaxin Peng, Shihui Ying, Zhiyu Hu
Neurocomputing3
2013 Groupwise Registration via Graph Shrinkage on the Image Manifold
abstract
Recently, group wise registration has been investigated for simultaneous alignment of all images without selecting any individual image as the template, thus avoiding the potential bias in image registration. However, none of current group wise registration method fully utilizes the image distribution to guide the registration. Thus, the registration performance usually suffers from large inter-subject variations across individual images. To solve this issue, we propose a novel group wise registration algorithm for large population dataset, guided by the image distribution on the manifold. Specifically, we first use a graph to model the distribution of all image data sitting on the image manifold, with each node representing an image and each edge representing the geodesic pathway between two nodes (or images). Then, the procedure of warping all images to their population center turns to the dynamic shrinking of the graph nodes along their graph edges until all graph nodes become close to each other. Thus, the topology of image distribution on the image manifold is always preserved during the group wise registration. More importantly, by modeling the distribution of all images via a graph, we can potentially reduce registration error since every time each image is warped only according to its nearby images with similar structures in the graph. We have evaluated our proposed group wise registration method on both synthetic and real datasets, with comparison to the two state-of-the-art group wise registration methods. All experimental results show that our proposed method achieves the best performance in terms of registration accuracy and robustness.
Shihui Ying, Guorong Wu 0001, Qian Wang 0001, Dinggang Shen
CVPR1
2013 Soft shape registration under Lie group frame
abstract
In this study, the authors address a two‐dimensional (2D) shape registration problem on data with anisotropic‐scale deformation and noise. First, the model is formulated under the iterative closest point (ICP) framework, which is one of the most popular methods for shape registration. To overcome the effect of noise, the expectation maximisation algorithm is used to improve the model. Then, the structure of Lie groups is adopted to parameterise the proposed model, which provides a unified framework to deal with the shape registration problems. Such representation makes it possible to introduce some suitable constraints to the model, which improves the robustness of the algorithm. Thereby, the 2D shape registration problem is turned to an optimisation problem on the matrix Lie group. Furthermore, a sequence of quadratic programming is designed to approximate the solution for the model. Finally, several comparative experiments are carried out to validate that the authors’ algorithm performs well in terms of robustness, especially in the presence of outliers.
Yaxin Peng, Shihui Ying
IET Comput. Vis.3
2011 Iwasawa decomposition: a new approach to 2D affine registration problem
Shihui Ying, Yaxin Peng, Zhijie Wen
Pattern Anal. Appl.1
2010 Scaling iterative closest point algorithm for registration of m-D point sets
Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Jianru Xue
J. Vis. Commun. Image Represent.4
2010 Affine iterative closest point algorithm for point set registration
Shaoyi Du, Nanning Zheng 0001, Shihui Ying
Pattern Recognit. Lett.3
2009 Lie Group Framework of Iterative Closest Point Algorithm for n-d Data Registration
abstract
The iterative closet point (ICP) method is a dominant method for data registration that has attracted extensive attention. In this paper, a unified mathematical model of ICP based on Lie group representation is established. Under the framework, the registration problem is formulated into an optimization problem over a certain Lie group. In order to simplify the model and to reduce the dimension of parameter space, the translation part of geometric transformation is eliminated by calibrating the centers of two data sets under registration. As a result, a fast algorithm by solving an iterative linear system is designed for the optimization problem on Lie groups. Moreover, PCA and ICA methods are jointly applied to estimate the initial registration to achieve the global minimum. Finally, several illustrations and comparison experiments are presented to test the performance of the proposed algorithm.
Shihui Ying, Shaoyi Du, Hong Qiao
Int. J. Pattern Recognit. Artif. Intell.1
2009 A Scale Stretch Method Based on ICP for 3D Data Registration
abstract
In this paper, we are concerned with the registration of two 3D data sets with large-scale stretches and noises. First, by incorporating a scale factor into the standard iterative closest point (ICP) algorithm, we formulate the registration into a constraint optimization problem over a 7D nonlinear space. Then, we apply the singular value decomposition (SVD) approach to iteratively solving such optimization problem. Finally, we establish a new ICP algorithm, named Scale-ICP algorithm, for registration of the data sets with isotropic stretches. In order to achieve global convergence for the proposed algorithm, we propose a way to select the initial registrations. To demonstrate the performance and efficiency of the proposed algorithm, we give several comparative experiments between Scale-ICP algorithm and the standard ICP algorithm.
Shihui Ying, Shaoyi Du, Hong Qiao
IEEE Trans Autom. Sci. Eng.1
2007 AN Extension of the ICP Algorithm Considering Scale Factor
abstract
The ICP algorithm is accurate and fast for registration between two point sets in a same scale, but it doesn't handle the case with different scales. This paper instead introduces a novel approach named the scaling iterative closest point (SICP) algorithm which integrates a scale matrix with boundaries into the original ICP algorithm for scaling registration. This method uses a simple iterative algorithm with the SVD algorithm and the properties of parabola incorporated to compute the translation, rotation and scale transformations at each iterative step, and its convergence is rapid with only a few iterations. The SICP algorithm is independent of shape representation and feature extraction; thereby it is general for scaling registration. Experimental results demonstrate its robustness and fast speed compared with the standard ICP algorithm.
Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Qubo You, Yang Wu 0001
ICIP (5)3
2007 ICP with Bounded Scale for Registration of M-D Point Sets
abstract
The iterative closest point (ICP) algorithm is an accurate and fast approach for registration between two point sets in a same scale, but it doesn't handle the case with different scales. This paper instead introduces a novel approach named the iterative closest point with bounded scale (ICPBS) algorithm which integrates a scale with boundaries into the traditional ICP algorithm. This proposed technique uses the singular value decomposition algorithm and the properties of parabola to compute the similar transformation at each iterative step, and yields more satisfying robust results than the traditional ICP method in registration between two m-D point sets with different scales. Experimental results demonstrate the presented method is robust and fast for practical use.
Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Jishang Wei
ICME3