EDBT 2026 Demo / reviewers in the wild / expert
Youyong Kong
dblp:154/7641
· DBLP profile ↗
66ranked-venue papers
5as first author
47since 2021 · last 2026
0000-0003-2095-8470ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 1 first-author · 26 since 2021Artificial intelligence and machine learning · 30 · 2 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Intraoperative 2D/3D Registration via Spherical Similarity Learning and Differentiable Levenberg-Marquardt OptimizationabstractIntraoperative 2D/3D registration aligns preoperative 3D volumes with real-time 2D radiographs, enabling accurate localization of instruments and implants. A recent fully differentiable similarity learning framework approximates geodesic distances on SE(3), expanding the capture range of registration and mitigating the effects of substantial disturbances, but existing Euclidean approximations distort manifold structure and slow convergence. To address the above limitations, we explore similarity learning on non-Euclidean spherical feature spaces to improve the ability to capture and fit complex manifold features. We extract feature embeddings using a CNN-Transformer encoder, project them into spherical space, and approximate their geodesic distances with Riemannian geodesic distances in the bi-invariant SO(4) space. This enables the learning of a more expressive and geometrically consistent deep similarity metric, enhancing the network’s ability to distinguish subtle pose differences. Fully differentiable Levenberg-Marquardt optimization is adopted to replace the existing gradient descent method to accelerate the convergence of the search during inference phase. Experiments on real and synthetic datasets show superior accuracy in both patient-specific and patient-agnostic scenarios. Minheng Chen, Youyong Kong |
WACV | 2 |
| 2026 | SpineCLUE: Automatic vertebrae identification using contrastive learning and uncertainty estimation
Minheng Chen, Mingying Li, Junxian Wu 0002, Cheng Xue 0003, Youyong Kong |
Artif. Intell. Medicine | 8 |
| 2026 | Embedding uncertainty modeling for cold-start item recommendation
Youyong Kong |
Neurocomputing | 2 |
| 2026 | Neurobridge: Bridging functional and structural brain networks via neural coupling and consistency-Guided dynamic graph learning
Xiaoyun Liu, Yonggui Yuan, Youyong Kong |
Medical Image Anal. | 5 |
| 2025 | HePa: Heterogeneous Graph Prompting for All-Level Classification TasksabstractHeterogeneous graphs, which are common in real-world downstream tasks, have recently sparked a wave of research interest. The performance of end-to-end heterogeneous graph neural networks (HGNNs) greatly relies on supervised training for specific tasks. To reduce the labeling cost, the "pretrain-finetune" paradigm has been widely adopted, but it leads to a knowledge gap between the pre-trained model and downstream tasks. In an effort to address this gap, the "pretrain-prompt" paradigm has emerged as a promising approach. This involves fine-tuning randomly initialized learnable vectors in downstream tasks. However, this approach may result in an insufficient representation of downstream task features. Existing techniques for heterogeneous graph prompting restructure the heterogeneous graph to align with the homogeneous graph prompting scheme. This can potentially introduce the same limitations as homogeneous graph prompt learning. In this paper, we propose HePa, short for Heterogeneous Graph Prompting for all-level classification tasks. It not only includes a unified prompt template-graph adapted for heterogeneous graphs but also introduces a novel pre-prompt token optimized during the pre-training phase to convey task information downstream. With these designs, HePa can complete all levels of classification tasks toward few-shot scenarios while activating in-context learning. Finally, we conducted a comprehensive experimental analysis of HePa on three benchmark datasets. Jia Jinghong, Lei Song 0013, Youyong Kong |
AAAI | 4 |
| 2025 | Exploring Rationale Learning for Continual Graph LearningabstractCatastrophic forgetting poses a significant challenge for graph neural networks in continuously updating their knowledge base with data streams. To address this issue, much of the research has focused on node-level continual learning using parameter regularization or rehearsal-based strategies, while little attention given to graph-level tasks. Furthermore, current paradigms for continual graph learning may inadvertently capture spurious correlations for specific tasks through shortcuts, thereby exacerbating the forgetting of previous knowledge when new tasks are introduced. To tackle these challenges, we propose a novel paradigm, Rationale Learning GNN (RL-GNN), for graph-level continual graph learning. Specifically, we harness the invariant learning principle to incorporate environmental interventions into both the current and historical distributions, aiming to uncover rationales by minimizing empirical risk across all environments. The rationale serves as the sole factor guiding the learning process. Therefore, continual graph learning is redefined as capturing these invariant rationales within task sequences, alleviating catastrophic forgetting caused by spurious features. Extensive experiments on real-world datasets with varying task lengths demonstrate the effectiveness of our RL-GNN in continuous knowledge assimilation and reduction of catastrophic forgetting. Lei Song 0013, Qinghua Si, Shihan Guan, Youyong Kong |
AAAI | 5 |
| 2025 | Learning Heterogeneous Tissues with Mixture of Experts for Gigapixel Whole Slide ImagesabstractAnalyzing gigapixel Whole Slide Images (WSIs) is challenging due to the complex pathological tissue environment and the absence of target-driven domain knowledge. Previous methods incorporated pathological priors to mitigate this issue but relied on additional inference steps and specialized workflows, restricting scalability and the model’s capacity to identify novel outcome-related factors. To address these challenges, we propose a plug-and-play Pathology-Aware Mixture-of-Experts (PAMoE) module, which based on mixture of experts to learn pathology-related knowledge and extract useful information. We train the experts to become ‘specialists’ in specific intratumoral tissues by learning to route each tissue to its mapped expert. In addition, to reduce the impact of irrelevant content on the model, we introduce a new routing rule that discards patches in which none of the experts express interest, which helps the model better capture the relationships between relevant patches. Through a comprehensive evaluation of PAMoE on survival task, we demonstrate that 1) Our module enhances the performance of baseline models in most cases, and 2) The sparse expert processing across different tissues enhances the learning of patch representations by addressing tissue heterogeneity. Source code is available at https://github.com/wjx-error/PAMoE. Junxian Wu 0002, Minheng Chen, Xinyi Ke, Tianwang Xun, Xiaoming Jiang, Lizhi Shao, Youyong Kong |
CVPR | 8 |
| 2025 | M2F2Net: Multi-stage Mixed Feature Fusion Network For Remote Sensing Change DetectionabstractRemote sensing image change detection (CD) seeks to analyze and discern changes in surface objects through the use of multi-temporal remote sensing imagery. However, as image resolution advances, existing methods often fall short in capturing comprehensive visual feature representations, and their networks are prone to spatial degradation. This results in incomplete boundary detection and difficulties in identifying subtle changes. To overcome these challenges, this paper introduces a Siamese U-Net architecture incorporating Multistage Mixed Feature Fusion (M2F2Net). The proposed model leverages a convolutional neural network (CNN) as the main encoder for local feature extraction, while employing a Transformer-based auxiliary encoder to capture global features. Furthermore, we introduce a Feature Fusion Module (FFM) to facilitate the efficient integration of local and global information. Additionally, we propose a novel convolutional unit and a Spatial Attention Module (SAM) designed to enhance the extraction of image features. Experimental results confirm that the proposed approach delivers substantial improvements across multiple evaluation criteria, while offering a superior accuracy compared to existing state-of-the-art change detection methods. Binhao Gu, Lei Song 0013, Youyong Kong, Binjie Gu |
ICASSP | 3 |
| 2025 | Soft Augmentation for Graph ClassificationabstractGraph data augmentation proves to be an effective approach for enhancing the performance of graph classification. However, due to the complex structure of graphs, the semantic meanings of graphs are sensitive to minor modifications, while the labels of augmented graphs remain identical, thus limiting the potential benefits of graph data augmentation. To address this limitation, we propose Graph Soft Augmentation (GSA), a method to smooth the labels of augmented graphs. Instead of assigning hard and fixed labels to augmented graphs as traditional graph data augmentations, which do not consider the changed semantics, GSA smooths the labels of augmented graphs. GSA can be divided into two stages. In the first stage, a graph similarity network is trained with the original dataset until convergence. In the second stage, GSA adopts a general graph data augmentation to construct augmented graphs. The labels of augmented graphs are then smoothed based on the similarities computed by the similarity network. Finally, the resulting augmented graphs, along with their smoothed labels, are incorporated into the graph classification network as training samples. Experimental results on a variety of publicly available datasets reveal the effectiveness of our GSA. Weihuang Zheng, Youyong Kong |
ICASSP | 4 |
| 2025 | Topology-Aware Dynamic Reweighting for Distribution Shifts on GraphabstractGraph Neural Networks (GNNs) are widely used for node classification tasks but often fail to generalize when training and test nodes come from different distributions, limiting their practicality. To address this challenge, recent approaches have adopted invariant learning and sample reweighting techniques from the out-of-distribution (OOD) generalization field. However, invariant learning-based methods face difficulties when applied to graph data, as they rely on the impractical assumption of obtaining real environment labels and strict invariance, which may not hold in real-world graph structures. Moreover, current sample reweighting methods tend to overlook topological information, potentially leading to suboptimal results. In this work, we introduce the Topology-Aware Dynamic Reweighting (TAR) framework to address distribution shifts by leveraging the inherent graph structure. TAR dynamically adjusts sample weights through gradient flow on the graph edges during training. Instead of relying on strict invariance assumptions, we theoretically prove that our method is able to provide distributional robustness, thereby enhancing the out-of-distribution generalization performance on graph data. Our framework's superiority is demonstrated through standard testing on extensive node classification OOD datasets, exhibiting marked improvements over existing methods. Weihuang Zheng, Jiayun Wu, Peng Cui 0001, Youyong Kong |
ICML | 6 |
| 2025 | Suit the Node Pair to the Case: A Multi-Scale Node Pair Grouping Strategy for Graph-MLP DistillationabstractGraph Neural Network (GNN) is powerful in solving various graph-related tasks, while its message passing mechanism may lead to latency during inference time. Multi-Layer-Perceptron (MLP) can achieve fast inference speed but with limited performance. One solution to fill this gap is through Knowledge Distillation. However, current distillation methods follow a ''node-to-node'' paradigm, while considering the complex relationships between different node pairs, direct distillation fails to capture these multiple-granularity features in GNN. Furthermore, current methods which focuses on the alignment of logits in the final layer ignores further learning within layers inside student MLP. Therefore, in this paper, we introduce a multi-scale knowledge distillation method (MSN-GDM) aiming to capture multiple knowledge from GNN to MLP. We firstly propose a multi-scale node-pair grouping strategy to assign node pairs to different-scale groups according to node pair similarity metrics. The similarity metrics consider both node features and topological structures of the given node pair. Then based on the preprocessed node-set groups, we design a multi-scale distillation method that can capture comprehensive knowledge in the corresponding node-set groups. The hierarchical weighted sum of each layer is applied as the final output. Extensive experiments on eight real-world datasets demonstrate the effectiveness of our proposed method. Weihuang Zheng, Youyong Kong |
IJCAI | 4 |
| 2025 | DiffuER: Auxiliary Regularized Diffusion Model for Text Generation with Semantic ConsistencyabstractRecent advances in diffusion models have demonstrated their remarkable potential in text generation, achieving performance that rivals or even exceeds autoregressive language models. However, existing diffusion language models face two critical challenges. First, they lack explicit mechanisms for semantic constraints, which is crucial for maintaining semantic consistency. Second, their non-autoregressive generation process at each timestep leads to insufficient token dependencies, resulting in semantic incompleteness or repetition in the generated text. To address these issues, we propose DiffuER (Diffusion with Embedding and Reconstruction loss). By utilizing word embeddings from a pre-trained masked language model as auxiliary regularization and leveraging an encoder-decoder module to reconstruct the original text from the generated sequence, DiffuER strengthens the model’s ability to capture semantics and maintain semantic consistency. Comprehensive experiments on four text generation tasks demonstrate DiffuER’s superiority over autoregressive, non-autoregressive models and diffusion-based baselines in terms of generation quality, diversity, and semantic consistency. Notably, our method achieves a 52.21 increase in generation diversity compared to autoregressive models, and a 38.7 decrease in perplexity compared to diffusion models. Our code is available at https://github.com/miaoGao/DiffuER. Chuanqi Shi, Baixuan Li, Yikemaiti Sataer, Youyong Kong |
IJCNN | 5 |
| 2025 | Prompt-driven graph distillation: Enabling single-layer MLPs to outperform Deep Graph Neural Networks in graph-based tasksabstractGraph distillation endeavors to transfer knowledge from large, complex teacher models, such as Graph Neural Networks (GNNs), to smaller, more efficient student models such as Multi-Layer Perceptrons (MLPs). Our work is motivated by a critical observation: as the complexity and depth of the teacher GNN increase, the performance of the distilled student MLP models tends to decline significantly. This issue is not limited to a specific distillation method but is a prevalent challenge across various GNN to MLP approaches. To address this gap, we introduce a novel prompt-driven graph distillation framework that enhances the student model’s input space by appending or integrating learned prompts with the original features. These prompts, derived from the teacher’s input, hidden and output knowledge, provide additional context that assists the MLP student in distilling information, thereby circumventing the need to directly capture any teacher’s knowledge into the simplest single-layer MLP. Our empirical experiment on benchmark graph datasets reveals a counter-intuitive phenomenon: the more capable the teacher, the greater the student’s ability to utilize prompts to outperform the teacher. This finding highlights the strength of our approach: a well-prompted student can indeed surpass its teacher, at achieving the best performance with an accuracy of 87.2% on the Cora standard split with a single-layer MLP, while also maintaining efficiency and robustness. In addition to testing on benchmark graph datasets, we applied our framework to the TUH EEG Epilepsy Corpus (TUEP), condensing complex GNN models to simple MLPs while achieving good performance and efficiency in epilepsy classification. Shihan Guan, Laurent Albera, Lei Song 0013, Youyong Kong, Huazhong Shu, Régine Le Bouquin-Jeannès |
Neurocomputing | 5 |
| 2025 | MSARAE: Multiscale adversarial regularized autoencoders for cortical network classification
Yihui Zhu, Yonggui Yuan, Youyong Kong |
Medical Image Anal. | 6 |
| 2025 | Multi-Scale Spatial-Temporal Attention Networks for Functional Connectome ClassificationabstractMany neuropsychiatric disorders are considered to be associated with abnormalities in the functional connectivity networks of the brain. The research on the classification of functional connectivity can therefore provide new perspectives for understanding the pathology of disorders and contribute to early diagnosis and treatment. Functional connectivity exhibits a nature of dynamically changing over time, however, the majority of existing methods are unable to collectively reveal the spatial topology and time-varying characteristics. Furthermore, despite the efforts of limited spatial-temporal studies to capture rich information across different spatial scales, they have not delved into the temporal characteristics among different scales. To address above issues, we propose a novel Multi-Scale Spatial-Temporal Attention Networks (MSSTAN) to exploit the multi-scale spatial-temporal information provided by functional connectome for classification. To fully extract spatial features of brain regions, we propose a Topology Enhanced Graph Transformer module to guide the attention calculations in the learning of spatial features by incorporating topology priors. A Multi-Scale Pooling Strategy is introduced to obtain representations of brain connectome at various scales. Considering the temporal dynamic characteristics between dynamic functional connectome, we employ Locality Sensitive Hashing attention to further capture long-term dependencies in time dynamics across multiple scales and reduce the computational complexity of the original attention mechanism. Experiments on three brain fMRI datasets of MDD and ASD demonstrate the superiority of our proposed approach. In addition, benefiting from the attention mechanism in Transformer, our results are interpretable, which can contribute to the discovery of biomarkers. The code is available at https://github.com/LIST-KONG/MSSTAN. Youyong Kong, Wenhan Wang, Yonggui Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Wavelet-Based Dual-Task NetworkabstractIn image processing, wavelet transform (WT) offers multiscale image decomposition, generating a blend of low-resolution approximation images and high-resolution detail components. Drawing parallels to this concept, we view feature maps in convolutional neural networks (CNNs) as a similar mix, but uniquely within the channel domain. Inspired by multitask learning (MTL) principles, we propose a wavelet-based dual-task (WDT) framework. This novel framework employs WT in the channel domain to split a single task into two parallel tasks, thereby reforming traditional single-task CNNs into dynamic dual-task networks. Our WDT framework integrates seamlessly with various popular network architectures, enhancing their versatility and efficiency. It offers a more rational approach to resource allocation in CNNs, balancing between low-frequency and high-frequency information. Rigorous experiments on Cifar10, ImageNet, HMDB51, and UCF101 validate our approach's effectiveness. Results reveal significant improvements in the performance of traditional CNNs on classification tasks, and notably, these enhancements are achieved with fewer parameters and computations. In summary, our work presents a pioneering step toward redefining the performance and efficiency of CNN-based tasks through WT. Fuzhi Wu, Jiasong Wu, Chen Zhang 0024, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Multiscale Low-Frequency Memory Network for Improved Feature Extraction in Convolutional Neural NetworksabstractDeep learning and Convolutional Neural Networks (CNNs) have driven major transformations in diverse research areas. However, their limitations in handling low-frequency in-formation present obstacles in certain tasks like interpreting global structures or managing smooth transition images. Despite the promising performance of transformer struc-tures in numerous tasks, their intricate optimization com-plexities highlight the persistent need for refined CNN en-hancements using limited resources. Responding to these complexities, we introduce a novel framework, the Mul-tiscale Low-Frequency Memory (MLFM) Network, with the goal to harness the full potential of CNNs while keep-ing their complexity unchanged. The MLFM efficiently preserves low-frequency information, enhancing perfor-mance in targeted computer vision tasks. Central to our MLFM is the Low-Frequency Memory Unit (LFMU), which stores various low-frequency data and forms a parallel channel to the core network. A key advantage of MLFM is its seamless compatibility with various prevalent networks, requiring no alterations to their original core structure. Testing on ImageNet demonstrated substantial accuracy improvements in multiple 2D CNNs, including ResNet, MobileNet, EfficientNet, and ConvNeXt. Furthermore, we showcase MLFM's versatility beyond traditional image classification by successfully integrating it into image-to-image translation tasks, specifically in semantic segmenta-tion networks like FCN and U-Net. In conclusion, our work signifies a pivotal stride in the journey of optimizing the ef-ficacy and efficiency of CNNs with limited resources. This research builds upon the existing CNN foundations and paves the way for future advancements in computer vision. Our codes are available at https://github.com/AlphaWuSeu/MLFM. Fuzhi Wu, Jiasong Wu, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
AAAI | 3 |
| 2024 | ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images
Xiangtian Xue, Jiasong Wu, Youyong Kong, Lotfi Senhadji, Huazhong Shu |
ECCV (46) | 3 |
| 2024 | Embedded Feature Similarity Optimization with Specific Parameter Initialization for 2D/3D Medical Image RegistrationabstractWe present a novel deep learning-based framework: Embedded Feature Similarity Optimization with Specific Parameter Initialization (SOPI) for 2D/3D medical image registration which is a most challenging problem due to the difficulty such as dimensional mismatch, heavy computation load and lack of golden evaluation standard. The framework we design includes a parameter specification module to efficiently choose initialization pose parameter and a fine-registration module to align images. The proposed framework takes extracting multi-scale features into consideration using a novel composite connection encoder with special training techniques. We compare the method with both learning-based methods and optimization-based methods on a in-house CT/X-ray dataset as well as simulated data to further evaluate performance. Our experiments demonstrate that the method in this paper has improved the registration performance, and thereby outperforms the existing methods in terms of accuracy and running time. We also show the potential of the proposed method as an initial pose estimator. The code is available at https://github.com/m1nhengChen/SOPI Minheng Chen, Zhirun Zhang, Shuheng Gu, Youyong Kong |
ICASSP | 4 |
| 2024 | Leveraging Tumor Heterogeneity: Heterogeneous Graph Representation Learning for Cancer Survival Prediction in Whole Slide ImagesabstractSurvival prediction is a significant challenge in cancer management. Tumor micro-environment is a highly sophisticated ecosystem consisting of cancer cells, immune cells, endothelial cells, fibroblasts, nerves and extracellular matrix. The intratumor heterogeneity and the interaction across multiple tissue types profoundly impacts the prognosis. However, current methods often neglect the fact that the contribution to prognosis differs with tissue types. In this paper, we propose ProtoSurv, a novel heterogeneous graph model for WSI survival prediction. The learning process of ProtoSurv is not only driven by data but also incorporates pathological domain knowledge, including the awareness of tissue heterogeneity, the emphasis on prior knowledge of prognostic-related tissues, and the depiction of spatial interaction across multiple tissues. We validate ProtoSurv across five different cancer types from TCGA (i.e., BRCA, LGG, LUAD, COAD and PAAD), and demonstrate the superiority of our method over the state-of-the-art methods. Junxian Wu 0002, Xinyi Ke, Xiaoming Jiang, Huanwen Wu, Youyong Kong, Lizhi Shao |
NeurIPS | 5 |
| 2024 | AHMN: A multi-modal network for long MOOC videos chapter segmentation
Jiasong Wu, Youyong Kong, Huazhong Shu, Lotfi Senhadji |
Multim. Tools Appl. | 3 |
| 2024 | CSLNSpeech: Solving the extended speech separation problem with the help of Chinese sign language
Jiasong Wu, Taotao Li, Fanman Meng, Youyong Kong, Guanyu Yang 0001, Lotfi Senhadji, Huazhong Shu |
Speech Commun. | 5 |
| 2024 | STANet: Spatio-Temporal Adaptive Network and Clinical Prior Embedding Learning for 3D+T CMR SegmentationabstractThe segmentation of cardiac structure in magnetic resonance images (CMR) is paramount in diagnosing and managing cardiovascular illnesses, given its 3D+Time (3D+T) sequence. The existing deep learning methods are constrained in their ability to 3D+T CMR segmentation, due to: (1) Limited motion perception. The complexity of heart beating renders the motion perception in 3D+T CMR, including the long-range and cross-slice motions. The existing methods' local perception and slice-fixed perception directly limit the performance of 3D+T CMR perception. (2) Lack of labels. Due to the expensive labeling cost of the 3D+T CMR sequence, the labels of 3D+T CMR only contain the end-diastolic and end-systolic frames. The incomplete labeling scheme causes inefficient supervision. Hence, we propose a novel spatio-temporal adaptation network with clinical prior embedding learning (STANet) to ensure efficient spatio-temporal perception and optimization on 3D+T CMR segmentation. (1) A spatio-temporal adaptive convolution (STAC) treats the 3D+T CMR sequence as a whole for perception. The long-distance motion correlation is embedded into the structural perception by learnable weight regularization to balance long-range motion perception. The structural similarity is measured by cross-attention to adaptively correlate the cross-slice motion. (2) A clinical prior embedding learning strategy (CPE) is proposed to optimize the partially labeled 3D+T CMR segmentation dynamically by embedding clinical priors into optimization. STANet achieves outstanding performance with Dice of 0.917 and 0.94 on two public datasets (ACDC and STACOM), which indicates STANet has the potential to be incorporated into computer-aided diagnosis tools for clinical application. Xiaoming Qi, Yuting He 0001, Yaolei Qi, Youyong Kong, Guanyu Yang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Topology Uncertainty Modeling For Imbalanced Node Classification on GraphsabstractMost existing graph neural networks work under a class-balanced assumption, while ignoring class-imbalanced scenarios that widely exist in real-world graphs. Although there are many methods in other fields that can alleviate this issue, they do not consider the special topology of the non-Euclidean graph. Hence, we propose Graph Topology Uncertainty (GraphTU), a novel probabilistic class-imbalanced solution specifically for graphs. Firstly, an invisible "uncertain gap" between under-represented minorities in training set and authentic minorities in unseen set is modeled by estimating statistical variances in topology. We extend the training distribution for minorities by sampling in this gap through a non-parametric way. Moreover, a gradient-guided mask is introduced to prevent biased statistics. Extensive experiments demonstrate the superior performance of GraphTU. Jiayi Gao, Youyong Kong |
ICASSP | 4 |
| 2023 | Topgformer: Topological-Based Graph Transformer for Mapping Brain Structural Connectivity to Functional ConnectivityabstractExploring the mapping between structural connectivity (SC) and functional connectivity (FC) is of essential importance to understanding the working mechanism of the human brain. Traditional methods are difficult to represent the complex relationship of high-order interaction between SC and FC. Recent learning-based methods can not well capture the important long-range interactions and edge information of brain connectivity. To address the issue, we propose a novel Topological-based Graph Transformer (TopGFormer) to generate functional connectivity from the structure connectivity with sufficient consideration of topological properties of brain connectivity. We propose a Topological Multi-Head Attention block that simultaneously uses centrality encoding and adjacency matrix to capture node and edge importance respectively. The centrality encoding reflects the importance of brain regions while the adjacency matrix captures the linking relationship of different regions. Experiments on HCP resting and emotion task datasets demonstrate the performance of the proposed method. Dalu Guo, Youyong Kong |
ICASSP | 4 |
| 2023 | Graph Contrastive Learning with Learnable Graph AugmentationabstractGraph contrastive learning has gained popularity due to its success in self-supervised graph representation learning. Augmented views in contrastive learning greatly determine the quality of the learned representations. Handcrafted data augmentations in previous work require tedious trial-and- errors per dataset, which is time-consuming and resource-intensive. Here, we propose a Graph Contrastive learning framework with Learnable graph Augmentation called GraphCLA. Specifically, learnable graph augmentation trains augmented views by minimizing the mutual information (MI) between the input graphs and their augmented graphs. This paper designs a kernel contrastive loss function based on an end-to-end differentiable graph kernel to learn augmented views. In addition, this paper utilizes a min-max optimization strategy to learn challenging augmented graphs and to learn discriminative representations adversarially. Finally, we compared GraphCLA with state-of-the-art self-supervised learning baselines and experimentally validate the effectiveness of GraphCLA. Xinyan Pu, Huazhong Shu, Jean-Louis Coatrieux, Youyong Kong |
ICASSP | 5 |
| 2023 | Brainnetformer: Decoding Brain Cognitive States with Spatial-Temporal Cross AttentionabstractLearning about the cognitive state of the brain has always been a popular topic. Based on the fact that fluctuations of brain signals and functional connectome (FC) relate to specific human behaviors, deep learning based methods have shown promising results on the prediction of such behaviors by analyzing biological signals. Existing methods either model from static perspectives or apply spatial-temporal graph convolution to extract dynamic properties. However, the static information and dynamic information can reflect global brain activities and local brain activities respectively. Thus, we propose BrainNetFormer to incorporate both static and dynamic properties for human behavior prediction. To be specific, a spatial cross attention module and a temporal cross attention module are introduced for information fusion. In addition, since a specific behavior of subjects can be decomposed into a series of subtasks, we introduce a sub-task regularization loss to assist in training and empower the model to recognize subtasks at each moment. Experiments on the HCP-Task dataset demonstrate the superior performance of the proposed model. Leheng Sheng, Wenhan Wang, Zhiyi Shi, Jichao Zhan, Youyong Kong |
ICASSP | 5 |
| 2023 | RH-BrainFS: Regional Heterogeneous Multimodal Brain Networks Fusion StrategyabstractMultimodal fusion has become an important research technique in neuroscience that completes downstream tasks by extracting complementary information from multiple modalities. Existing multimodal research on brain networks mainly focuses on two modalities, structural connectivity (SC) and functional connectivity (FC). Recently, extensive literature has shown that the relationship between SC and FC is complex and not a simple one-to-one mapping. The coupling of structure and function at the regional level is heterogeneous. However, all previous studies have neglected the modal regional heterogeneity between SC and FC and fused their representations via "simple patterns", which are inefficient ways of multimodal fusion and affect the overall performance of the model. In this paper, to alleviate the issue of regional heterogeneity of multimodal brain networks, we propose a novel Regional Heterogeneous multimodal Brain networks Fusion Strategy (RH-BrainFS). Briefly, we introduce a brain subgraph networks module to extract regional characteristics of brain networks, and further use a new transformer-based fusion bottleneck module to alleviate the issue of regional heterogeneity between SC and FC. To the best of our knowledge, this is the first paper to explicitly state the issue of structural-functional modal regional heterogeneity and to propose a
solution. Extensive experiments demonstrate that the proposed method outperforms several state-of-the-art methods in a variety of neuroscience tasks. Hongting Ye, Yalu Zheng, Youyong Kong, Yonggui Yuan |
NeurIPS | 5 |
| 2023 | Multi-scale self-attention mixup for graph classification
Youyong Kong, Jiasong Wu |
Pattern Recognit. Lett. | 1 |
| 2023 | Multi-Connectivity Representation Learning Network for Major Depressive Disorder DiagnosisabstractThe pathophysiology of major depressive disorder (MDD) has been demonstrated to be highly associated with the dysfunctional integration of brain activity. Existing studies only fuse multi-connectivity information in a one-shot approach and ignore the temporal property of functional connectivity. A desired model should utilize the rich information in multiple connectivities to help improve the performance. In this study, we develop a multi-connectivity representation learning framework to integrate multi-connectivity topological representation from structural connectivity, functional connectivity and dynamic functional connectivities for automatic diagnosis of MDD. Briefly, structural graph, static functional graph and dynamic functional graphs are first computed from the diffusion magnetic resonance imaging (dMRI) and resting state functional magnetic resonance imaging (rsfMRI). Secondly, a novel Multi-Connectivity Representation Learning Network (MCRLN) approach is developed to integrate the multiple graphs with modules of structural-functional fusion and static-dynamic fusion. We innovatively design a Structural-Functional Fusion (SFF) module, which decouples graph convolution to capture modality-specific features and modality-shared features separately for an accurate brain region representation. To further integrate the static graphs and dynamic functional graphs, a novel Static-Dynamic Fusion (SDF) module is developed to pass the important connections from static graphs to dynamic graphs via attention values. Finally, the performance of the proposed approach is comprehensively examined with large cohorts of clinical data, which demonstrates its effectiveness in classifying MDD patients. The sound performance suggests the potential of the MCRLN approach for the clinical use in diagnosis. The code is available at https://github.com/LIST-KONG/MultiConnectivity-master. Youyong Kong, Wenhan Wang, Xiaoyun Liu, Shuwen Gao, Zhenghua Hou, Chunming Xie, Zhijun Zhang 0010, Yonggui Yuan |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Feature Space Message Passing Network for Medical Image Semantic SegmentationabstractAccurate semantic segmentation of medical images is of significant importance for subsequent processing and analysis. The encoder-decoder deep learning framework has been widely applied for numerous medical image segmentation tasks. However, most existing approaches are restricted by the limited receptive field for failing to capture longrange dependencies, meanwhile lacking global features for spatial information recovery. To solve both problems, we propose a novel feature space message passing network (FSMPN) framework. At first, a dynamic message passing block (DMPB) is proposed to perform the long-range interactions for better feature learning between voxels. Secondly, a skipped graph connection (SGC) module is developed to explicitly transfer learned graph with features from encoder stage to decoder stage to help recover spatial information. The proposed FSMPN was able to achieve superior performance on different types of medical image datasets compared to other popular models. Junxiao Sun, Shuyi Niu, Yan Zhang 0094, Youyong Kong |
ICASSP | 5 |
| 2022 | Spatio-Temporal Attention Graph Convolution Network for Functional Connectome ClassificationabstractNumerous evidence has demonstrated the pathophysiology of a number of mental disorders is intimately associated with abnormal changes of dysfunctional integration of brain network. Functional connectome (FC) exhibits a strong discriminative power for mental disorder identification. However, existing methods are insufficient for modeling both spatial correlation and temporal dynamics of FC. In this study, we propose a novel Spatio-Temporal Attention Graph Convolution Network (STAGCN) for FC classification. In spatial domain, we develop attention enhanced graph convolutional network to take advantage of brain regions’ topological features. Moreover, a novel multi-head self-attention approach is proposed to capture the temporal relationships among different dynamic FC. Extensive experiments on two tasks of mental disorder diagnosis demonstrate the superior performance of the proposed STAGCN. Wenhan Wang, Youyong Kong, Zhenghua Hou, Yonggui Yuan |
ICASSP | 2 |
| 2022 | Temporal Cross-Graph Network for Brain Functional Activity PredictionabstractPrediction of brain functional activity is of great significance for neuroscience research. The brain functional activities at different regions are highly related, and their relationships can be captured with functional connectivity and structural connectivity. The existing works are challenging to integrate two connectivity information for functional activity prediction. In this paper, we propose a Temporal Cross-Graph Network (TCGN) for predicting brain functional activity, which can comprehensively exploit multi-modal spatial dependence and temporal patterns. In particular, a novel cross-graph convolution module is developed to capture the spatial features of brain structural and functional connectivity. A temporal fusion module is designed to learn the pattern of dynamic functional connectivity to guide the prediction. Specially, a multi-task loss function is proposed to incorporate functional activity and dynamic functional connectivity. Extensive experiments on the Human Connectome Project dataset demonstrate the effectiveness of the proposed framework. Xinyu Yuan, Wenhan Wang, Youyong Kong, Jiasong Wu, Guanyu Yang 0001, Huazhong Shu |
ICASSP | 3 |
| 2022 | Iterative Seeded Region Growing for Brain Tissue SegmentationabstractBrain tissue segmentation from magnetic resonance imaging (MRI) is of significant importance for clinical application and cognitive research. The promising deep learning based methods heavily depend on the quality and quantity of training datasets, and also ignore the domain knowledge. To overcome this issue, this paper proposes a novel Iterative Seeded Region Growing (ISRG) approach for brain tissue segmentation with only one reference image. After super-voxel generation and matching, we first select the high confidence seeded regions based on the high similarity between individual brain images. Then, we obtain initial the voxel-wise tissue probabilities with a proposed fully convolutional network (named TPUNet). Thirdly, the seeded regions are updated according to the voxel-wise tissue probabilities. The second and the third steps are iteratively performed until the segmentation labels of the entire image are obtained. The proposed approach is evaluated on IBSR18 dataset and achieves better results compared with other methods. Junxiao Sun, Guanyu Yang 0001, Huazhong Shu, Youyong Kong |
ICIP | 6 |
| 2022 | Graph Attention Mixup Transformer for Graph Classification
Xinyan Pu, Youyong Kong |
ICONIP (6) | 4 |
| 2022 | GCN2CAPS: Graph Convolutional Network to Capsule Network For Wide-Field Robust Graph LearningabstractGraph Neural Networks (GNNs) have achieved remarkable performance in extracting structure-aware node representations for graph signal data. However, existing GNNs overly emphasize the consistency of neighbor nodes in the limited receptive field and severely overlook the wide-field information. In this paper, we propose a novel GCN2Caps that transfers multi-field node representations from graph convolutional network (GCN) to capsule network for wide-field graph learning. Specifically, multi-field GCN is employed to extend the receptive field and extract multi-field features. To perform multi-field interactions, GCN2Caps then explores the inherent relationships between multi-field features and generates the wide-field features through the capsule mechanism. On the basis of the wide-field features, a wide-field min-cut graph constraint is introduced to the loss function to execute complementary constraints on the original graph structure. Extensive experiments on three real-world datasets demonstrate that GCN2Caps significantly outperforms stateof-the-art baselines on semi-supervised node classification and the robustness of GCN2Caps is further validated. Shuyi Niu, Junxiao Sun, Youyong Kong, Huazhong Shu |
ICPR | 3 |
| 2022 | Hierarchical Diffusion Scattering Graph Neural NetworkabstractGraph neural network (GNN) is popular now to solve the tasks in non-Euclidean space and most of them learn deep embeddings by aggregating the neighboring nodes. However, these methods are prone to some problems such as over-smoothing because of the single-scale perspective field and the nature of low-pass filter. To address these limitations, we introduce diffusion scattering network (DSN) to exploit high-order patterns. With observing the complementary relationship between multi-layer GNN and DSN, we propose Hierarchical Diffusion Scattering Graph Neural Network (HDS-GNN) to efficiently bridge DSN and GNN layer by layer to supplement GNN with multi-scale information and band-pass signals. Our model extracts node-level scattering representations by intercepting the low-pass filtering, and adaptively tunes the different scales to regularize multi-scale information. Then we apply hierarchical representation enhancement to improve GNN with the scattering features. We benchmark our model on nine real-world networks on the transductive semi-supervised node classification task. The experimental results demonstrate the effectiveness of our method. Xinyan Pu, Jiasong Wu, Huazhong Shu, Youyong Kong |
IJCAI | 6 |
| 2022 | XMorpher: Full Transformer for Deformable Medical Image Registration via Cross Attention
Yuting He 0001, Youyong Kong, Jean-Louis Coatrieux, Huazhong Shu, Guanyu Yang 0001, Shuo Li 0001 |
MICCAI (6) | 3 |
| 2022 | Convolutional modulation theory: A bridge between convolutional neural networks and signal modulation theory
Fuzhi Wu, Jiasong Wu, Youyong Kong, Guanyu Yang 0001, Huazhong Shu, Guy Carrault, Lotfi Senhadji |
Neurocomputing | 3 |
| 2022 | Multi-Stage Graph Fusion Networks for Major Depressive Disorder DiagnosisabstractMajor depressive disorder (MDD) is a common and severe psychiatric illness marked by loss of interest and low energy, which result in the highest burden of disability among all mental disorders. Clinical MDD diagnosis still utilizes the phenomenological approach of syndrome-based interview, which leads to a high rate of misdiagnosis. Therefore, it is highly imperative to explore effective biomarkers to enable precise personalized diagnosis. There still exist two main challenges due to complexity of MDD and individual differences. On the one hand, discriminative features need to be investigated to better reflect the characteristics of MDD. On the other hand, the performance from shallow and static learning models is still not satisfactory. To overcome these issues, we propose a novel Multi-Stage Graph Fusion Networks (MSGFN) for major depressive disorder diagnosis. At first, functional connectivity is calculated to better characterize interactions between white matter and gray matter. Second, multi-stage features are obtained by a deep subspace learning model, and a number of graphs are constructed under the self-expression constraints at each stage. Finally, a novel graph convolutional fusion module is proposed with graph convolutional operations to integrate features and graph at each stage. Extensive experiments demonstrate the superior performance of the proposed framework. Our source code is available on:https://github.com/LIST-KONG/MSGFN-master. Youyong Kong, Shuyi Niu, Heren Gao, Yingying Yue, Huazhong Shu, Chunming Xie, Zhijun Zhang 0010, Yonggui Yuan |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Few-Shot Learning for Deformable Medical Image Registration With Perception-Correspondence Decoupling and Reverse TeachingabstractDeformable medical image registration estimates corresponding deformation to align the regions of interest (ROIs) of two images to a same spatial coordinate system. However, recent unsupervised registration models only have correspondence ability without perception, making misalignment on blurred anatomies and distortion on task-unconcerned backgrounds. Label-constrained (LC) registration models embed the perception ability via labels, but the lack of texture constraints in labels and the expensive labeling costs causes distortion internal ROIs and overfitted perception. We propose the first few-shot deformable medical image registration framework, Perception-Correspondence Registration (PC-Reg), which embeds perception ability to registration models only with few labels, thus greatly improving registration accuracy and reducing distortion. 1) We propose the Perception-Correspondence Decoupling which decouples the perception and correspondence actions of registration to two CNNs. Therefore, independent optimizations and feature representations are available avoiding interference of the correspondence due to the lack of texture constraints. 2) For few-shot learning, we propose Reverse Teaching which aligns labeled and unlabeled images to each other to provide supervision information to the structure and style knowledge in unlabeled images, thus generating additional training data. Therefore, these data will reversely teach our perception CNN more style and structure knowledge, improving its generalization ability. Our experiments on three datasets with only five labels demonstrate that our PC-Reg has competitive registration accuracy and effective distortion-reducing ability. Compared with LC-VoxelMorph( λ = 1), we achieve the 12.5%, 6.3% and 1.0% Reg-DSC improvements on three datasets, revealing our framework with great potential in clinical application. Yuting He 0001, Rongjun Ge, Jian Yang 0009, Youyong Kong, Huazhong Shu, Guanyu Yang 0001, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Landmark Localization for Cephalometric Analysis Using Multiscale Image Patch-Based Graph Convolutional NetworksabstractAccurate and robust cephalometric image analysis plays an essential role in orthodontic diagnosis, treatment assessment and surgical planning. This paper proposes a novel landmark localization method for cephalometric analysis using multiscale image patch-based graph convolutional networks. In detail, image patches with the same size are hierarchically sampled from the Gaussian pyramid to well preserve multiscale context information. We combine local appearance and shape information into spatialized features with an attention module to enrich node representations in graph. The spatial relationships of landmarks are built with the incorporation of three-layer graph convolutional networks, and multiple landmarks are simultaneously updated and moved toward the targets in a cascaded coarse-to-fine process. Quantitative results obtained on publicly available cephalometric X-ray images have exhibited superior performance compared with other state-of-the-art methods in terms of mean radial error and successful detection rate within various precision ranges. Our approach performs significantly better especially in the clinically accepted range of 2 mm and this makes it suitable in cephalometric analysis and orthognathic surgery. Yuanxiu Zhang, Youyong Kong, Chen Zhang 0024, Jean-Louis Coatrieux, Huazhong Shu |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Semi-Supervised Medical Image Semantic Segmentation with Multi-scale Graph Cut LossabstractMost semantic segmentation methods are based on supervised convolutional neural networks which require large amounts of labeled data. However, the acquisition of a large number of high-quality labels is time-consuming and of high annotation cost for medical images. In this paper, we propose a semi-supervised learning framework based on a novel multi-scale graph cut loss function. Firstly, the multi-scale features obtained from the segmentation network are utilized to construct the graph in non-Euclidean space. Then the long-distance information between voxels at different scales can be captured through the graph embedding module. After that, the graph cut loss is calculated according to the final latent features. Only a few labeled data is needed in our proposed method, which is of significance in the practical clinic. The experiments on the BrainWeb20 dataset and the IBSR18 dataset demonstrate the effectiveness of the proposed method compared to the well-known state-of-the-art methods. Junxiao Sun, Yan Zhang 0094, Jiasong Wu, Youyong Kong |
ICIP | 5 |
| 2021 | CPNet: Cycle Prototype Network for Weakly-Supervised 3D Renal Compartments Segmentation on CT Images
Song Wang 0002, Yuting He 0001, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001 |
MICCAI (2) | 3 |
| 2021 | GSCFN: A graph self-construction and fusion network for semi-supervised brain tissue segmentation in MRI
Yan Zhang 0094, Youyong Kong, Jiasong Wu, Jian Yang 0009, Huazhong Shu, Gouenou Coatrieux |
Neurocomputing | 3 |
| 2021 | Meta grayscale adaptive network for 3D integrated renal structures segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Rongjun Ge, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2021 | Examinee-Examiner Network: Weakly Supervised Accurate Coronary Lumen Segmentation Using Centerline ConstraintabstractAccurate coronary lumen segmentation on coronary-computed tomography angiography (CCTA) images is crucial for quantification of coronary stenosis and the subsequent computation of fractional flow reserve. Many factors including difficulty in labeling coronary lumens, various morphologies in stenotic lesions, thin structures and small volume ratio with respect to the imaging field complicate the task. In this work, we fused the continuity topological information of centerlines which are easily accessible, and proposed a novel weakly supervised model, Examinee-Examiner Network (EE-Net), to overcome the challenges in automatic coronary lumen segmentation. First, the EE-Net was proposed to address the fracture in segmentation caused by stenoses by combining the semantic features of lumens and the geometric constraints of continuous topology obtained from the centerlines. Then, a Centerline Gaussian Mask Module was proposed to deal with the insensitiveness of the network to the centerlines. Subsequently, a weakly supervised learning strategy, Examinee-Examiner Learning, was proposed to handle the weakly supervised situation with few lumen labels by using our EE-Net to guide and constrain the segmentation with customized prior conditions. Finally, a general network layer, Drop Output Layer, was proposed to adapt to the class imbalance by dropping well-segmented regions and weights the classes dynamically. Extensive experiments on two different data sets demonstrated that our EE-Net has good continuity and generalization ability on coronary lumen segmentation task compared with several widely used CNNs such as 3D-UNet. The results revealed our EE-Net with great potential for achieving accurate coronary lumen segmentation in patients with coronary artery disease. Code at http://github.com/qiyaolei/Examinee-Examiner-Network. Yaolei Qi, Yuting He 0001, Zehang Li, Youyong Kong, Jean-Louis Coatrieux, Huazhong Shu, Guanyu Yang 0001, Shengxian Tu |
IEEE Trans. Image Process. | 6 |
| 2020 | Deep Complementary Joint Model for Complex Scene Registration and Few-Shot Segmentation on Medical Images
Yuting He 0001, Guanyu Yang 0001, Youyong Kong, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux, Jean-Louis Dillenseger, Shuo Li 0001 |
ECCV (18) | 4 |
| 2020 | Deep octonion networks
Jiasong Wu, Fuzhi Wu, Youyong Kong, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 4 |
| 2020 | Compressed sensing MR image reconstruction via a deep frequency-division network
Jiulou Zhang, Yunbo Gu, Youyong Kong, Yang Chen 0008, Huazhong Shu, Jean-Louis Coatrieux |
Neurocomputing | 5 |
| 2020 | Dense biased networks with deep priori anatomy and hard region adaptation: Semi-supervised learning for fine renal artery segmentation
Yuting He 0001, Guanyu Yang 0001, Jian Yang 0009, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
Medical Image Anal. | 5 |
| 2020 | Learning Deep Landmarks for Imbalanced ClassificationabstractWe introduce a deep imbalanced learning framework called learning DEep Landmarks in laTent spAce (DELTA). Our work is inspired by the shallow imbalanced learning approaches to rebalance imbalanced samples before feeding them to train a discriminative classifier. Our DELTA advances existing works by introducing the new concept of rebalancing samples in a deeply transformed latent space, where latent points exhibit several desired properties including compactness and separability. In general, DELTA simultaneously conducts feature learning, sample rebalancing, and discriminative learning in a joint, end-to-end framework. The framework is readily integrated with other sophisticated learning concepts including latent points oversampling and ensemble learning. More importantly, DELTA offers the possibility to conduct imbalanced learning with the assistancy of structured feature extractor. We verify the effectiveness of DELTA not only on several benchmark data sets but also on more challenging real-world tasks including click-through-rate (CTR) prediction, multi-class cell type classification, and sentiment analysis with sequential inputs. Feng Bao 0002, Yue Deng 0001, Youyong Kong, Zhiquan Ren, Jin-Li Suo, Qionghai Dai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Unsupervised Three-Dimensional Image Registration Using a Cycle Convolutional Neural NetworkabstractIn this paper, an unsupervised cycle image registration convolutional neural network named CIRNet is developed for 3D medical image registration. Different from most deep learning based registration methods that require known spatial transforms, our proposed method is trained in an unsupervised way and predicts the dense displacement vector field. The CIRNet is composed by two image registration modules which have the same architecture and share the parameters. A cycle identical loss is designed in the CIRNet to provide additional constraints to ensure the accuracy of the predicted dense displacement vector field. The method is evaluated by the registration in 4D (3D+t) cardiac CT and MRI images respectively. Quantitative evaluation results demonstrate that our method performs better than the other two existing image registration algorithms. Especially, compared to the traditional image registration methods, our proposed network can finish 3D image registration in less than one second. Ziwei Lu, Jean-Louis Coatrieux, Guanyu Yang 0001, Tiancong Hua, Liyu Hu, Youyong Kong, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Huazhong Shu |
ICIP | 6 |
| 2019 | A Multi-Task Convolutional Neural Network for Renal Tumor Segmentation and Classification Using Multi-Phasic CT ImagesabstractAccounting for nearly 2% of all adults, renal cell carcinomas are sensitive to laparoscopic partial nephrectomy (LPN) which needs an accurate diagnosis and localization before operation. Faced with various intensity distribution, erratic location, irregular shape, etc, the image classification and semantic segmentation on CT scans of renal tumor are challenges. This paper presents a multi-task network, segmentation and classification convolutional neural network (SCNet), for preoperative assessment of renal tumor. Via the combination of two tasks, semantic features are fed to the classification network and classification results give segmentation network feedbacks in return. Besides, a 2-step segmentation strategy is conducted to the segmentation module which improves the result by 2.8%. Our experimental results of classification and segmentation achieve 100% accuracy and 0.882 dice coefficient of tumor region respectively, which are better than the results of a single classification network and segmentation network. Tan Pan, Huazhong Shu, Jean-Louis Coatrieux, Guanyu Yang 0001, Chuanxia Wang, Ziwei Lu, Zhongwen Zhou, Youyong Kong, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger |
ICIP | 8 |
| 2019 | Brain Tissue Segmentation based on Graph Convolutional NetworksabstractIn neuroscience research, brain tissue segmentation from magnetic resonance imaging is of significant importance. A challenging issue is to provide an accurate segmentation due to the tissue heterogeneity, which is caused by noise, bias filed and partial volume effects. To overcome these problems, we propose a novel brain MRI segmentation algorithm, the originality of which stands on the combination of supervoxels with graph convolutional networks. Supervoxels are generated from the 3D MRI image with the help of an improved simple linear iterative clustering algorithm. A graph is then built from these supervoxels through the K nearest neighbor algorithm, before being sent to GCNs for tissues classification. The proposed method is evaluated on the two common datasets- the BrainWeb18 dataset and the Internet Brain Segmentation Repository 18 dataset. Experiments demonstrate the performance of our method and that it is better than well-known state-of-the-art methods such as FMRIB software library, statistical parametric mapping, adaptive graph filter. Yan Zhang 0094, Youyong Kong, Jiasong Wu, Gouenou Coatrieux, Huazhong Shu |
ICIP | 2 |
| 2019 | DPA-DenseBiasNet: Semi-supervised 3D Fine Renal Artery Segmentation with Dense Biased Network and Deep Priori Anatomy
Yuting He 0001, Guanyu Yang 0001, Yang Chen 0008, Youyong Kong, Jiasong Wu, Lijun Tang, Xiaomei Zhu, Jean-Louis Dillenseger, Shaobo Zhang 0008, Huazhong Shu, Jean-Louis Coatrieux, Shuo Li 0001 |
MICCAI (6) | 4 |
| 2018 | Multi-Session Parcellation of the Human Brain Using Resting-State fMRIabstractThe parcellation of human brain is important to discover the neural mechanisms behind human behavior. Despite many statistical models are proposed to compute parcellation using resting-state functional magnetic resonance imaging (rs-fMRI), there still remains challenges for obtaining more reliable parcellations for individuals. To address these challenges, we design a multi-session parcellation approach based on data of several sessions for each subject. Initially, we cluster original brain data into a certain number of homogeneous supervoxels to reduce computational cost of the subsequent stages. Secondly, we use spectral clustering method to cluster our supervoxels into several ROIs to obtain a reasonable parcellation of each subject. Thirdly, we propose a method based on adjacency matrix to merge parcellations from different sessions to obtain a more reproductive parcellation of each subject. The performance of the proposed algorithm is evaluated with two commonly used metrics, including silhouette width and Dice's coefficient. The experiments demonstrate the effectiveness of the proposed approach with a better reproducibility. The proposed algorithm has high potential to generate reliable human brain parcellations for analyzing individual brain network more reliably. Renhao Lei, Junxiao Sun, Youyong Kong |
CSCWD | 4 |
| 2018 | Automatic Segmentation of Kidney and Renal Tumor in CT Images Based on 3D Fully Convolutional Neural Network with Pyramid Pooling ModuleabstractRenal cancer is one of ten most common cancers in human beings. The laparoscopic partial nephrectomy (LPN) becomes the main therapeutic approach in treating renal cancer. Accurate kidney and tumor segmentation in CT images is a prerequisite step in the surgery planning. However, automatic and accurate kidney and renal tumor segmentation in CT images remains a challenge. In this paper, we propose a new method to perform a precise segmentation of kidney and renal tumor in CT angiography images. This method relies on a three-dimensional (3D) fully convolutional network (FCN) which combines a pyramid pooling module (PPM). The proposed network is implemented as an end-to-end learning system directly on 3D volumetric images. It can make use of the 3D spatial contextual information to improve the segmentation of the kidney as well as the tumor lesion. The experiments conducted on 140 patients show that these target structures can be segmented with a high accuracy. The resulting average dice coefficients obtained for kidney and renal tumor are equal to 0.931 and 0.802 respectively. These values are higher than those obtained from the other two neural networks. Guanyu Yang 0001, Tan Pan, Youyong Kong, Jiasong Wu, Huazhong Shu, Limin Luo 0001, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Lijun Tang, Xiaomei Zhu |
ICPR | 4 |
| 2018 | PCANet: An energy perspective
Jiasong Wu, Shijie Qiu, Youyong Kong, Longyu Jiang, Yang Chen 0008, Wankou Yang, Lotfi Senhadji, Huazhong Shu |
Neurocomputing | 3 |
| 2017 | MomentsNet: A simple learning-free method for binary image recognitionabstractIn this paper, we propose a new simple and learning-free deep learning network named MomentsNet, whose convolution layer, nonlinear processing layer and pooling layer are constructed by Moments kernels, binary hashing and block-wise histogram, respectively. Twelve typical moments (including geometrical moment, Zernike moment, Tchebichef moment, etc.) are used to construct the MomentsNet whose recognition performance for binary image is studied. The results reveal that MomentsNet has better recognition performance than its corresponding moments in almost all cases and ZernikeNet achieves the best recognition performance among MomentsNet constructed by twelve moments. ZernikeNet also shows better recognition performance on a binary image database than that of PCANet, which is a learning-based deep learning network. Jiasong Wu, Shijie Qiu, Youyong Kong, Yang Chen 0008, Lotfi Senhadji, Huazhong Shu |
ICIP | 3 |
| 2017 | Discriminant Kernel Assignment for Image CodingabstractThis paper proposes discriminant kernel assignment (DKA) in the bag-of-features framework for image representation. DKA slightly modifies existing kernel assignment to learn width-variant Gaussian kernel functions to perform discriminant local feature assignment. When directly applying gradient-descent method to solve DKA, the optimization may contain multiple time-consuming reassignment implementations in iterations. Accordingly, we introduce a more practical way to locally linearize the DKA objective and the difficult task is cast as a sequence of easier ones. Since DKA only focuses on the feature assignment part, it seamlessly collaborates with other discriminative learning approaches, e.g., discriminant dictionary learning or multiple kernel learning, for even better performances. Experimental evaluations on multiple benchmark datasets verify that DKA outperforms other image assignment approaches and exhibits significant efficiency in feature coding. Yue Deng 0001, Yanyu Zhao, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai |
IEEE Trans. Cybern. | 4 |
| 2017 | A Hierarchical Fused Fuzzy Deep Neural Network for Data ClassificationabstractDeep learning (DL) is an emerging and powerful paradigm that allows large-scale task-driven feature learning from big data. However, typical DL is a fully deterministic model that sheds no light on data uncertainty reductions. In this paper, we show how to introduce the concepts of fuzzy learning into DL to overcome the shortcomings of fixed representation. The bulk of the proposed fuzzy system is a hierarchical deep neural network that derives information from both fuzzy and neural representations. Then, the knowledge learnt from these two respective views are fused altogether forming the final data representation to be classified. The effectiveness of the model is verified on three practical tasks of image categorization, high-frequency financial data prediction and brain MRI segmentation that all contain high level of uncertainties in the raw data. The fuzzy dDL paradigm greatly outperforms other nonfuzzy and shallow learning approaches on these tasks. Yue Deng 0001, Zhiquan Ren, Youyong Kong, Feng Bao 0002, Qionghai Dai |
IEEE Trans. Fuzzy Syst. | 3 |
| 2017 | Deep Direct Reinforcement Learning for Financial Signal Representation and TradingabstractCan we train the computer to beat experienced traders for financial assert trading? In this paper, we try to address this challenge by introducing a recurrent deep neural network (NN) for real-time financial signal representation and trading. Our model is inspired by two biological-related learning concepts of deep learning (DL) and reinforcement learning (RL). In the framework, the DL part automatically senses the dynamic market condition for informative feature learning. Then, the RL module interacts with deep representations and makes trading decisions to accumulate the ultimate rewards in an unknown environment. The learning system is implemented in a complex NN that exhibits both the deep and recurrent structures. Hence, we propose a task-aware backpropagation through time method to cope with the gradient vanishing issue in deep training. The robustness of the neural system is verified on both the stock and the commodity future markets under broad testing conditions. Yue Deng 0001, Feng Bao 0002, Youyong Kong, Zhiquan Ren, Qionghai Dai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Deep and Structured Robust Information Theoretic Learning for Image AnalysisabstractThis paper presents a robust information theoretic (RIT) model to reduce the uncertainties, i.e., missing and noisy labels, in general discriminative data representation tasks. The fundamental pursuit of our model is to simultaneously learn a transformation function and a discriminative classifier that maximize the mutual information of data and their labels in the latent space. In this general paradigm, we, respectively, discuss three types of the RIT implementations with linear subspace embedding, deep transformation, and structured sparse learning. In practice, the RIT and deep RIT are exploited to solve the image categorization task whose performances will be verified on various benchmark data sets. The structured sparse RIT is further applied to a medical image analysis task for brain magnetic resonance image segmentation that allows group-level feature selections on the brain tissues. Yue Deng 0001, Feng Bao 0002, XueSong Deng, Ruiping Wang 0001, Youyong Kong, Qionghai Dai |
IEEE Trans. Image Process. | 5 |
| 2015 | Discriminative Clustering and Feature Selection for Brain MRI SegmentationabstractAutomatic segmentation of brain tissues from MRI is of great importance for clinical application and scientific research. Recent advancements in supervoxel-level analysis enable robust segmentation of brain tissues by exploring the inherent information among multiple features extracted on the supervoxels. Within this prevalent framework, the difficulties still remain in clustering uncertainties imposed by the heterogeneity of tissues and the redundancy of the MRI features. To cope with the aforementioned two challenges, we propose a robust discriminative segmentation method from the view of information theoretic learning. The prominent goal of the method is to simultaneously select the informative feature and to reduce the uncertainties of supervoxel assignment for discriminative brain tissue segmentation. Experiments on two brain MRI datasets verified the effectiveness and efficiency of the proposed approach. Youyong Kong, Yue Deng 0001, Qionghai Dai |
IEEE Signal Process. Lett. | 1 |
| 2015 | Sparse Coding-Inspired Optimal Trading System for HFT IndustryabstractThe financial industry has witnessed an exceptionally fast progress of incorporating information processing techniques in designing knowledge-based automated systems for high-frequency trading (HFT). This paper proposes a sparse coding-inspired optimal trading (SCOT) system for real-time high-frequency financial signal representation and trading. Mathematically, SCOT simultaneously learns the dictionary, sparse features, and the trading strategy in a joint optimization, yielding optimal feature representations for the specific trading objective. The learning process is modeled as a bilevel optimization and solved by the online gradient descend method with fast convergence. In this dynamic context, the system is tested on the real financial market to trade the index futures in the Shanghai exchange center. Yue Deng 0001, Youyong Kong, Feng Bao 0002, Qionghai Dai |
IEEE Trans. Ind. Informatics | 2 |