Suixue Wang

dblp:195/2085 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0009-0001-1216-2299ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 SSL-CST: Cell Segmentation for Single-Cell Spatial Transcriptome Based on Self-Supervised Learning
abstract
The continuous advancements in life science technology have enabled spatial transcriptome technology to achieve an impressive level of resolution at the single-cell level. This technology has emerged as a crucial method for studying the cellular composition and differentiation states of tissues, investigating cell-cell interactions, and unraveling the molecular mechanisms underlying diseases and developmental processes. A key component in this analysis is the accurate segmentation of cells. However, existing segmentation methods often fail to fully leverage the valuable information provided by spatial transcriptomics, leading to inaccurate cell segmentation. In this study, we introduce SSL-CST, a cell segmentation for single-cell spatial transcriptome method based on self-supervised learning. SSL-CST employs a pre-trained model for foundational contour segmentation. Following the denoising process, it utilizes a self-supervised neural network to correct the cell boundaries to obtain accurate cell boundaries. Through this approach, SSL-CST outperforms other state-of-the-art methods in various tests conducted on multiple datasets. The improved segmentation provided by SSL-CST further enhances the analysis of single-cell spatial expression, providing effective tools for biological discovery.
Weiliang Huo, Suixue Wang, Qingchen Zhang 0001
AAAI3
2026 SAMGTD: Spatial-Aware Masked Graph Transformer-Diffusion Model for Enhanced Cell Type Deconvolution in Spatial Transcriptomics
abstract
Recent advances in spatial transcriptomics have enabled the integration of gene expression profiles with precise spatial coordinates, which have facilitated the exploration of tumor occurrence and development mechanisms, as well as the development of more effective targeted and immunotherapy approaches for tumor treatment. Deciphering cell type represents a critical challenge in spatial transcriptomics research. Existing methods are limited by the pervasive “dropout” events in spatial transcriptomics, hindering their ability to fully capture the relationship between spatial location and gene expression, thereby compromising the performance of cell type deconvolution. To address these limitations, we propose a spatial-aware masked graph transformer-diffusion model (SAMGTD) for enhanced cell type deconvolution in spatial transcriptomics. For spatial transcriptomics, the masked graph transformer model is designed to adaptively capture complex dependencies between spatial locations and gene expression. It employs a masking strategy that guides the model to focus on important local information during training, while the multi-head attention mechanism captures global context. More importantly, the spatial diffusion model is constructed to achieve the dual enhancement of spatial transcriptomics, including denoising and data imputation. It incorporates the multi-head attention mechanism and residual blocks, effectively addressing the “dropout” issue commonly encountered in spatial transcriptomics. For scRNA-seq, we construct a variational autoencoder to reduce noise interference while preserving key gene expression information. Finally, we construct a spatial-aware contrastive learning model to integrate scRNA-seq and spatial transcriptomics for cell type deconvolution. Experiments conducted on three datasets demonstrate that SAMGTD outperforms baseline methods.
Suixue Wang, Qingchen Zhang 0001, Xiulong Liu 0001
AAAI2
2025 An End-to-End Deep Reinforcement Learning Framework for 3D Reconstruction of Spatial Transcriptomics
abstract
Reconstructing three-dimensional (3D) structures from consecutive tissue slices is essential to uncover tissue architecture and intercellular interactions. However, some of the current 3D reconstruction methods rely on two-stage pipelines and precise point-to-point registration, and they are overly reliant on expert knowledge. Others assume globally consistent spatial omics patterns, and they fail to account for spatial heterogeneity. To overcome the above limitations, we propose STAlignDRL, a deep reinforcement learning framework for spatial transcriptomics 3D reconstruction which can achieve end-to-end consecutive slice alignment. In detail, the proposed algorithm works in three steps. First, we perform data preprocessing to obtain the features of the slices and extract similar local contours and action coefficients between the slices. Next, we formulate the slice alignment operation of 3D reconstruction as a sequential decision-making process. Finally, we construct an end-to-end fully connected deep reinforcement learning network to achieve high-precision slice alignment. Notably, we are the first to propose a reinforcement learning-based method in the context of spatial transcriptomics 3D reconstruction, providing new insights into the 3D reconstruction of consecutive spatial transcriptomics slices. To validate our method, we conduct experiments on the Breast Cancer, DLPFC 3, and Mouse Hippocampus datasets, and compare it with the PASTE, STAligner, ST-GEARS, STitch3D, and SANTO methods on the Mouse Hippocampus dataset. The experimental results show that STAlign-DRL outperforms several recent state-of-the-art methods.
Huaiji Wang, Suixue Wang, Weiliang Huo, Xiangjun Hu 0001, Xiangfei Zhang, Qingchen Zhang 0001
BIBM2
2025 Higher-order Logical Knowledge Representation Learning
abstract
Real-world knowledge graphs abound with higher-order logical relations that simple triples, limited to pairwise connections, fail to represent. Thus, capturing higher-order logical relations involving multiple entities has garnered significant attention. However, existing methods ignore the structural information in higher-order relations. To this end, we propose a higher-order logical knowledge representation learning method, named LORE, which leverages network motifs, the patterns/subgraphs that naturally capture the structural information in graphs, to extract higher-order features and ultimately, learn effective representations of knowledge graphs. Compared to existing approaches, LORE aggregates the attribute features of entities with the extracted higher-order logical relations to form enhanced representations of knowledge graphs. In particular, three aggregators (i.e., Hadamard, Connection, and Summation) are proposed and employed. Extensive experiments have been conducted on six real-world datasets for two downstream tasks (i.e., entity classification and link prediction). The results show that LORE outperforms baselines significantly and consistently.
Suixue Wang, Weiliang Huo, Qingchen Zhang 0001
IJCAI1
2025 POMP: Pathology-omics Multimodal Pre-training Framework for Cancer Survival Prediction
abstract
Cancer survival prediction is an important direction in precision medicine, aiming to help clinicians tailor treatment regimens for patients. With the rapid development of high-throughput sequencing and computational pathology technologies, survival prediction has shifted from clinical features to joint modeling of multi-omics data and pathology images. However, existing multimodal learning methods struggle to effectively learn pathology-omics interactions due to the lack of proper alignment of multimodal data before fusion. In this paper, we propose POMP, a pathology-omics multimodal pre-training framework jointly learned with three training tasks for integrating pathological images and omics data for cancer survival prediction. To better perform cross-modal learning, we introduce a pathology-omics contrastive learning method to align the pathology and omics information. POMP leverages the principle of pre-trained models and explores the benefit of aligning multimodal information from the same patient, achieving state-of-the-art results on six cancer datasets from the Cancer Genome Atlas (TCGA). We also show that our contrastive learning method allows us to exploit the cosine similarity of pathological images and omics data as the survival risk score, which can further boost prediction performance compared with other commonly used methods. The code is available at https://github.com/SuixueWang/POMP.
Suixue Wang, Huiyuan Lai, Weiliang Huo, Qingchen Zhang 0001
IJCAI1
2025 MASTER: A Multi-granularity Invariant Structure Clustering Scheme for Multi-view Clustering
abstract
Deep multi-view clustering has attracted increasing attention in the pattern mining of data. However, most of them perform self-learning mechanisms in a single space, ignoring the fruitful structural information hidden in different-level feature spaces. Meanwhile, they conduct the reconstruction constraint to learn generalized representations of samples, failing to explore the discriminative ability of complementary and consistent information. To address the challenges, a multi-granularity invariant structure clustering scheme (MASTER) is proposed to define a bottom-up process that extracts multi-level information in sample, neighborhood, and category granularities from low-level, high-level, and semantics feature space, respectively. Specifically, it leverages the self-learning reconstruction with information-theoretic overclustering to capture invariant sample structure in the low-level feature space. Then, it models data diffusion of the clustering process in the reliable neighborhood to capture invariant local structure in the high-level feature space. Meanwhile, it defines dual divergences induced by the space geometry to capture invariant global structure in the semantics space. Finally, extensive experiments on 8 real-world datasets show that MASTER achieves state-of-the-art performance compared to 11 baselines.
Suixue Wang, Qingchen Zhang 0001, Peng Li 0027, Weiliang Huo
IJCAI1
2025 CSF-GAN: Cross-modal Semantic Fusion-based Generative Adversarial Network for Text-guided Image Inpainting
abstract
Most visual-guided image inpainting methods based on generative adversarial networks (GANs) struggle when the missing region has weak correlations with the surrounding visual context. Recently, diffusion-based methods guided by textual context have been proposed to address this limitation by leveraging additional semantic information to restore corrupted objects. However, these models typically involve more parameters and exhibit slower generation speeds compared to GAN-based approaches. To address this problem, we propose a novel text-guided image inpainting model, the cross-modal semantic fusion generative adversarial network (CSF-GAN). CSF-GAN is designed as a one-stage GAN with the following key contributions. First, a novel semantic fusion module (SFM) is introduced to integrate sentence- and word-level textual context into the inpainting process, enabling more effective guidance from multi-granularity semantic information. Second, a newly designed word-level local discriminator provides detailed feedback to the generator, enhancing the accuracy of generated content in alignment with word-level semantics. Third, two loss functions, the inpainting loss and edge loss, are employed to enhance both structural coherence and textural realism in the generated results. Extensive experiments on two benchmark datasets demonstrate that CSF-GAN outperforms state-of-the-art methods.
Suixue Wang, Qingchen Zhang 0001, Liang Zhao 0005, Weiliang Huo, Sijia Hou, Chunjiang Fu
IJCAI2
2024 ContraMAE: Contrastive alignment masked autoencoder framework for cancer survival prediction
abstract
With the rapid advancement in multimodal fusion technology, the integration of pathological images with genomics data has achieved promising results in cancer survival prediction. However, most existing multimodal models are not pre-trained by combining pathology and genomics modalities, ignoring the inherent task-agnostic associations between different modalities. While some self-supervised methods align multimodal information through pre-training objectives such as correlation and mean square error, they lack in-depth multimodal interaction. To address these issues, we propose ContraMAE, a contrastive alignment masked autoencoder framework, to fuse pathological images and genomics data for cancer survival prediction. Concretely, we introduce a contrastive objective to align multimodality and construct their intrinsic consistency. Besides, we design two reconstruction objectives to capture the complex relationships between multi-modalities by mutually compensating for the information that each side lacks. In survival prediction, the pathology and genomics encodings from the ContraMAE encoder are concatenated as the final representation to generate a survival risk score. Experimental results demonstrate that ContraMAE outperforms existing state-of-the-art methods on five cancer datasets sourced from The Cancer Genome Atlas (TCGA). The code is available at https://github.com/SuixueWang/ContraMAE.
Suixue Wang, Huiyuan Lai, Qingchen Zhang 0001
BIBM1
2023 MLLCD: A Meta Learning-based Method for Lung Cancer Diagnosis Using Histopathology Images
abstract
Lung cancer is a leading cause of death. An accurate early lung cancer diagnosis can improve a patient’s survival chances. Histopathological images are essential for cancer diagnosis. With the development of deep learning in the past decade, many scholars have used deep learning to learn the features of histopathological images and achieve lung cancer classification. However, deep learning requires a large quantity of annotated data to train the model to achieve a good classification effect, and collecting many annotated pathological images is time-consuming and expensive. Faced with the scarcity of pathological data, we present a meta-learning method for lung cancer diagnosis (called MLLCD). In detail, the MLLCD works in three steps. First, we preprocess all data using the bilinear interpolation method and then design the base learner which units a convolutional neural network(CNN) and transformer to distill local features and global features of pathology images with different resolutions. Finally, we train and update the base learner with a model-agnostic meta-learning (MAML) algorithm. Clinical Proteomic Tumor Analysis Consortium (CPTAC) cancer patient data demonstrate that our proposed model achieves the receiver operating characteristic (ROC) values of 0.94 for lung cancer diagnosis.
Xiangjun Hu 0001, Suixue Wang, Hang Li 0006, Qingchen Zhang 0001
BIBM2
2023 HC-MAE: Hierarchical Cross-attention Masked Autoencoder Integrating Histopathological Images and Multi-omics for Cancer Survival Prediction
abstract
Accurate cancer survival prediction enables clinicians to tailor treatment regimens based on individual patient prognoses, effectively mitigating over-treatment and inefficient medical resource allocation. Recently, the integration of histopathological images and multi-omics data, together with deep learning, has become increasingly applied to predict cancer survival. However, current deep learning-based integration methods ignore the spatial relationships across various fields of view within gigapixel histopathological images, since they mainly focus on a specific field of view. Inspired by the hierarchical image pyramid transformer (HIPT), we propose a hierarchical cross-attention masked autoencoder (HC-MAE) to integrate histopathological images and multi-omics data for cancer survival prediction. Specifically, HC-MAE aggregates the representations learned from different fields of view, effectively capturing the fine-grained details and the spatial relationships within histopathological images. We conduct experiments to compare the HC-MAE method with current state-of-the-art methods on six cancer datasets sourced from The Cancer Genome Atlas (TCGA). The experimental results demonstrate that HC-MAE achieves superior performance on five out of six cancer datasets, significantly outperforming the compared methods. The code is available at https://github.com/SuixueWang/HC-MAE.
Suixue Wang, Xiangjun Hu 0001, Qingchen Zhang 0001
BIBM1