VLDB 2026 Research / reviewers in the wild / expert
Weiliang Huo
dblp:279/9417
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0002-3977-6538ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSL-CST: Cell Segmentation for Single-Cell Spatial Transcriptome Based on Self-Supervised LearningabstractThe continuous advancements in life science technology have enabled spatial transcriptome technology to achieve an impressive level of resolution at the single-cell level. This technology has emerged as a crucial method for studying the cellular composition and differentiation states of tissues, investigating cell-cell interactions, and unraveling the molecular mechanisms underlying diseases and developmental processes. A key component in this analysis is the accurate segmentation of cells. However, existing segmentation methods often fail to fully leverage the valuable information provided by spatial transcriptomics, leading to inaccurate cell segmentation. In this study, we introduce SSL-CST, a cell segmentation for single-cell spatial transcriptome method based on self-supervised learning. SSL-CST employs a pre-trained model for foundational contour segmentation. Following the denoising process, it utilizes a self-supervised neural network to correct the cell boundaries to obtain accurate cell boundaries. Through this approach, SSL-CST outperforms other state-of-the-art methods in various tests conducted on multiple datasets. The improved segmentation provided by SSL-CST further enhances the analysis of single-cell spatial expression, providing effective tools for biological discovery. Weiliang Huo, Suixue Wang, Qingchen Zhang 0001 |
AAAI | 1 |
| 2026 | GATCL: An Adaptive Contrastive Learning Framework Based on MHGAT for Spatial Domain Identification in Spatial TranscriptomicsabstractRecent advances in spatial transcriptomics have enabled the simultaneous measurement of gene expression profiles and spatial location information, offering a more comprehensive and in-depth view for studying the tissue microenvironment. Spatial domain identification is a crucial step in analyzing spatial transcriptomics. However, current methods have poor accuracy and visualization because they lack self-adaptability to different tissue data, and moreover, they cannot effectively extract spatial location information. To address these issues, we propose an adaptive graph contrastive learning framework based on multi-head graph attention networks (GATCL) for spatial domain identification. Specifically, we design a data augmentation module to mask and shuffle the pre-processed gene expression data to generate more differentiated negative samples. In addition, we construct the multi-head graph attention networks (MHGAT) to encode gene expression profiles and spatial location information. More importantly, we design an adaptive graph contrastive learning model that works both with positive and negative samples from spatial transcriptomics. We introduce the attention pooling mechanism to dynamically and adaptively aggregate the spots' neighborhood information, and to improve the model's generalization ability for different spatial transcriptomics data. Furthermore, we design a discriminator that adds spectral normalization to bilinear functions. Experimental results on DLPFC, breast cancer, and mouse somatosensory cortex datasets demonstrate that the average Adjusted Rand Index (ARI) scores are 0.5746, 0.6182, and 0.6496, respectively, significantly outperforming baseline methods. More importantly, GATCL provides a more detailed visualization of different spatial transcriptomics data. Weiliang Huo, Qingchen Zhang 0001, Xiulong Liu 0001 |
AAAI | 2 |
| 2026 | Cell interactions inference for single-cell spatial transcriptomes with GraphCIM
Weiliang Huo, Qingchen Zhang 0001 |
Expert Syst. Appl. | 1 |
| 2025 | An End-to-End Deep Reinforcement Learning Framework for 3D Reconstruction of Spatial TranscriptomicsabstractReconstructing three-dimensional (3D) structures from consecutive tissue slices is essential to uncover tissue architecture and intercellular interactions. However, some of the current 3D reconstruction methods rely on two-stage pipelines and precise point-to-point registration, and they are overly reliant on expert knowledge. Others assume globally consistent spatial omics patterns, and they fail to account for spatial heterogeneity. To overcome the above limitations, we propose STAlignDRL, a deep reinforcement learning framework for spatial transcriptomics 3D reconstruction which can achieve end-to-end consecutive slice alignment. In detail, the proposed algorithm works in three steps. First, we perform data preprocessing to obtain the features of the slices and extract similar local contours and action coefficients between the slices. Next, we formulate the slice alignment operation of 3D reconstruction as a sequential decision-making process. Finally, we construct an end-to-end fully connected deep reinforcement learning network to achieve high-precision slice alignment. Notably, we are the first to propose a reinforcement learning-based method in the context of spatial transcriptomics 3D reconstruction, providing new insights into the 3D reconstruction of consecutive spatial transcriptomics slices. To validate our method, we conduct experiments on the Breast Cancer, DLPFC 3, and Mouse Hippocampus datasets, and compare it with the PASTE, STAligner, ST-GEARS, STitch3D, and SANTO methods on the Mouse Hippocampus dataset. The experimental results show that STAlign-DRL outperforms several recent state-of-the-art methods. Huaiji Wang, Suixue Wang, Weiliang Huo, Xiangjun Hu 0001, Xiangfei Zhang, Qingchen Zhang 0001 |
BIBM | 4 |
| 2025 | Higher-order Logical Knowledge Representation LearningabstractReal-world knowledge graphs abound with higher-order logical relations that simple triples, limited to pairwise connections, fail to represent. Thus, capturing higher-order logical relations involving multiple entities has garnered significant attention. However, existing methods ignore the structural information in higher-order relations. To this end, we propose a higher-order logical knowledge representation learning method, named LORE, which leverages network motifs, the patterns/subgraphs that naturally capture the structural information in graphs, to extract higher-order features and ultimately, learn effective representations of knowledge graphs. Compared to existing approaches, LORE aggregates the attribute features of entities with the extracted higher-order logical relations to form enhanced representations of knowledge graphs. In particular, three aggregators (i.e., Hadamard, Connection, and Summation) are proposed and employed. Extensive experiments have been conducted on six real-world datasets for two downstream tasks (i.e., entity classification and link prediction). The results show that LORE outperforms baselines significantly and consistently. Suixue Wang, Weiliang Huo, Qingchen Zhang 0001 |
IJCAI | 2 |
| 2025 | POMP: Pathology-omics Multimodal Pre-training Framework for Cancer Survival PredictionabstractCancer survival prediction is an important direction in precision medicine, aiming to help clinicians tailor treatment regimens for patients. With the rapid development of high-throughput sequencing and computational pathology technologies, survival prediction has shifted from clinical features to joint modeling of multi-omics data and pathology images. However, existing multimodal learning methods struggle to effectively learn pathology-omics interactions due to the lack of proper alignment of multimodal data before fusion. In this paper, we propose POMP, a pathology-omics multimodal pre-training framework jointly learned with three training tasks for integrating pathological images and omics data for cancer survival prediction. To better perform cross-modal learning, we introduce a pathology-omics contrastive learning method to align the pathology and omics information. POMP leverages the principle of pre-trained models and explores the benefit of aligning multimodal information from the same patient, achieving state-of-the-art results on six cancer datasets from the Cancer Genome Atlas (TCGA). We also show that our contrastive learning method allows us to exploit the cosine similarity of pathological images and omics data as the survival risk score, which can further boost prediction performance compared with other commonly used methods. The code is available at https://github.com/SuixueWang/POMP. Suixue Wang, Huiyuan Lai, Weiliang Huo, Qingchen Zhang 0001 |
IJCAI | 4 |
| 2025 | MASTER: A Multi-granularity Invariant Structure Clustering Scheme for Multi-view ClusteringabstractDeep multi-view clustering has attracted increasing attention in the pattern mining of data. However, most of them perform self-learning mechanisms in a single space, ignoring the fruitful structural information hidden in different-level feature spaces. Meanwhile, they conduct the reconstruction constraint to learn generalized representations of samples, failing to explore the discriminative ability of complementary and consistent information. To address the challenges, a multi-granularity invariant structure clustering scheme (MASTER) is proposed to define a bottom-up process that extracts multi-level information in sample, neighborhood, and category granularities from low-level, high-level, and semantics feature space, respectively. Specifically, it leverages the self-learning reconstruction with information-theoretic overclustering to capture invariant sample structure in the low-level feature space. Then, it models data diffusion of the clustering process in the reliable neighborhood to capture invariant local structure in the high-level feature space. Meanwhile, it defines dual divergences induced by the space geometry to capture invariant global structure in the semantics space. Finally, extensive experiments on 8 real-world datasets show that MASTER achieves state-of-the-art performance compared to 11 baselines. Suixue Wang, Qingchen Zhang 0001, Peng Li 0027, Weiliang Huo |
IJCAI | 5 |
| 2025 | CSF-GAN: Cross-modal Semantic Fusion-based Generative Adversarial Network for Text-guided Image InpaintingabstractMost visual-guided image inpainting methods based on generative adversarial networks (GANs) struggle when the missing region has weak correlations with the surrounding visual context. Recently, diffusion-based methods guided by textual context have been proposed to address this limitation by leveraging additional semantic information to restore corrupted objects. However, these models typically involve more parameters and exhibit slower generation speeds compared to GAN-based approaches. To address this problem, we propose a novel text-guided image inpainting model, the cross-modal semantic fusion generative adversarial network (CSF-GAN). CSF-GAN is designed as a one-stage GAN with the following key contributions. First, a novel semantic fusion module (SFM) is introduced to integrate sentence- and word-level textual context into the inpainting process, enabling more effective guidance from multi-granularity semantic information. Second, a newly designed word-level local discriminator provides detailed feedback to the generator, enhancing the accuracy of generated content in alignment with word-level semantics. Third, two loss functions, the inpainting loss and edge loss, are employed to enhance both structural coherence and textural realism in the generated results. Extensive experiments on two benchmark datasets demonstrate that CSF-GAN outperforms state-of-the-art methods. Suixue Wang, Qingchen Zhang 0001, Liang Zhao 0005, Weiliang Huo, Sijia Hou, Chunjiang Fu |
IJCAI | 5 |