VLDB 2026 Research / reviewers in the wild / expert
Wanwan Shi
dblp:358/8362
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-4730-4610ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Level Domain Adaptation and Contrastive Domain Isolation with Bilinear Fusion for Patient Drug Response PredictionabstractAccurate prediction of patient drug response is critical for precision cancer medicine but remains constrained by limited clinical data. While in vitro cell line data offer a scalable alternative, effective cross-domain transfer remains challenging. Many existing methods tend to overlook heterogeneous domain shifts across biological contexts, underrepresent the intrinsic differences between cell lines and patient tissues, and insufficiently capture high-order gene-drug interactions. To address these challenges, we propose MACB-DRP, a hierarchical transfer learning framework comprising three complementary stages that progressively coordinate adaptation across tissue, drug, and sample levels while enabling representation separation. The framework begins with tissue-aware domain adaptation, leveraging cancer-type classification and unsupervised alignment to preserve biologically meaningful structure across domains. It then incorporates drug-conditioned adversarial transfer for distribution alignment, coupled with bilinear fusion to model nonlinear and high-order gene-drug interactions. Finally, contrastive anchoring with feature-matched pairs enables fine-grained sample-level alignment, while feature-mismatched negatives preserve irreducible biological disparities. Experimental evaluation demonstrates that MACB-DRP achieves comprehensive predictive performance for patient drug responses, with robust results across multiple cancer types and nine drugs, and further reveals hierarchical structure across drugs and tissues in the visualization. These findings highlight the potential of biologically guided domain adaptation for improving translational pharmacogenomics. Yuting Bai, Hanwen Lv, Wanwan Shi, Zhiyi Zou, Jiawei Luo 0001 |
AAAI | 3 |
| 2026 | Topology-Aware Contrastive Learning for Spatially Variable Gene Identification
Wanwan Shi, Juping Li, Nguyen Hoang Tu, Jiawei Luo 0001 |
ICIC (30) | 2 |
| 2026 | Image-Enhanced Multi-Modal Contrastive Transformer for Subcellular Spatial TranscriptomicsabstractRecent advances in spatial molecular imaging technologies have enabled gene expression profiling alongside high-resolution imaging, providing unprecedented opportunities to resolve molecular heterogeneity at subcellular resolution. However, these technologies fail to fully capture cellular characteristics due to the limited number of genes they can detect, which hinder downstream analysis. Spatial imaging data provide high-resolution and fine-grained morphology information, developing computational methods that effectively integrate image features with transcriptomic profiles is crucial for enabling comprehensive subcellular data analysis. In this study, we present SIMMT, an image-enhanced multi-modal contrastive transformer framework for identifying spatial domains and enhancing subcellular data. In the framework, we design a dual transformer architecture to learn multi-modal representations for cells by modeling transcriptomics and morphological images respectively. To fully capture modality interactions within spatial contexts, we introduce a contrastive learning module that enhances cell representation by aligning tissue morphology and gene expression at the cell level. We tested SIMMT on subcellular spatial transcriptomics datasets from human lung cancer tissue, mouse brain tissue, human colorectal cancer tissue, and human ovarian cancer tissue. The results demonstrated that SIMMT consistently outperformed state-of-the-art methods in spatial clustering and gene expression pattern analysis. Our method also effectively demonstrated its ability to identify tumor spatial heterogeneity and uncover potential gene biomarkers in the human bronchiolar adenoma (BA) dataset. Wanwan Shi, Ying Liu 0027, Qiu Xiao, Yuting Bai, Xinling Zeng, Chee Keong Kwoh 0001, Jiawei Luo 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | High-Frequency-Aware Graph Integration for Subcellular Spatial TranscriptomicsabstractRecent advances in spatial transcriptomics have enabled subcellular-resolution profiling of gene expression, offering unprecedented opportunities to investigate intracellular architecture and local microenvironmental interactions. Graph neural networks (GNNs) have shown great promise in modeling spatial transcriptomics data. However, existing GNN-based methods primarily focus on low-frequency signals, overlooking high-frequency signals critical for resolving transcriptional differences across subcellular compartments and cell boundaries. This limits their ability to characterize fine-grained structural and functional heterogeneity within tissues, hindering accurate spatial domain identification. In this study, we propose HiFi-ST, a High-Frequency-Aware Graph Integration framework for subcellular spatial transcriptomics. HiFi-ST employs a high-pass filter to extract high-frequency transcriptional differences, which are then integrated with spatial contexts through a transformer-based architecture. A contrastive learning module is designed to enhance cell representation by aligning spatial organization with transcriptional heterogeneity. Comprehensive experiments on subcellular datasets demonstrated that HiFi-ST consistently outperformed six state-of-the-art methods in spatial clustering, gene expression enhancement, and niche identification. Wanwan Shi, Yahui Long, Ying Liu 0027, Qiu Xiao, Yuting Bai, Xiaoyi Peng, Xiangtao Chen, Jiawei Luo 0001 |
BIBM | 1 |
| 2025 | scGANCL: Bidirectional Generative Adversarial Network for Imputing scRNA-Seq Data With Contrastive LearningabstractThe advent of single-cell RNA sequencing (scRNA-seq) has offering unprecedented insights at the single-cell level. This groundbreaking technology has opened new pathways for understanding cellular diversity and revealing novel insights into disease mechanisms. However, the analysis of scRNA-seq data is challenging, primarily due to dropout events caused by technical noise. Developing effective imputation methods is crucial for the reliable and informative analysis of scRNA-seq data. While deep learning-based approaches have been proposed for scRNA-seq data imputation, they often fall short of optimal performance, especially in identifying rare cell types. Here we propose a novel self-supervised deep learning model named scGANCL for scRNA-seq data imputation. scGANCL combines bidirectional generative adversarial network (BiGAN) with contrastive learning (CL) to enhance imputation performance. To fully exploit gene expression profiles, a contrastive learning module is introduced to enhance the representation learning of cells by minimizing the discrepancy between the distributions of real and generated data. Comprehensive experiments have been conducted on ten simulated and seven real datasets to validate scGANCL's effectiveness. The results demonstrated scGANCL consistently outperformed seven state-of-the-art methods across various downstream tasks. Ablation studies further validated the contribution of each component to the overall performance of the model. Wanwan Shi, Yahui Long, Jiawei Luo 0001, Ying Liu 0027, Zehao Xiong, Zhongyuan Xu |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | scMID: A Deep Multi-Omics Integration Framework for Comprehensive Single-Cell Data AnalysisabstractBiological research on single cells has witnessed remarkable progress in recent years, with downstream analyses playing a crucial role in uncovering cellular functions and mechanisms. Traditional single-cell analyses, which predominantly rely on single-omics data such as single-cell RNA sequencing, are inherently limited. These methods can only capture one aspect of cellular information, overlooking the complex interplay between different molecular layers, and thus are prone to introducing biases in results. The advent of single-cell multi-omics sequencing technologies has revolutionized this landscape. By enabling the integration of diverse molecular profiles, including transcriptomics, epigenomics, and proteomics, these technologies offer a more holistic view of cellular functions. However, existing integration methods often lack the ability to handle the complexity and heterogeneity of multi-omics data, limiting their application in in-depth single-cell studies. In this study, we propose an analysis method based on single-cell multi-omics data integration and dropout pattern (scMID). Specifically, scMID utilizes omics-independent deep autoencoders for the alignment of multi-omics data, employs GCN algorithm for data integration, and calculates the gene importance by combining the gene similarity obtained from the binarized dropout pattern. Meanwhile, scMID proposes a dual-strategy for feature gene screening, aiming to identify genes with high biological significance that best match the structural characteristics of reference data. Experimental results demonstrate that scMID significantly improves the accuracy of single-cell clustering in downstream analyses, breaking through the limitations of traditional feature selection methods and providing a superior analytical framework for decoding complex biological information. Qiu Xiao, Wanwan Shi, Ying Zuo, Fei Guo 0001, Jiawei Luo 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | A multi-modality and multi-granularity collaborative learning framework for identifying spatial domains and spatially variable genesabstractMOTIVATION: Recent advances in spatial transcriptomics technologies have provided multi-modality data integrating gene expression, spatial context, and histological images. Accurately identifying spatial domains and spatially variable genes is crucial for understanding tissue structures and biological functions. However, effectively combining multi-modality data to identify spatial domains and determining SVGs closely related to these spatial domains remains a challenge. RESULTS: In this study, we propose spatial transcriptomics multi-modality and multi-granularity collaborative learning (spaMMCL). For detecting spatial domains, spaMMCL mitigates the adverse effects of modality bias by masking portions of gene expression data, integrates gene and image features using a shared graph convolutional network, and employs graph self-supervised learning to deal with noise from feature fusion. Simultaneously, based on the identified spatial domains, spaMMCL integrates various strategies to detect potential SVGs at different granularities, enhancing their reliability and biological significance. Experimental results demonstrate that spaMMCL substantially improves the identification of spatial domains and SVGs. AVAILABILITY AND IMPLEMENTATION: The code and data of spaMMCL are available on Github: Https://github.com/liangxiao-cs/spaMMCL. Baiyun Chen, Wei Liu 0296, Wanwan Shi, Yongwang Wang, Xiangtao Chen, Jiawei Luo 0001 |
Bioinform. | 6 |
| 2023 | scSRL: Siamese Representation Learning-based method for analyzing single-cell RNA-seq dataabstractSingle-cell RNA sequencing (scRNA-seq) technology is utilized to analyze cellular heterogeneity, perform cellular-level biological research and derive novel insights from complex cellular systems. However, the raw scRNA-seq data is not directly suitable for downstream task analysis due to its high variability, sparsity and dimensionality. Therefore, in this study, we propose a new self-supervised framework based on siamese representation learning, named scSRL which can fully explore the intrinsic properties of cells by maximizing the similarity between positive pairs. These positive pairs are constructed by multiple data augmentation operations to further increase data diversity and better learn latent representation. Moreover, our method employs a gradient stopping strategy to mitigate collapsing in the siamese network. It is worth noting that the scSRL focuses on aggregating cells with similar functions without introducing negative samples, which can avoid additional computational cost. Finally, We evaluated scSRL on 10 real datasets for downstream tasks such as clustering, classification and visualization, and it consistently exhibited outstanding performance in all these fundamental tasks. Meanwhile, we did pseudotime inference experiments in two embryonic development datasets, and the scSRL model can accurately reconstruct cell trajectory and describe cell developmental process. scSRL is currently an open-source method, available at https://github.com/zysun17/scSRL. Zhaoyang Sun, Ying Liu 0027, Wanwan Shi, Jiawei Luo 0001 |
BIBM | 4 |
| 2023 | Spatial-MGCN: a novel multi-view graph convolutional network for identifying spatial domains with attention mechanismabstractMOTIVATION: Recent advances in spatial transcriptomics technologies have enabled gene expression profiles while preserving spatial context. Accurately identifying spatial domains is crucial for downstream analysis and it requires the effective integration of gene expression profiles and spatial information. While increasingly computational methods have been developed for spatial domain detection, most of them cannot adaptively learn the complex relationship between gene expression and spatial information, leading to sub-optimal performance. RESULTS: To overcome these challenges, we propose a novel deep learning method named Spatial-MGCN for identifying spatial domains, which is a Multi-view Graph Convolutional Network (GCN) with attention mechanism. We first construct two neighbor graphs using gene expression profiles and spatial information, respectively. Then, a multi-view GCN encoder is designed to extract unique embeddings from both the feature and spatial graphs, as well as their shared embeddings by combining both graphs. Finally, a zero-inflated negative binomial decoder is used to reconstruct the original expression matrix by capturing the global probability distribution of gene expression profiles. Moreover, Spatial-MGCN incorporates a spatial regularization constraint into the features learning to preserve spatial neighbor information in an end-to-end manner. The experimental results show that Spatial-MGCN outperforms state-of-the-art methods consistently in several tasks, including spatial clustering and trajectory inference. Jiawei Luo 0001, Ying Liu 0027, Wanwan Shi, Zehao Xiong, Cong Shen 0002, Yahui Long |
Briefings Bioinform. | 4 |
| 2023 | scGCL: an imputation method for scRNA-seq data based on graph contrastive learningabstractMOTIVATION: Single-cell RNA-sequencing (scRNA-seq) is widely used to reveal cellular heterogeneity, complex disease mechanisms and cell differentiation processes. Due to high sparsity and complex gene expression patterns, scRNA-seq data present a large number of dropout events, affecting downstream tasks such as cell clustering and pseudo-time analysis. Restoring the expression levels of genes is essential for reducing technical noise and facilitating downstream analysis. However, existing scRNA-seq data imputation methods ignore the topological structure information of scRNA-seq data and cannot comprehensively utilize the relationships between cells. RESULTS: Here, we propose a single-cell Graph Contrastive Learning method for scRNA-seq data imputation, named scGCL, which integrates graph contrastive learning and Zero-inflated Negative Binomial (ZINB) distribution to estimate dropout values. scGCL summarizes global and local semantic information through contrastive learning and selects positive samples to enhance the representation of target nodes. To capture the global probability distribution, scGCL introduces an autoencoder based on the ZINB distribution, which reconstructs the scRNA-seq data based on the prior distribution. Through extensive experiments, we verify that scGCL outperforms existing state-of-the-art imputation methods in clustering performance and gene imputation on 14 scRNA-seq datasets. Further, we find that scGCL can enhance the expression patterns of specific genes in Alzheimer's disease datasets. AVAILABILITY AND IMPLEMENTATION: The code and data of scGCL are available on Github: https://github.com/zehaoxiong123/scGCL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zehao Xiong, Jiawei Luo 0001, Wanwan Shi, Ying Liu 0027, Zhongyuan Xu |
Bioinform. | 3 |