VLDB 2026 Research / reviewers in the wild / expert
Wenwen Min
dblp:174/0484
· DBLP profile ↗
46ranked-venue papers
11as first author
42since 2021 · last 2026
0000-0002-2558-2911ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 36 · 8 first-author · 32 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpaCRD: Multimodal Deep Fusion of Histology and Spatial Transcriptomics for Cancer Region DetectionabstractAccurate detection of cancer tissue regions (CTR) enables deeper analysis of the tumor microenvironment and offers crucial insights into treatment response. Traditional CTR detection methods, which typically rely on the rich cellular morphology in histology images, are susceptible to a high rate of false positives due to morphological similarities across different tissue regions. The groundbreaking advances in spatial transcriptomics (ST) provide detailed cellular phenotypes and spatial localization information, offering new opportunities for more accurate cancer region detection. However, current methods are unable to effectively integrate histology images with ST data, especially in the context of cross-sample and cross-platform/batch settings for accomplishing the CTR detection. To address this challenge, we propose SpaCRD, a transfer learning-based method that deeply integrates histology images and ST data to enable reliable CTR detection across diverse samples, platforms, and batches. Once trained on source data, SpaCRD can be readily generalized to accurately detect cancerous regions across samples from different platforms and batches. The core of SpaCRD is a category-regularized variational reconstruction-guided bidirectional cross-attention fusion network, which enables the model to adaptively capture latent co-expression patterns between histological features and gene expression from multiple perspectives. Extensive benchmark analysis on 23 matched histology-ST datasets spanning various disease types, platforms, and batches demonstrates that SpaCRD consistently outperforms existing eight state-of-the-art methods in CTR detection. Shuailin Xue, Jun Wan 0005, Wenwen Min |
AAAI | 4 |
| 2026 | SpaAGMF: Adaptive Gated Multi-scale Fusion of Histology and Spatial Transcriptomics for Cancer Region Classification
Yunfeng Mao, Wenwen Min |
ISBRA (1) | 2 |
| 2026 | SpaBiT: enhancing spatial transcriptomics resolution via bidirectional attention transformersabstractMOTIVATION: Spatial transcriptomics (STs) enables the precise mapping of gene expression within tissue architecture, however its application is often limited by low spatial resolution and sparse sampling. While existing deep learning methods leverage histology images, spatial coordinates, or low-resolution expression data to predict high-density profiles, these methods are limited in either capturing the intrinsic constraints between histological context and spatial topology or ignoring the complex local neighborhood relationships between spots. RESULTS: To address these limitations, we propose SpaBiT, a multimodal framework designed to enhance ST resolution via a bidirectional attention mechanism. At its core, SpaBiT employs a bidirectional cross-attention module to facilitate precise information exchange between image features and neighborhood-aware representations learned via a graph attention network. This design explicitly models the synergistic constraints between local morphology and spatial graph topology, yielding high-fidelity, high-density gene expression maps. SpaBiT exhibits competitive performance in reconstructing complex spatial gene expression, outperforming the benchmark models utilized in this study across various quantitative metrics, providing a robust tool for deciphering complex tissue microenvironments. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/wenwenmin/SpaBiT. Wenwen Min |
Bioinform. | 3 |
| 2026 | SpaMGCL: Neighborhood-aware masked graph contrastive learning for spatial transcriptomics analysisabstractThe rapid development of spatial transcriptomics (ST) has enabled the simultaneous acquisition of gene expression profiles and spatial coordinates, providing unprecedented opportunities to understand the structural organization of tissues. However, accurate spatial domain identification remains a critical challenge, as existing methods often struggle to capture the complex interactions between gene expression patterns and spatial topology. To address this issue, we propose SpaMGCL ( Spa tial M asked G raph C ontrastive L earning), a novel self-supervised learning framework for spatial domain identification. The proposed framework integrates a masking mechanism with graph contrastive learning, leveraging both gene expression information and spatial constraints to learn highly discriminative latent representations. Specifically, SpaMGCL introduces a masked graph contrastive learning framework to enhance spatial representation learning. First, a masking strategy is employed to simulate missing data, encouraging the model to learn robust representations by reconstructing masked regions. Subsequently, spatial perturbations are applied to generate contrastive views, from which a shared encoder extracts hierarchical features. Meanwhile, contrastive learning is utilized to obtain discriminative spatial representations. Furthermore, SpaMGCL incorporates a neighborhood-aware mechanism that utilizes local spatial context as contrastive anchors, enabling the model to effectively characterize spatially coherent patterns in tissue structures. This design improves intra-domain compactness while enhancing inter-domain separability. We comprehensively evaluate SpaMGCL on seven publicly available ST datasets. Experimental results show that SpaMGCL outperforms existing state-of-the-art methods in spatial domain identification, demonstrating its superior effectiveness and robustness. The implementation is publicly available at https://github.com/wenwenmin/SpaMGCL . Jinjie Zhao, Donghai Fang, Wenwen Min |
Expert Syst. Appl. | 3 |
| 2026 | FGTBT: Frequency-guided task-balancing transformer for unified facial landmark detection
Jun Wan 0005, Xinyu Xiong, Zhihui Lai 0001, Jie Zhou 0009, Wenwen Min |
Inf. Sci. | 6 |
| 2026 | Multi-view masked graph representation learning with semantic alignment for spatial multi-omics clustering
Jinjie Zhao, Jun Wan 0005, Wenwen Min |
Pattern Recognit. | 3 |
| 2026 | BS-LDM: Effective Bone Suppression in High-Resolution Chest X-Ray Images With Conditional Latent Diffusion ModelsabstractLung diseases represent a significant global health challenge, with Chest X-Ray (CXR) being a key diagnostic tool due to its accessibility and affordability. Nonetheless, the detection of pulmonary lesions is often hindered by overlapping bone structures in CXR images, leading to potential misdiagnoses. To address this issue, we develop an end-to-end framework called BS-LDM, designed to effectively suppress bone in high-resolution CXR images. This framework is based on conditional latent diffusion models and incorporates a multi-level hybrid loss-constrained vector-quantized generative adversarial network which is crafted for perceptual compression, ensuring the preservation of details. To further enhance the framework's performance, we utilize offset noise in the forward process, and a temporal adaptive thresholding strategy in the reverse process. These additions help minimize discrepancies in generating low-frequency information of soft tissue images. Additionally, we have compiled a high-quality bone suppression dataset named SZCH-X-Rays. This dataset includes 818 pairs of high-resolution CXR and soft tissue images collected from our partner hospital. Moreover, we processed 241 data pairs from the JSRT dataset into negative images, which are more commonly used in clinical practice. Our comprehensive experiments and downstream evaluations reveal that BS-LDM excels in bone suppression, underscoring its clinical value. Yifei Sun 0005, Zhanghao Chen, Wenming Deng, Jin Liu 0012, Wenwen Min, Ahmed El-Azab, Changmiao Wang, Ruiquan Ge |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | RTGMFF: Enhanced fMRI-Based Brain Disorder Diagnosis via ROI-Driven Text Generation and Multimodal Feature FusionabstractFunctional magnetic resonance imaging (fMRI) is a powerful tool for probing brain function, yet reliable clinical diagnosis is hampered by low signal-to-noise ratios, inter-subject variability, and the limited frequency awareness of prevailing CNN- and Transformer-based models. Moreover, most fMRI datasets lack textual annotations that could contextualize regional activation and connectivity patterns. We introduce RTGMFF, a framework that unifies automatic ROI-level text generation with multimodal feature fusion for brain-disorder diagnosis. RTGMFF consists of three components: (i) ROI-driven fMRI text generation deterministically condenses each subject's activation, connectivity, age, and sex into reproducible text tokens; (ii) Hybrid frequency-spatial encoder fuses a hierarchical waveletmamba branch with a cross-scale Transformer encoder to capture frequency-domain structure alongside long-range spatial dependencies; and (iii) Adaptive semantic alignment module embeds the ROI token sequence and visual features in a shared space, using a regularized cosine-similarity loss to narrow the modality gap. Extensive experiments on the ADHD-200 and ABIDE benchmarks show that RTGMFF surpasses current methods in diagnostic accuracy, achieving notable gains in sensitivity, specificity, and area under the ROC curve. Code is available at https://github.com/BeistMedAI/RTGMFF. Junhao Jia, Yifei Sun 0005, Yunyou Liu, Changmiao Wang, Fei-wei Qin, Yong Peng 0001, Wenwen Min |
BIBM | 8 |
| 2025 | RE-SAM2: Boosting Few-Shot Medical Image Segmentation via Reinforcement Learning and Ensemble LearningabstractDeep learning models for medical image segmentation often encounter difficulties when there is a lack of annotated data. While current few-shot segmentation methods have reduced these challenges, they frequently fail to fully utilize the information in the limited samples available. Additionally, they typically depend on large quantities of unlabeled data with pseudo-labels for domain adaptation. In response to these issues, we propose RE-SAM2, a novel framework for fewshot medical image segmentation that combines reinforcement learning with ensemble learning. The central concept involves retaining the reward model from reinforcement learning after training and integrating it into the model through ensemble learning techniques. Unlike previous methods, RE-SAM2 does not require extra unlabeled data and achieves notable improvements in segmentation accuracy with limited supervision. Experiments conducted on benchmark datasets reveal that RE-SAM2 surpasses current leading approaches. The code is available at the link https://github.com/zzzzz37/RE-SAM2. Shougan Teng, Wenwen Min, Changmiao Wang, Zhenbing Liu |
BIBM | 3 |
| 2025 | Inferring Super-Resolved Gene Expression by Integrating Histology Images and Spatial Transcriptomics with HISTEX
Shuailin Xue, Changmiao Wang, Xiaomao Fan, Wenwen Min |
MICCAI (13) | 4 |
| 2025 | sTPLS: identifying common and specific correlated patterns under multiple biological conditionsabstractThe rapidly emerging large-scale data in diverse biological research fields present valuable opportunities to explore the underlying mechanisms of tissue development and disease progression. However, few existing methods can simultaneously capture common and condition-specific association between different types of features across different biological conditions, such as cancer types or cell populations. Therefore, we developed the sparse tensor-based partial least squares (sTPLS) method, which integrates multiple pairs of datasets containing two types of features but derived from different biological conditions. We demonstrated the effectiveness and versatility of sTPLS through simulation study and three biological applications. By integrating the pairwise pharmacogenomic data, sTPLS identified 11 gene-drug comodules with high biological functional relevance specific for seven cancer types and two comodules that shared across multi-type cancers, such as breast, ovarian, and colorectal cancers. When applied to single-cell data, it uncovered nine gene-peak comodules representing transcriptional regulatory relationships specific for five cell types and three comodules shared across similar cell types, such as intermediate and naïve B cells. Furthermore, sTPLS can be directly applied to tensor-structured data, successfully revealing shared and distinct cell communication patterns mediated by the MK signaling pathway in coronavirus disease 2019 patients and healthy controls. These results highlight the effectiveness of sTPLS in identifying biologically meaningful relationships across diverse conditions, making it useful for multi-omics integrative analysis. Wenwen Min |
Briefings Bioinform. | 2 |
| 2025 | Inferring single-cell resolution spatial gene expression via fusing spot-based spatial transcriptomics, location, and histology using GCNabstractSpatial transcriptomics (ST technology allows for the detection of cellular transcriptome information while preserving the spatial location of cells. This capability enables researchers to better understand the cellular heterogeneity, spatial organization, and functional interactions in complex biological systems. However, current technological methods are limited by low resolution, which reduces the accuracy of gene expression levels. Here, we propose scstGCN, a multimodal information fusion method based on Vision Transformer and Graph Convolutional Network that integrates histological images, spot-based ST data and spatial location information to infer super-resolution gene expression profiles at single-cell level. We evaluated the accuracy of the super-resolution gene expression profiles generated on diverse tissue ST datasets with disease and healthy by scstGCN along with their performance in identifying spatial patterns, conducting functional enrichment analysis, and tissue annotation. The results show that scstGCN can predict super-resolution gene expression accurately and aid researchers in discovering biologically meaningful differentially expressed genes and pathways. Additionally, scstGCN can segment and annotate tissues at a finer granularity, with results demonstrating strong consistency with coarse manual annotations. Our source code and all used datasets are available at https://github.com/wenwenmin/scstGCN and https://zenodo.org/records/12800375. Shuailin Xue, Wenwen Min |
Briefings Bioinform. | 4 |
| 2025 | SpaICL: image-guided curriculum strategy-based graph contrastive learning for spatial transcriptomics clusteringabstractSpatial transcriptomics, by capturing both gene expression and spatial information, holds great promise for unraveling the complex organization of tissues. In this study, we introduce SpaICL, an image-guided curriculum strategy-based graph contrastive learning framework for spatial transcriptomics clustering. SpaICL integrates gene expression, spatial coordinates, and histological image features to construct a low-dimensional latent representation that enhances the de-lineation of spatial functional domains. The model employs a complementary masking strategy and a shared graph neural network encoder to generate dual embeddings, while a dual cross-attention mechanism aligns local and global features across multiple modalities. Additionally, the curriculum learning module further facilitates the gradual integration of neighborhood information, effectively mitigating the over-smoothing issues associated with fixed adjacency matrices. We evaluated the performance of SpaICL on five benchmark spatial transcriptomics datasets, achieving superior results compared to existing baseline methods. Moreover, SpaICL demonstrates significant potential in downstream analytical applications. The code of SpaICL is available at https://github.com/wenwenmin/SpaICL. Jingcheng Zhao, Wenwen Min |
Briefings Bioinform. | 2 |
| 2025 | IGCLAPS: an interpretable graph contrastive learning method with adaptive positive sampling for scRNA-seq data analysisabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology enables biological research at single-cell resolution. Cell clustering is a crucial task in scRNA-seq data analysis since it provides insights into cell heterogeneity. Although existing methods have made significant progress in this task, it remains challenging to fully utilize the relationship among cells. RESULTS: We propose Interpretable Graph Contrastive Learning method with Adaptive Positive Sampling (IGCLAPS), a novel end-to-end graph contrastive clustering method for scRNA-seq data analysis. Specifically, IGCLAPS learns low-dimensional embeddings with a graph transformer, based on which a dual-head graph contrastive learning module is used to perform dimension reduction and cell clustering simultaneously. Besides, an accurate definition of positive sample pairs is crucial in contrastive learning, we devise an adaptive positive sampling module, which dynamically identifies true positive sample pairs based on both expression similarity and soft cluster labels generated by the contrastive learning module. Extensive experiments on a series of real datasets including cell clustering, visualization, and differential expression analysis demonstrate that IGCLAPS can effectively enhance clustering performance and generate interpretable gene expression patterns of scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The source codes of IGCLAPS are available at https://github.com/ZhengWeihuaYNU/IGCLAPS. Wenwen Min, Shunfang Wang |
Bioinform. | 2 |
| 2025 | Geometry-informed multimodal fusion network for enhancing high-density spatial transcriptomics from histology images
Zhiceng Shi, Shuailin Xue, Wenwen Min |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Improving cell-type composition inference in spatial transcriptomics with SpaDAMAabstractAccurate determination of cell-type composition in disease-relevant tissues is essential for identifying potential disease targets and understanding tissue heterogeneity. Most current spatial transcriptomics (ST) technologies lack single-cell resolution, which makes precise cell-type composition identification challenging. Several deconvolution methods have been developed to address this limitation by relying on single-cell RNA sequencing (scRNA-seq) data from the same tissue as a reference to estimate the cell type composition in ST data spots. However, these methods often overlook the inherent differences between scRNA-seq and ST data. To overcome this challenge, we introduce a Domain-Adversarial Masked Autoencoder (SpaDAMA) method. SpaDAMA leverages Domain-Adversarial Learning (DAL) to facilitate effective knowledge transfer from the source domain (pseudo-ST data generated from scRNA-seq) to the target domain (real ST data). Through adversarial training, SpaDAMA harmonizes the distributions of both datasets and maps them onto a unified latent representation, thereby reducing discrepancies in data modalities. Furthermore, to strengthen the model's capability in extracting reliable features from real ST data, SpaDAMA employs masking strategies that effectively minimize noise and mitigate spatial artifacts. We validated SpaDAMA on 32 simulated datasets and 4 real-world datasets, demonstrating its superior performance in cell-type deconvolution and providing a promising tool for spatial transcriptomic analyses. Wenwen Min |
PLoS Comput. Biol. | 4 |
| 2025 | SpaMask: Dual masking graph autoencoder with contrastive learning for spatial transcriptomicsabstractUnderstanding the spatial locations of cell within tissues is crucial for unraveling the organization of cellular diversity. Recent advancements in spatial resolved transcriptomics (SRT) have enabled the analysis of gene expression while preserving the spatial context within tissues. Spatial domain characterization is a critical first step in SRT data analysis, providing the foundation for subsequent analyses and insights into biological implications. Graph neural networks (GNNs) have emerged as a common tool for addressing this challenge due to the structural nature of SRT data. However, current graph-based deep learning approaches often overlook the instability caused by the high sparsity of SRT data. Masking mechanisms, as an effective self-supervised learning strategy, can enhance the robustness of these models. To this end, we propose SpaMask, dual masking graph autoencoder with contrastive learning for SRT analysis. Unlike previous GNNs, SpaMask masks a portion of spot nodes and spot-to-spot edges to enhance its performance and robustness. SpaMask combines Masked Graph Autoencoders (MGAE) and Masked Graph Contrastive Learning (MGCL) modules, with MGAE using node masking to leverage spatial neighbors for improved clustering accuracy, while MGCL applies edge masking to create a contrastive loss framework that tightens embeddings of adjacent nodes based on spatial proximity and feature similarity. We conducted a comprehensive evaluation of SpaMask on eight datasets from five different platforms. Compared to existing methods, SpaMask achieves superior clustering accuracy and effective batch correction. Wenwen Min, Donghai Fang |
PLoS Comput. Biol. | 1 |
| 2025 | Weighted Sparse Partial Least Squares With Joint Sample and Feature Selection for Integrating Multi-Omics DataabstractSparse Partial Least Squares (sPLS) is a common dimensionality reduction technique for data fusion, which projects data samples from two views by seeking linear combinations with a small number of variables with the maximum variance. However, sPLS extracts the combinations between two data sets with all data samples so that it cannot detect latent subsets of samples. To extend the application of sPLS by identifying a specific subset of samples and remove outliers, we propose an $\ell _\infty /\ell _{0}$-norm constrained weighted sparse PLS ($\ell _\infty /\ell _{0}$-wsPLS) method for joint sample and feature selection, where the $\ell _\infty /\ell _{0}$-norm constrains are used to select a subset of samples. We prove that the $\ell _\infty /\ell _{0}$-norm constrains have the Kurdyka-Łojasiewicz property so that a globally convergent algorithm is developed to solve it. Moreover, multi-view data with a same set of samples can be available in various real problems. To this end, we extend the $\ell _\infty /\ell _{0}$-wsPLS model and propose two multi-view wsPLS models for multi-view data fusion. We develop an efficient iterative algorithm for each multi-view wsPLS model and show its convergence property. As well as numerical and biomedical data experiments demonstrate the efficiency of the proposed methods. Wenwen Min, Taosheng Xu, Chris Ding |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Multi-Slice Spatial Transcriptomics Data Integration Analysis with STG3NetabstractWith the rapid development of Spatially Resolved Transcriptomics (SRT) technology, which allows for the mapping of gene expression within tissue sections, the integrative analysis of SRT datasets from multiple tissue slices has become increasingly important. However, batch effects between multiple slices pose significant challenges in analyzing SRT data. To address these challenges, we have developed a plug-and-play batch correction method called Global Nearest Neighbor (G2N) anchor pairs selection. G2N effectively mitigates batch effects by selecting representative anchor pairs across slices. Building upon G2N, we propose STG3Net, which cleverly combines masked graph convolutional autoencoders as backbone modules. These autoencoders, integrated with generative adversarial learning, enable STG3Net to achieve robust multi-slice spatial domain identification and batch correction. We comprehensively evaluate the feasibility of STG3Net on three multiple SRT datasets from different platforms, considering accuracy, consistency, and the F1LISI metric (a measure of batch effect correction efficiency). Compared to existing methods, STG3Net achieves the best overall performance while preserving the biological variability and connectivity between slices. Source code and all public datasets used in this paper are available at https://github.com/wenwenmin/ STG3Net and https://zenodo.org/records/12737170. Donghai Fang, Wenwen Min |
BIBM | 3 |
| 2024 | Masked Graph Autoencoders with Contrastive Augmentation for Spatially Resolved Transcriptomics DataabstractWith the rapid advancement of Spatial Resolved Transcriptomics (SRT) technology, it is now possible to comprehensively measure gene transcription while preserving the spatial context of tissues. Spatial domain identification and gene denoising are key objectives in SRT data analysis. We propose a Masked Graph Autoencoder with Contrastively augmentation (STMGAC) to learn low-dimensional latent representations for domain identification of Spatial Transcriptomics (ST). In the latent space, persistent signals for representations are obtained through self-distillation to guide self-supervised matching. At the same time, positive and negative anchor pairs are constructed using triplet learning to augment the discriminative ability. We evaluated the performance of STMGAC on five datasets, achieving results superior to those of existing baseline methods. All code and public datasets used in this paper are available at https://github.com/wenwenmin/STMGAC and https://zenodo.org/records/13253801. Donghai Fang, Dongting Xie, Wenwen Min |
BIBM | 4 |
| 2024 | Masked adversarial neural network for cell type deconvolution in spatial transcriptomicsabstractAccurately determining cell type composition in disease-relevant tissues is crucial for identifying disease targets. Most existing spatial transcriptomics (ST) technologies cannot achieve single-cell resolution, making it challenging to accurately determine cell types. To address this issue, various deconvolution methods have been developed. Most of these methods use single-cell RNA sequencing (scRNA-seq) data from the same tissue as a reference to infer cell types in ST data spots. However, they often overlook the differences between scRNA-seq and ST data. To overcome this limitation, we propose a Masked Adversarial Neural Network (MACD). MACD employs adversarial learning to align real ST data with simulated ST data generated from scRNA-seq data. By mapping them into a unified latent space, it can minimize the differences between the two types of data. Additionally, MACD uses masking techniques to effectively learn the features of real ST data and mitigate noise. We evaluated MACD on 32 simulated datasets, demonstrating its accuracy in performing cell type deconvolution. All code and public datasets used in this paper are available at https://github.com/wenwenmin/MACD. Shunfang Wang, Wenwen Min |
BIBM | 4 |
| 2024 | Masked Conditional Diffusion Model with GNN for Spatial Transcriptomics Data ImputationabstractSpatially resolved transcriptomics represents a significant advancement in single-cell analysis by offering both gene expression data and their corresponding physical locations. However, this high degree of spatial resolution entails a drawback, as the resulting spatial transcriptomic data at the cellular level is notably plagued by a high incidence of missing values. Furthermore, most existing imputation methods either overlook the spatial information between spots or compromise the overall gene expression data distribution. To address these challenges, our primary focus is on effectively utilizing the spatial location information within spatial transcriptomic data to impute missing values, while preserving the overall data distribution. We introduce stMCDI, a masked conditional diffusion model for spatial transcriptomics data imputation, which employs a denoising network trained using randomly masked data portions as guidance, with the unmasked data serving as conditions. Additionally, it utilizes a GNN encoder to integrate the spatial position information, thereby enhancing model performance. Compared with baseline methods, our model achieves state-of-the-art performance in all evaluation metrics on six real-world datasets. The results obtained from spatial transcriptomics datasets elucidate the performance of our methods relative to existing approaches. Our code can be accessed at https://github.com/wenwenmin/stMCDI. Wenwen Min, Shunfang Wang, Changmiao Wang, Taosheng Xu |
BIBM | 2 |
| 2024 | scASDC: Attention Enhanced Structural Deep Clustering for Single-cell RNA-seq DataabstractSingle-cell RNA sequencing (scRNA-seq) data analysis is pivotal for understanding cellular heterogeneity. However, the high sparsity and complex noise patterns inherent in scRNA-seq data present significant challenges for traditional clustering methods. To address these issues, we propose a deep clustering method, Attention-Enhanced Structural Deep Embedding Graph Clustering (scASDC), which integrates multiple advanced modules to improve clustering accuracy and robustness. Our approach employs a multi-layer graph convolutional network (GCN) to capture high-order structural relationships between cells, termed as the graph autoencoder module. We introduce a ZINB-based autoencoder module that extracts content information from the data and learns latent representations of gene expression. These modules are further integrated through an attention fusion mechanism, ensuring effective combination of gene expression and structural information at each layer of the GCN. Additionally, a self-supervised learning module is incorporated to enhance the robustness of the learned embeddings. Extensive experiments demonstrate that scASDC outperforms existing state-of-the-art methods, providing a robust and effective solution for single-cell clustering tasks. All code and public datasets used in this paper are available at https://github.com/wenwenmin/scASDC. Wenwen Min, Taosheng Xu, Guangsheng Wu, Shunfang Wang |
BIBM | 1 |
| 2024 | High-Resolution Spatial Transcriptomics from Histology Images using HisToSGEabstractSpatial transcriptomics (ST) is a groundbreaking genomic technology that enables spatial localization analysis of gene expression within tissue sections. However, it is significantly limited by high costs and sparse spatial resolution. An alternative, more cost-effective strategy is to use deep learning methods to predict high-density gene expression profiles from histological images. However, existing methods struggle to capture rich image features effectively or rely on low-dimensional positional coordinates, making it difficult to accurately predict high- resolution gene expression profiles. To address these limitations, we developed HisToSGE which employs a Pathology Image Large Model (PILM) to extract rich image features from histological images and utilizes a feature learning block to robustly generate high-resolution gene expression profiles. We evaluated HisToSGE on four ST datasets, comparing its performance with five state-of-the-art baseline methods. The results demonstrate that HisToSGE excels in generating high-resolution gene expression profiles and performing downstream tasks such as spatial domain identification. All code and public datasets used in this paper are available at https://github.com/wenwenmin/HisToSGE. Zhiceng Shi, Shuailin Xue, Wenwen Min |
BIBM | 4 |
| 2024 | Pretrained-Guided Conditional Diffusion Models for Microbiome Data AnalysisabstractEmerging evidence indicates that human cancers are intricately linked to human microbiomes, forming an inseparable connection. However, due to limited sample sizes and significant data loss during collection for various reasons, some machine learning methods have been proposed to address the issue of missing data. These methods have not fully utilized the known clinical information of patients to enhance the accuracy of data imputation. Therefore, we introduce mbVDiT, a novel pre-trained conditional diffusion model for microbiome data imputation and denoising, which uses the unmasked data and patient metadata as conditional guidance for imputating missing values. It is also uses VAE to integrate the the other public microbiome datasets to enhance model performance. The results on the microbiome datasets from three different cancer types demonstrate the performance of our methods in comparison with existing methods. Source code and all public datasets used in this paper are available at Github (https://github.com/wenwenmin/mbVDiT) and Zenodo (https://zenodo.org/records/13254073). Xinyuan Shi, Wenwen Min |
BIBM | 3 |
| 2024 | CCLNet: Causal and Contrastive Learning Framework for Enhanced Pulmonary Embolism DetectionabstractThe fusion of multimodal medical data is crucial for helping doctors make accurate treatment decisions. For example, combining Computed Tomography Pulmonary Angiography (CTPA) with Electronic Health Records (EHR) can significantly improve the accuracy of Pulmonary Embolism (PE) detection, thereby increasing patient survival rates. Although multimodal learning has advantages in PE diagnosis, the heterogeneity of multimodal data poses a significant challenge to accurate diagnosis. The natural semantic and structural differences between data modalities make it difficult to effectively integrate their information. In addition, within a single modality, the existence of redundant and irrelevant information introduces unnecessary variability, making the data more complex, and making stable diagnosis challenging. To address these issues, we propose a new framework called CCLNet, which includes a contrastive learning component for addressing inter-modality heterogeneity and a causal learning component for handling intra-modality heterogeneity. Specifically, we achieve precise alignment between visual and tabular modalities by using global-level information to soften labels during contrastive learning. In addition, by using causal intervention methods to eliminate the influence of heterogeneous factors within the modality, we can accurately reveal the causal relationship between features and targets, thereby improving the accuracy and stability of the model. Experimental results demonstrate that our method performs excellently, achieving the best results. Our code is available at https://github.com/LeavingStarW/CLPE. Ruiquan Ge, Jianxun Yu, Fei-wei Qin, Nannan Li 0001, Wenwen Min, Ahmed El-Azab, Changmiao Wang |
BIBM | 7 |
| 2024 | Polymorphic multi-head attention aggregation network for skin lesion segmentationabstractSkin cancer is one of the most common malignant tumors, and accurately segmenting the lesion area from dermoscopic images is of great clinical significance. With the rapid development of computer technology, Transformer-based models have dominated the field of automatic skin lesion segmentation. However, Transformer-based models typically focus on capturing global dependencies, lacking explicitly encoded convolutional layers as in RNNs or CNNs. This paper proposes a polymorphic multi-head attention aggregation network for skin lesion segmentation (PMAA-Net). It establishes a polymorphic multi-head attention mechanism (PMA), where each self-attention head is designed with convolutional layers that can encode gradients and textures to obtain various local features. Meanwhile, two attention heads are designed to capture spatial and channel-wise features respectively. This allows the model to capture local features and spatial-dimensional features. We further introduce a multi-head attention aggregation module (AGG) to aggregate multiple attention heads. This transforms the attention maps into a high-rank mixture distribution, significantly enhancing the feature representation capability of the attention heads. Our experimental results on three publicly available skin lesion segmentation datasets show that PMAA-Net outperforms other mainstream methods, especially those focusing on extracting local structural features. The codes are available at https://github.com/yuli0501/PMAA-Net. Chaowei Chen, Wenwen Min, Che Zhao, Shunfang Wang |
BIBM | 3 |
| 2024 | Contrastive Masked Graph Autoencoders for Spatial Transcriptomics Data Analysis
Donghai Fang, Yichen Gao, Wenwen Min |
ISBRA (1) | 5 |
| 2024 | Spatial Gene Expression Prediction from Histology Images with STco
Zhiceng Shi, Changmiao Wang, Wenwen Min |
ISBRA (1) | 4 |
| 2024 | stEnTrans: Transformer-Based Deep Learning for Spatial Transcriptomics Enhancement
Shuailin Xue, Changmiao Wang, Wenwen Min |
ISBRA (1) | 4 |
| 2024 | SpaDiT: diffusion transformer for spatial gene expression prediction using scRNA-seqabstractThe rapid development of spatially resolved transcriptomics (SRT) technologies has provided unprecedented opportunities for exploring the structure of specific organs or tissues. However, these techniques (such as image-based SRT) can achieve single-cell resolution, but can only capture the expression levels of tens to hundreds of genes. Such spatial transcriptomics (ST) data, carrying a large number of undetected genes, have limited its application value. To address the challenge, we develop SpaDiT, a deep learning framework for spatial reconstruction and gene expression prediction using scRNA-seq data. SpaDiT employs scRNA-seq data as an a priori condition and utilizes shared genes between ST and scRNA-seq data as latent representations to construct inputs, thereby facilitating the accurate prediction of gene expression in ST data. SpaDiT enhances the accuracy of spatial gene expression predictions over a variety of spatial transcriptomics datasets. We have demonstrated the effectiveness of SpaDiT by conducting extensive experiments on both seq-based and image-based ST data. We compared SpaDiT with eight highly effective baseline methods and found that our proposed method achieved an 8%-12% improvement in performance across multiple metrics. Source code and all datasets used in this paper are available at https://github.com/wenwenmin/SpaDiT and https://zenodo.org/records/12792074. Wenwen Min |
Briefings Bioinform. | 3 |
| 2024 | Multimodal contrastive learning for spatial gene expression prediction using histology imagesabstractIn recent years, the advent of spatial transcriptomics (ST) technology has unlocked unprecedented opportunities for delving into the complexities of gene expression patterns within intricate biological systems. Despite its transformative potential, the prohibitive cost of ST technology remains a significant barrier to its widespread adoption in large-scale studies. An alternative, more cost-effective strategy involves employing artificial intelligence to predict gene expression levels using readily accessible whole-slide images stained with Hematoxylin and Eosin (H&E). However, existing methods have yet to fully capitalize on multimodal information provided by H&E images and ST data with spatial location. In this paper, we propose mclSTExp, a multimodal contrastive learning with Transformer and Densenet-121 encoder for Spatial Transcriptomics Expression prediction. We conceptualize each spot as a "word", integrating its intrinsic features with spatial context through the self-attention mechanism of a Transformer encoder. This integration is further enriched by incorporating image features via contrastive learning, thereby enhancing the predictive capability of our model. We conducted an extensive evaluation of highly variable genes in two breast cancer datasets and a skin squamous cell carcinoma dataset, and the results demonstrate that mclSTExp exhibits superior performance in predicting spatial gene expression. Moreover, mclSTExp has shown promise in interpreting cancer-specific overexpressed genes, elucidating immune-related genes, and identifying specialized spatial domains annotated by pathologists. Our source code is available at https://github.com/shizhiceng/mclSTExp. Wenwen Min, Zhiceng Shi, Jun Wan 0005, Changmiao Wang |
Briefings Bioinform. | 1 |
| 2024 | Precise facial landmark detection by Dynamic Semantic Aggregation Transformer
Jun Wan 0005, Yujia Wu, Zhihui Lai 0001, Wenwen Min, Jun Liu 0036 |
Pattern Recognit. | 5 |
| 2024 | Boundary-Aware Gradient Operator Network for Medical Image SegmentationabstractMedical image segmentation is a crucial task in computer-aided diagnosis. Although convolutional neural networks (CNNs) have made significant progress in the field of medical image segmentation, the convolution kernels of CNNs are optimized from random initialization without explicitly encoding gradient information, leading to a lack of specificity for certain features, such as blurred boundary features. Furthermore, the frequently applied down-sampling operation also loses the fine structural features in shallow layers. Therefore, we propose a boundary-aware gradient operator network (BG-Net) for medical image segmentation, in which the gradient convolution (GConv) and the boundary-aware mechanism (BAM) modules are developed to simulate image boundary features and the remote dependencies between channels. The GConv module transforms the gradient operator into a convolutional operation that can extract gradient features; it attempts to extract more features such as images boundaries and textures, thereby fully utilizing limited input to capture more features representing boundaries. In addition, the BAM can increase the amount of global contextual information while suppressing invalid information by focusing on feature dependencies and the weight ratios between channels. Thus, the boundary perception ability of BG-Net is improved. Finally, we use a multi-modal fusion mechanism to effectively fuse lightweight gradient convolution and U-shaped branch features into a multilevel feature, enabling global dependencies and low-level spatial details to be effectively captured in a shallower manner. We conduct extensive experiments on eight datasets that broadly cover medical images to evaluate the effectiveness of the proposed BG-Net. The experimental results demonstrate that BG-Net outperforms the state-of-the-art methods, particularly those focused on boundary segmentation. Wenwen Min, Shunfang Wang |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | TransVCOX: Bridging Transformer Encoder and Pre-trained VAE for Robust Cancer Multi-Omics Survival AnalysisabstractTraditional survival analysis models, such as the COX proportional hazards model, face challenges in processing multimodal data, identifying nonlinear relationships, and recognizing complex data patterns. The rise of deep learning, particularly Transformers and variational autoencoders (VAEs), has showcased its potential in analyzing cancer multi-omics data comprehensively. However, many individual cancer datasets suffer from limited sample sizes, preventing some deep learning models from extracting in-depth data representations and resulting in subpar performance. To address this issue, we advocate the adoption of pre-training and fine-tuning techniques, which effectively mitigate performance deficits due to sparse cancer data samples. We introduce TransVCOX, a deep survival analysis model integrating a Transformer encoder with VAE. This model leverages pre-training and fine-tuning approaches to predict patients’ survival risk using cancer multi-omics data. Rigorous tests on eight unique cancer datasets from TCGA revealed: (1) TransVCOX outperforms other deep learning and conventional COX models. (2) Pre-training significantly reduces model overfitting and enhances performance. (3) VAE encoding, compared to positional encoding, offers a richer decision-making foundation. (4) The performance boost doesn’t linearly correlate with the addition of Transformer blocks. These findings underline TransVCOX’s promising capability for predicting cancer patients’ survival risks using multi-omics data. The implementation can be accessed at https://github.com/wenwenmin/TransVCOX. Wenwen Min, Shunfang Wang |
BIBM | 2 |
| 2023 | Multimodal attention-based variational autoencoder for clinical risk predictionabstractPrediction of survival risk in cancer patients is crucial for understanding the underlying mechanisms of canceration in different stages. Previous studies mainly relied on single-modal omics data due to technological constraints. However, with the increasing availability of cancer omics data, researchers have focused on the use of multi-omics and multimodal data for survival analysis. The application of deep learning methods has become an option for the prediction of clinical risk. Recent advances in the attention mechanism and the variational autoencoder (VAE) have made them promising for analyzing cancer omics data. However, VAE has limitations in disregarding the importance of different features between modalities, and the introduction of an attention mechanism could address this limitation. In this study, we propose a Multimodal Attention-based VAE (MAVAE) deep learning framework using cross-modal multihead attention to integrate cancer multi-omics data for clinical risk prediction. We evaluated our approach on eight TCGA datasets. We find that (1) MAVAE outperforms traditional machine learning and recent deep learning methods; (2) Multi-modal data yields better classification performance than single-modal data; (3) The multi-head attention mechanism improves the decision-making process; (4) Clinical and genetic data are the most important modal data. Our implementation of MAVAE is available at https://github.com/wenwenmin/MAVAE. Taosheng Xu, Jun Wan 0005, Wenwen Min |
BIBM | 5 |
| 2023 | TsImpute: an accurate two-step imputation method for single-cell RNA-seq dataabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) technology has enabled discovering gene expression patterns at single cell resolution. However, due to technical limitations, there are usually excessive zeros, called "dropouts," in scRNA-seq data, which may mislead the downstream analysis. Therefore, it is crucial to impute these dropouts to recover the biological information. RESULTS: We propose a two-step imputation method called tsImpute to impute scRNA-seq data. At the first step, tsImpute adopts zero-inflated negative binomial distribution to discriminate dropouts from true zeros and performs initial imputation by calculating the expected expression level. At the second step, it conducts clustering with this modified expression matrix, based on which the final distance weighted imputation is performed. Numerical results based on both simulated and real data show that tsImpute achieves favorable performance in terms of gene expression recovery, cell clustering, and differential expression analysis. AVAILABILITY AND IMPLEMENTATION: The R package of tsImpute is available at https://github.com/ZhengWeihuaYNU/tsImpute. Wenwen Min, Shunfang Wang |
Bioinform. | 2 |
| 2023 | Precise Facial Landmark Detection by Reference Heatmap TransformerabstractMost facial landmark detection methods predict landmarks by mapping the input facial appearance features to landmark heatmaps and have achieved promising results. However, when the face image is suffering from large poses, heavy occlusions and complicated illuminations, they cannot learn discriminative feature representations and effective facial shape constraints, nor can they accurately predict the value of each element in the landmark heatmap, limiting their detection accuracy. To address this problem, we propose a novel Reference Heatmap Transformer (RHT) by introducing reference heatmap information for more precise facial landmark detection. The proposed RHT consists of a Soft Transformation Module (STM) and a Hard Transformation Module (HTM), which can cooperate with each other to encourage the accurate transformation of the reference heatmap information and facial shape constraints. Then, a Multi-Scale Feature Fusion Module (MSFFM) is proposed to fuse the transformed heatmap features and the semantic features learned from the original face images to enhance feature representations for producing more accurate target heatmaps. To the best of our knowledge, this is the first study to explore how to enhance facial landmark detection by transforming the reference heatmap information. The experimental results from challenging benchmark datasets demonstrate that our proposed method outperforms the state-of-the-art methods in the literature. Jun Wan 0005, Jun Liu 0036, Jie Zhou 0009, Zhihui Lai 0001, LinLin Shen, Ping Xiong 0001, Wenwen Min |
IEEE Trans. Image Process. | 8 |
| 2023 | Structured Sparse Non-Negative Matrix Factorization With $\ell _{2,0}$ℓ2,0-NormabstractNon-negative matrix factorization (NMF) is a powerful tool for dimensionality reduction and clustering. However, the interpretation of the clustering result from NMF is difficult, especially for the high-dimensional biological data without effective feature selection. To address this problem, we introduce a row-sparse NMF with$\ell _{2,0}$-norm constraint (NMF$\_\ell _{20}$), where the basis matrix$\bm {W}$is constrained by using the$\ell _{2,0}$-norm constraint such that$\bm {W}$has a row-sparsity pattern with feature selection. However, it is a challenge to solve the model, because the$\ell _{2,0}$-norm constraint is a non-convex and non-smooth function. Fortunately, we prove that the$\ell _{2,0}$-norm constraint satisfies the Kurdyka-Łojasiewicz property. Based on this finding, we present a proximal alternating linearized minimization algorithm and its monotone accelerated version to solve the NMF$\_\ell _{20}$model. In addition, we further present a orthogonal NMF with$\ell _{2,0}$-norm constraint (ONMF$\_\ell _{20}$) to enhance the clustering performance by using a non-negative orthogonal constraint. The ONMF$\_\ell _{20}$model is solved by transforming into a series of constrained and penalized matrix factorization problems. The convergence and guarantees for these proposed algorithms are proved and the computational complexity is well evaluated. The results on numerical and scRNA-seq datasets demonstrate the efficiency of our methods in comparison with existing methods. Wenwen Min, Taosheng Xu, Tsung-Hui Chang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | A Novel Sparse Graph-Regularized Singular Value Decomposition Model and Its Application to Genomic Data AnalysisabstractLearning the gene coexpression pattern is a central challenge for high-dimensional gene expression analysis. Recently, sparse singular value decomposition (SVD) has been used to achieve this goal. However, this model ignores the structural information between variables (e.g., a gene network). The typical graph-regularized penalty can be used to incorporate such prior graph information to achieve more accurate discovery and better interpretability. However, the existing approach fails to consider the opposite effect of variables with negative correlations. In this article, we propose a novel sparse graph-regularized SVD model with absolute operator (AGSVD) for high-dimensional gene expression pattern discovery. The key of AGSVD is to impose a novel graph-regularized penalty ($| \boldsymbol {u}|^{T} \boldsymbol {L}| \boldsymbol {u}|$). However, such a penalty is a nonconvex and nonsmooth function, so it brings new challenges to model solving. We show that the nonconvex problem can be efficiently handled in a convex fashion by adopting an alternating optimization strategy. The simulation results on synthetic data show that our method is more effective than the existing SVD-based ones. In addition, the results on several real gene expression data sets show that the proposed methods can discover more biologically interpretable expression patterns by incorporating the prior gene network. Wenwen Min, Tsung-Hui Chang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | TSCCA: A tensor sparse CCA method for detecting microRNA-gene patterns from multiple cancersabstractExisting studies have demonstrated that dysregulation of microRNAs (miRNAs or miRs) is involved in the initiation and progression of cancer. Many efforts have been devoted to identify microRNAs as potential biomarkers for cancer diagnosis, prognosis and therapeutic targets. With the rapid development of miRNA sequencing technology, a vast amount of miRNA expression data for multiple cancers has been collected. These invaluable data repositories provide new paradigms to explore the relationship between miRNAs and cancer. Thus, there is an urgent need to explore the complex cancer-related miRNA-gene patterns by integrating multi-omics data in a pan-cancer paradigm. In this study, we present a tensor sparse canonical correlation analysis (TSCCA) method for identifying cancer-related miRNA-gene modules across multiple cancers. TSCCA is able to overcome the drawbacks of existing solutions and capture both the cancer-shared and specific miRNA-gene co-expressed modules with better biological interpretations. We comprehensively evaluate the performance of TSCCA using a set of simulated data and matched miRNA/gene expression data across 33 cancer types from the TCGA database. We uncover several dysfunctional miRNA-gene modules with important biological functions and statistical significance. These modules can advance our understanding of miRNA regulatory mechanisms of cancer and provide insights into miRNA-based treatments for cancer. Wenwen Min, Tsung-Hui Chang |
PLoS Comput. Biol. | 1 |
| 2021 | Group-Sparse SVD Models via $L_1$L1- and $L_0$L0-norm Penalties and their Applications in Biological DataabstractSparse Singular Value Decomposition (SVD) models have been proposed for biclustering high dimensional gene expression data to identify block patterns with similar expressions. However, these models do not take into account prior group effects upon variable selection. To this end, we first propose group-sparse SVD models with group Lasso (GL1-SVD) and group L0-norm penalty (GL0-SVD) for non-overlapping group structure of variables. However, such group-sparse SVD models limit their applicability in some problems with overlapping structure. Thus, we also propose two group-sparse SVD models with overlapping group Lasso (OGL1-SVD) and overlapping group L0-norm penalty (OGL0-SVD). We first adopt an alternating iterative strategy to solve GL1-SVD based on a block coordinate descent method, and GL0-SVD based on a projection method. The key of solving OGL1-SVD is a proximal operator with overlapping group Lasso penalty. We employ an alternating direction method of multipliers (ADMM) to solve the proximal operator. Similarly, we develop an approximate method to solve OGL0-SVD. Applications of these methods and comparison with competing ones using simulated data demonstrate their effectiveness. Extensive applications of them onto several real gene expression data with gene prior group knowledge identify some biologically interpretable gene modules. Wenwen Min, Juan Liu 0007 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Gene Functional Module Discovery via Integrating Gene Expression and PPI Network Data
Juan Liu 0007, Wenwen Min |
ICIC (2) | 3 |
| 2018 | Edge-group sparse PCA for network-guided high dimensional data analysisabstractMotivation: Principal component analysis (PCA) has been widely used to deal with high-dimensional gene expression data. In this study, we proposed an Edge-group Sparse PCA (ESPCA) model by incorporating the group structure from a prior gene network into the PCA framework for dimension reduction and feature interpretation. ESPCA enforces sparsity of principal component (PC) loadings through considering the connectivity of gene variables in the prior network. We developed an alternating iterative algorithm to solve ESPCA. The key of this algorithm is to solve a new k-edge sparse projection problem and a greedy strategy has been adapted to address it. Here we adopted ESPCA for analyzing multiple gene expression matrices simultaneously. By incorporating prior knowledge, our method can overcome the drawbacks of sparse PCA and capture some gene modules with better biological interpretations. Results: We evaluated the performance of ESPCA using a set of artificial datasets and two real biological datasets (including TCGA pan-cancer expression data and ENCODE expression data), and compared their performance with PCA and sparse PCA. The results showed that ESPCA could identify more biologically relevant genes, improve their biological interpretations and reveal distinct sample characteristics. Availability and implementation: An R package of ESPCA is available at http://page.amss.ac.cn/shihua.zhang/. Supplementary information: Supplementary data are available at Bioinformatics online. Wenwen Min, Juan Liu 0007 |
Bioinform. | 1 |
| 2018 | Network-Regularized Sparse Logistic Regression Models for Clinical Risk Prediction and Biomarker DiscoveryabstractMolecular profiling data (e.g., gene expression) has been used for clinical risk prediction and biomarker discovery. However, it is necessary to integrate other prior knowledge like biological pathways or gene interaction networks to improve the predictive ability and biological interpretability of biomarkers. Here, we first introduce a general regularized Logistic Regression (LR) framework with regularized term , which can reduce to different penalties, including Lasso, elastic net, and network-regularized terms with different . This framework can be easily solved in a unified manner by a cyclic coordinate descent algorithm which can avoid inverse matrix operation and accelerate the computing speed. However, if those estimated and have opposite signs, then the traditional network-regularized penalty may not perform well. To address it, we introduce a novel network-regularized sparse LR model with a new penalty to consider the difference between the absolute values of the coefficients. We develop two efficient algorithms to solve it. Finally, we test our methods and compare them with the related ones using simulated and real data to show their efficiency. Wenwen Min, Juan Liu 0007 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2015 | A novel two-stage method for identifying microRNA-gene regulatory modules in breast cancerabstractIn this paper, we propose a two-stage method for identifying miRNA-gene regulatory modules by integrating miRNA/mRNA expression profiles and miRNA genomic cluster data. We first adopt a Multiple-output Sparse Group Lasso (MSGL) regression model to predict the miRNA-gene regulatory network. Further, we propose a L0-penalized Singular Value Decomposition (L0-SVD) model to identify modules from the predicted network. We apply this method to miRNA and mRNA expression profiles of the breast cancer data from TCGA databases and identify ten miRNA-gene regulatory modules. We find that (1) the modules are significantly associated in a predicted miRNA-gene regulatory network; (2) the modules are significantly enriched in GO biological processes and KEGG pathways, respectively; (3) many miRNAs and genes in the modules are related with breast cancer. On average, 51% of the miRNAs and 30% of the genes are related with breast cancer. The results demonstrate that miRNA-gene regulatory modules provide insights into the mechanisms of the combinatorial regulation between miRNAs and genes. Wenwen Min, Juan Liu 0007, Fei Luo 0004 |
BIBM | 1 |