Wenju Hou

dblp:328/5707 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0001-9247-8324ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A pre-trained language model-based cross-modal fusion framework for predicting miRNA-drug resistance and sensitivity associations
abstract
MicroRNAs (miRNAs) are pivotal regulators of drug resistance and sensitivity in cancer cells, functioning as tumor suppressors or oncogenes that modulate the cellular response to anticancer drugs. While experimental identification of miRNA-mediated drug resistance and sensitivity is both costly and laborious, computational methods present a promising alternative. Recent advances in pre-trained language models (PLMs) offer new opportunities to leverage large-scale unlabeled biomolecular data for enhanced relationship prediction. In this study, we introduce PLMF-MDA, a PLM-based cross-modal fusion model designed to predict miRNA-drug resistance (MDR) and miRNA-drug sensitivity (MDS) associations. PLMF-MDA integrates miRNA and drug multimodal embeddings derived from PLMs and intrinsic feature extractors, and employs a cross-modal attention fusion module to adaptively capture key interactions between modalities. To evaluate the performance of the approach, we manually constructed two benchmark datasets. Experimental results demonstrate that the PLMF-MDA achieves superior prediction performance. Furthermore, case studies on anticancer drug docetaxel and gefitinib demonstrate its potential in discovering novel MDR (MDS) associations. All data and source code are available on GitHub: https://github.com/sheng-n/PLMF-MDA.
Nan Sheng, Yun-Zhi Liu, Wenju Hou, Lan Huang 0002, Yan Wang 0028
PLoS Comput. Biol.4
2025 GSAM-MRI: Frequency-Based Domain Randomization for Generalized MR Image Segmentation with Segment Anything Model
abstract
Magnetic resonance imaging (MRI) data segmentation plays a critical role in clinical diagnosis and treatment planning. However, the performance of deep learning-based segmentation models is often hindered by domain shifts caused by variations in imaging factors. To address this challenge, we propose GSAM-MRI, a Generalized Segment Anything Model for robust MRI segmentation in the scenario of single-source domain generalization (SDG). GSAM-MRI integrates multiple components to enhance domain generalization: (1) Frequency-based Domain Randomization module that simulates inter-site variability by perturbing the frequency domain; (2) Domain Adversarial Block that promotes domain-invariant feature learning through adversarial training; (3) General Embedding Generator that fuses multi-scale hierarchical features to produce dense prompt embeddings; Additionally, a hybrid loss function is employed for output consistency. Experiments on prostate segmentation and white matter hyperintensity segmentation tasks demonstrate that GSAM-MRI consistently outperforms state-of-the-art SDG methods and baseline models, achieving superior generalization across unseen domains.
Lan Huang 0002, Yinglu Sun, Xinfei Wang 0001, Qixing Yang, Xuping Xie, Wenju Hou, Chunjie Guo, Yan Wang 0028
BIBM7
2025 HATZFS predicts pancreatic cancer driver biomarkers by hierarchical reinforcement learning and zero-forcing set
Wenju Hou, Nan Sheng, Chunman Zuo, Yan Wang 0028
Expert Syst. Appl.2
2025 PCG-CAM: Enhanced class activation map using principal components of gradients and its applications in brain MRI
Lan Huang 0002, Yangguang Shao, Wenju Hou, Yan Wang 0028, Nan Sheng, Yinglu Sun, Yao Wang 0010
Inf. Sci.3
2025 Self-Supervised Contrastive Learning on Attribute and Topology Graphs for Predicting Relationships Among lncRNAs, miRNAs and Diseases
abstract
Exploring associations between long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is crucial for disease prevention, diagnosis and treatment. While determining these relationships experimentally is resource-intensive and time-consuming, computational methods have emerged as an attractive way. However, existing computational methods tend to focus on single tasks, neglecting the benefits of leveraging multiple biomolecular interactions and domain-specific knowledge for multi-task prediction. Furthermore, the scarcity of labeled data for lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) poses challenges for comprehensive node embedding learning. This paper proposes a multi-task prediction model (called SSCLMD) that employs self-supervised contrastive learning on attribute and topology graphs to identify potential LDAs, MDAs and LMIs. Firstly, domain knowledge of lncRNAs, miRNAs and diseases as well as their interactions are exploited to construct attribute graph and topology graph, respectively. Then, the nodes are encoded in the attribute and topology spaces to extract the specific and common feature. Meanwhile, the attention mechanism is performed to adaptively fuse the embedding from different views. SSCLMD incorporates contrastive self-supervised learning as a regularize to guide node embedding learning in both attribute and topology space without relying on labels. Severing as a regularize in multi-task learning paradigm, it to improves the model.s generalization capabilities. Extensive experiments on 2 manually curated datasets demonstrate that SSCLMD significantly outperforms baseline methods in LDA, MDA and LMI prediction tasks. Case studies on both old and new datasets further supported SSCLMD's ability to uncover novel disease-related lncRNAs and miRNAs.
Lan Huang 0002, Nan Sheng, Lei Wang 0121, Wenju Hou, Yan Wang 0028
IEEE J. Biomed. Health Informatics5
2024 Multi-view learning framework for predicting unknown types of cancer markers via directed graph neural networks fitting regulatory networks
abstract
The discovery of diagnostic and therapeutic biomarkers for complex diseases, especially cancer, has always been a central and long-term challenge in molecular association prediction research, offering promising avenues for advancing the understanding of complex diseases. To this end, researchers have developed various network-based prediction techniques targeting specific molecular associations. However, limitations imposed by reductionism and network representation learning have led existing studies to narrowly focus on high prediction efficiency within single association type, thereby glossing over the discovery of unknown types of associations. Additionally, effectively utilizing network structure to fit the interaction properties of regulatory networks and combining specific case biomarker validations remains an unresolved issue in cancer biomarker prediction methods. To overcome these limitations, we propose a multi-view learning framework, CeRVE, based on directed graph neural networks (DGNN) for predicting unknown type cancer biomarkers. CeRVE effectively extracts and integrates subgraph information through multi-view feature learning. Subsequently, CeRVE utilizes DGNN to simulate the entire regulatory network, propagating node attribute features and extracting various interaction relationships between molecules. Furthermore, CeRVE constructed a comparative analysis matrix of three cancers and adjacent normal tissues through The Cancer Genome Atlas and identified multiple types of potential cancer biomarkers through differential expression analysis of mRNA, microRNA, and long noncoding RNA. Computational testing of multiple types of biomarkers for 72 cancers demonstrates that CeRVE exhibits superior performance in cancer biomarker prediction, providing a powerful tool and insightful approach for AI-assisted disease biomarker discovery.
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Wenju Hou
Briefings Bioinform.8
2024 MMGAT: a graph attention network framework for ATAC-seq motifs finding
abstract
BACKGROUND: Motif finding in Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) data is essential to reveal the intricacies of transcription factor binding sites (TFBSs) and their pivotal roles in gene regulation. Deep learning technologies including convolutional neural networks (CNNs) and graph neural networks (GNNs), have achieved success in finding ATAC-seq motifs. However, CNN-based methods are limited by the fixed width of the convolutional kernel, which makes it difficult to find multiple transcription factor binding sites with different lengths. GNN-based methods has the limitation of using the edge weight information directly, makes it difficult to aggregate the neighboring nodes' information more efficiently when representing node embedding. RESULTS: To address this challenge, we developed a novel graph attention network framework named MMGAT, which employs an attention mechanism to adjust the attention coefficients among different nodes. And then MMGAT finds multiple ATAC-seq motifs based on the attention coefficients of sequence nodes and k-mer nodes as well as the coexisting probability of k-mers. Our approach achieved better performance on the human ATAC-seq datasets compared to existing tools, as evidenced the highest scores on the precision, recall, F1_score, ACC, AUC, and PRC metrics, as well as finding 389 higher quality motifs. To validate the performance of MMGAT in predicting TFBSs and finding motifs on more datasets, we enlarged the number of the human ATAC-seq datasets to 180 and newly integrated 80 mouse ATAC-seq datasets for multi-species experimental validation. Specifically on the mouse ATAC-seq dataset, MMGAT also achieved the highest scores on six metrics and found 356 higher-quality motifs. To facilitate researchers in utilizing MMGAT, we have also developed a user-friendly web server named MMGAT-S that hosts the MMGAT method and ATAC-seq motif finding results. CONCLUSIONS: The advanced methodology MMGAT provides a robust tool for finding ATAC-seq motifs, and the comprehensive server MMGAT-S makes a significant contribution to genomics research. The open-source code of MMGAT can be found at https://github.com/xiaotianr/MMGAT , and MMGAT-S is freely available at https://www.mmgraphws.com/MMGAT-S/ .
Wenju Hou, Lan Huang 0002, Nan Sheng, Qixing Yang, Shuangquan Zhang, Yan Wang 0028
BMC Bioinform.2
2022 Leveraging multidimensional features for policy opinion sentiment prediction
Wenju Hou, Ying Li 0004
Inf. Sci.1