EDBT 2026 Demo / reviewers in the wild / expert
Weidun Xie
dblp:289/8072
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-9009-2647ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DDintensity: Addressing imbalanced drug-drug interaction risk levels using pre-trained deep learning model embeddings
Weidun Xie, Xingjian Chen, Zetian Zheng, Ruoxuan Zhang, Chengbin Peng 0001, Monika Gullerova, Ka-Chun Wong |
Artif. Intell. Medicine | 1 |
| 2025 | cis-positional information in regulatory single nucleotide variation prioritizationabstractAbstract Duttke et al. have proved and generalized past observations on the positional preferences of regulatory genomics at multiple functional levels in July 2024 [1]. However, the explicit open-box distribution learning on those positional preferences are under-explored in the existing gene regulation methods including deep learning. Contributing towards such directions, we propose to develop regulatory positional distribution models. The positional distribution models can capture the spatial features of gene transcription which can improve different downstream applications such as rSNV prioritization. Experiments have been conducted to substantiate its claim in deleterious rSNV predictions of ClinVar. In particular, we have collected the ClinVar dataset (i.e., ‘variant_summary.txt.gz’ on 2024-09-18) and retrieved all deleterious SNVs (i.e., labelled as ‘Pathogenic’ and ‘Likely pathogenic’ in the ‘ClinicalSigificance’ column) around all human TSS locations from Ensembl (i.e., Ensembl Genes 112). In particular, it was surprising that, although CADD is already an ensemble approach built upon different state-of-the-arts methods [2], CADD can still be improved with statistical significance after cis-positional information has been incorporated across different situations where evolutionary conservation signals (PhastCons and PhyloP) have been integrated. Based on the results, we propose two future research directions. The first direction is to examine different statistical distributions for position-aware gene regulation modelling while the second direction is to incorporate those distributions into different downstream applications such as eQTL analysis and deleterious rSNV predictions. The outcomes will have broad implications across different downstream gene regulation modelling studies. [1] Duttke S.H. et al. ‘Position-dependent function of human sequence-specific transcription factors.’ Nature 2024;631:891–898. [2] Rentzsch, P. et al. Nucleic acids research 2019:47(D1):D886–D894. Ka-Chun Wong, Zhongyu Yao, Weidun Xie, Zhongshen Li, Tianchi Lu, Cho Ling |
Briefings Bioinform. | 3 |
| 2025 | When East Meets West: Cross-Domain Drug Interaction Annotations With Large Language Models and Bidirectional Neural NetworksabstractDrug combination therapy is a promising strategy for managing complex and co-existing diseases. However, drug-drug interactions (DDIs) can result in unexpected adverse effects, making it crucial to understand such interactions to prevent adverse drug reactions and develop new therapeutic strategies. Current DDI annotation methods heavily rely on atom-level graph structural features, overlooking valuable drug contextual representations within medical literature. Additionally, these methods are typically designed for a specific task, limiting their scalability to diverse medical scenarios. To address these limitations, we propose TEmbed-DDI, a novel framework that leverages contextual representations and pre-trained large language model embeddings to enhance feature extraction for DDI annotations. Specifically, we retrieve meaningful contextual texts for each drug to enrich semantic features and adopt pre-trained large language model embeddings to capture rich features from these long-range contextual representations. TEmbed-DDI is the first framework to incorporate LLM-powered embeddings for medical interaction annotations. Furthermore, a bidirectional neural network is integrated into TEmbed-DDI for the integrative Western and traditional Chinese medicine DDI annotation tasks. Comparative results demonstrate that TEmbed-DDI achieves state-of-the-art performance, with the highest AUC scores of 0.992 and 0.95 on the Western CHCH and DEEP interaction annotation benchmarks. Even for the newly constructed Traditional Chinese Medicine (TCM) DDI annotation benchmark, TEmbed-DDI consistently exhibits outstanding generalization capability, achieving an AUC of 0.956. Moreover, case studies further validate TEmbed-DDI's capability to annotate previously unknown interactions. These findings suggest that TEmbed-DDI can serve as a valuable tool in annotating previously unknown drug combinations for real-world applications, facilitating the development of efficacious therapies. Furthermore, as the first framework combining traditional Chinese medicine into DDI annotation tasks, its adaptability highlights the potential in supporting cross-domain medical research. TEmbed-DDI's design principles can inspire the development of flexible LLM-powered frameworks for drug combination discovery in the future. Ruoxuan Zhang, Weidun Xie, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | scPER2P: Parameter-Efficient Single-Cell LLM for Translated Proteome Profiles
Xingjian Chen, Zetian Zheng, Weidun Xie, Fuzhou Wang, Ka-Chun Wong |
ICONIP (5) | 4 |
| 2024 | HE2Gene: image-to-RNA translation via multi-task learning for spatial transcriptomics dataabstractMOTIVATION: Tissue context and molecular profiling are commonly used measures in understanding normal development and disease pathology. In recent years, the development of spatial molecular profiling technologies (e.g. spatial resolved transcriptomics) has enabled the exploration of quantitative links between tissue morphology and gene expression. However, these technologies remain expensive and time-consuming, with subsequent analyses necessitating high-throughput pathological annotations. On the other hand, existing computational tools are limited to predicting only a few dozen to several hundred genes, and the majority of the methods are designed for bulk RNA-seq. RESULTS: In this context, we propose HE2Gene, the first multi-task learning-based method capable of predicting tens of thousands of spot-level gene expressions along with pathological annotations from H&E-stained images. Experimental results demonstrate that HE2Gene is comparable to state-of-the-art methods and generalizes well on an external dataset without the need for re-training. Moreover, HE2Gene preserves the annotated spatial domains and has the potential to identify biomarkers. This capability facilitates cancer diagnosis and broadens its applicability to investigate gene-disease associations. AVAILABILITY AND IMPLEMENTATION: The source code and data information has been deposited at https://github.com/Microbiods/HE2Gene. Xingjian Chen, Jiecong Lin, Weidun Xie, Zetian Zheng, Ka-Chun Wong |
Bioinform. | 5 |
| 2022 | SRG-Vote: Predicting Mirna-Gene Relationships via Embedding and LSTM EnsembleabstractTargeted therapy for one for a set of genes has made it possible to apply precision medicine for different patients due to the existence of tumor heterogeneity. However, how to regulate those genes are still problematic. One of the natural regulators of genes is microRNAs. Thus, a better understanding of the miRNA-gene interaction mechanism might contribute to future diagnosis, prevention, and cancer therapy. The interactions between microRNA and genes play an essential role in molecular genetics. The in-vivo experiments validating the relationships between them are time-consuming, money-costly, and labor-intensive. With the development of high-throughput technology, we dealt with tons of biological data. However, extracting features from tremendous raw data and making a mathematical model is still a challenging topic. Machine learning and deep learning algorithms have become powerful tools in dealing with biological data. Inspired by this, in this paper, we propose a model that combines features/embedding extraction methods, deep learning algorithms, and a voting system. We leverage doc2vec to generate sequential embedding from molecular sequences. The role2vec, GCN, and GMM for geometrical embedding were generated from the complex network from similarity and pair-wise datasets. For the deep learning algorithms, we leveraged LSTM and Bi-LSTM according to different embedding and features. Finally, we adopted a voting system to balance results from different data sources. The results have shown that our voting system could achieve a higher AUC than the existing benchmark. The case studies demonstrate that our model could reveal potential relationships between miRNAs and genes. The source code, features, and predictive results can be downloaded at https://github.com/Xshelton/SRG-vote. Weidun Xie, Zetian Zheng, Qiuzhen Lin, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Subclass-Specific Prognosis and Treatment Efficacy Inference in Head and Neck Squamous CarcinomaabstractExploring the prognostic classification and biomarkers in Head and Neck Squamous Carcinoma (HNSC) is of great clinical significance. We hybridized three prominent strategies to comprehensively characterize the molecular features of HNSC. We constructed a 15-gene signature to predict patients' death risk with an average AUC of 0.744 for 1-, 3-, and 5-year on TCGA-HNSC training set, and average AUCs of 0.636, 0.584, 0.755 in GSE65858, GSE-112026, CPTAC-HNSCC datasets, respectively. By combined with NMF clustering and consensus clustering of fraction of tumor immune cell infiltration (ICI) in the tumor microenvironment (TME), we captured a more refined biological characteristics of HNSC, and observed a prognosis heterogeneity in high tumor immunity patients. By matching tumor subset-specific expression signatures to drug-induced cell line expression profiles from large-scale pharmacogenomic databases in the OCTAD workspace, we identified a group of HNSC patients featured with poor prognosis and demonstrated that the individuals in this group are likely to receive increased drug sensitivity to reverse differentially expressed disease signature genes. This trend is especially highlighted among those with higher death risk and tumour immunity. Zetian Zheng, Weidun Xie, Xingjian Chen, Fuzhou Wang, Xiangtao Li, Qiuzhen Lin, Ka-Chun Wong |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Metric Learning Based Vision Transformer for Product Matching
Wei Shao 0009, Fuzhou Wang, Weidun Xie, Ka-Chun Wong |
ICONIP (1) | 4 |
| 2021 | SG-LSTM-FRAME: a computational frame using sequence and geometrical information via LSTM to predict miRNA-gene associationsabstractMOTIVATION: MircroRNAs (miRNAs) regulate target genes and are responsible for lethal diseases such as cancers. Accurately recognizing and identifying miRNA and gene pairs could be helpful in deciphering the mechanism by which miRNA affects and regulates the development of cancers. Embedding methods and deep learning methods have shown their excellent performance in traditional classification tasks in many scenarios. But not so many attempts have adapted and merged these two methods into miRNA-gene relationship prediction. Hence, we proposed a novel computational framework. We first generated representational features for miRNAs and genes using both sequence and geometrical information and then leveraged a deep learning method for the associations' prediction. RESULTS: We used long short-term memory (LSTM) to predict potential relationships and proved that our method outperformed other state-of-the-art methods. Results showed that our framework SG-LSTM got an area under curve of 0.94 and was superior to other methods. In the case study, we predicted the top 10 miRNA-gene relationships and recommended the top 10 potential genes for hsa-miR-335-5p for SG-LSTM-core. We also tested our model using a larger dataset, from which 14 668 698 miRNA-gene pairs were predicted. The top 10 unknown pairs were also listed. AVAILABILITY: Our work can be download in https://github.com/Xshelton/SG_LSTM. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online. Weidun Xie, Jiawei Luo 0001, Chu Pan, Ying Liu 0027 |
Briefings Bioinform. | 1 |