VLDB 2026 Research / reviewers in the wild / expert
Lianlian Wu
dblp:286/6216
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-9611-4488ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MOCT: A Multi-Class Oblique Tree Algorithm for Synergistic Drug Combination PredictionabstractMachine learning has been successfully applied to drug combination prediction in recent years. However, in some situations, the class imbalance problem still shows highly negative impacts on the modeling process, which cannot be directly handled by traditional methods. In addition, the interpretability of models is another key point for biological and medical experts. In this study, a clustering-based oblique decision tree (MOCT) algorithm is proposed to extract interpretable knowledge for the multi-class datasets. It firstly clusters samples of different classes, and then a proper feature subspace is generated to split data and forms a nonleaf node. Unlike traditional decision trees, our MOCT only grows one none-leaf node in each layer to generate a concise tree structure. Datasets of drug combinations were collected from three cell lines with three classes (Additive, Antagonism, and Synergy) in experiments, and the results show that our MOCT algorithm is superior to other methods with better interpretability. Zhikai Lin, Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | EndoCADx: A Real-Time LVLM-Based CADx System for Multimodal Diagnosis of Gastric Lesions in White-Light EndoscopyabstractAccurate characterization and detailed documentation of gastric lesions during endoscopy are critical for early diagnosis and effective patient management. However, conventional computer-aided detection (CAD) systems primarily focus on lesion detection and lack the ability to generate comprehensive semantic descriptions, which limits their clinical utility. To address this gap, we have developed EndoCADx, a real-time computer-aided diagnosis (CADx) system integrating the Qwen2-VL large vision-language model. It was specifically designed for the real-time detection and description of gastric focal lesions in white-light endoscopic images. To fine-tune and evaluate the model, we constructed a domain-specific multimodal dataset comprising 7,543 expert-annotated image-text pairs across five common lesion types, named EndoGastro-7k. EndoCADx achieved a lesion detection accuracy of 91.6% and an accuracy in describing five key lesion characteristics of 84.7%, outperforming several state-of-the-art large vision-language models (LVLMs). Clinical evaluations further demonstrated its practical utility, with 86% of physicians endorsing its diagnostic completeness and 90% expressing a willingness to integrate it into routine practice. EndoCADx represents the first LVLM-based real-time CADx system tailored to gastric lesion analysis. It offers enhanced diagnostic accuracy, richer lesion interpretation, and improved support for endoscopic decision-making. Tingshun Xiong, Changhong Xiao, Ruiqing Jiang, Bin Mei, Lianlian Wu, Zhongyuan Wang 0001 |
BIBM | 8 |
| 2025 | Multi-task multi-view and iterative error-correcting random forest for acute toxicity prediction
Lianlian Wu, Guangyi Lin, Jiayu Zou, Bowei Yan, Kunhong Liu 0001, Xiaochen Bo |
Expert Syst. Appl. | 2 |
| 2025 | A feature pair-based neural network embedded decision tree for synergistic drug combination prediction
Jiayu Zou, Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo |
Pattern Recognit. | 2 |
| 2024 | CPDT: A Novel Cluster-based Paired Decision Tree for Identifying Biomedical Entity InteractionsabstractFor the interaction prediction task in the biomedical field, most machine learning algorithms overlook the relationships between entities within a pair by treating their features independently. To address this issue, this paper proposes a novel Cluster-based Paired Decision Tree model (CPDT), which pairs synonymous features of entity pairs to form paired feature spaces for simultaneous processing. It employs an adaptive grid-based clustering algorithm to partition these spaces in an axis-parallel manner, constructing interpretable decision boundaries. Moreover, the clustering algorithm leverages the probability density function to accommodate various data distributions in paired feature spaces, enhancing the effectiveness of sample partitioning. Experimental results demonstrate that CPDT performs well in two interaction prediction tasks: Drug Combination and Synthetic Lethality predictions. Furthermore, CPDT yields simple and interpretable decision rules that uncover potential patterns in biomedical interaction prediction. It also identifies molecules with medical significance, suggesting promising applications in the biomedical domain. Jiayu Zou, Lianlian Wu, Weiping Lin, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo |
BIBM | 2 |
| 2023 | A Multi-View Learning-Based Bayesian Ruleset Extraction Algorithm For Accurate Hepatotoxicity PredictionabstractThe interpretable machine learning method is important in drug discovery. Unlike traditional ensemble learning methods, this paper proposes an interpretable algorithm based on Bayesian rule extraction to obtain reliable and explainable results for hepatotoxicity prediction. To extract information from different types of omics data, our algorithm employs a multi-view learning strategy to enhance performance. Specifically, a random forest is trained in each view, and then the Bayesian rule extraction algorithm is designed to select an optimal rule subset, controlling the size and accuracy of the ruleset through probabilities. These rule sets are integrated through multi-view voting to get the final decisions. The performance of our algorithm is tested on the hepatotoxicity dataset, demonstrating that compared to traditional machine learning algorithms and rule-based algorithms, our approach maintains excellent performance while achieving high interpretability in most cases. Our python source code and the related Supplementary Materials are available at: github.com/MLDMXM2017/MV-BRS. Lianlian Wu, Yong Xu 0009, Kunhong Liu 0001, Xiaochen Bo |
BIBM | 2 |
| 2023 | Drug-target and Drug-disease Association Prediction based on Drug-target-disease Network and Multi-task LearningabstractTraditional drug-target and drug-disease associations prediction tasks have been performed independently, without fully exploiting the relationships between drugs and various other entities, leading to inaccurate predictions. With the emergence of large-scale heterogeneous biological networks, multi-task learning can effectively enhance the accuracy of association prediction based on the associations between entities. In this study, we propose a multi-task learning framework named DTD-MTL to predict drug-target and drug-disease associations simultaneously. Firstly, it utilizes a multi-layer relational graph convolutional network (RGCN) to learn the features of each node in the drug-target-disease network. Subsequently, it obtains the initial feature of an edge by concatenating the features of the two nodes on the same edge. To coordinate different prediction tasks, drug features are shared among different tasks. Afterwards, the autoencoder (AE) is used to extract features from different types of edges. In order to make the learned edge features more suitable for different prediction tasks, the distance covariance (DC) is utilized to eliminate the specificity between different types of edges, thereby leveraging the relationships between different tasks more effectively. Finally, the drug-target and drug-disease associations predictions are achieved based on the edge features extracted by the AE. Experimental results on a widely-used dataset show that DTD-MTL outperforms the state-of-the-art methods in the prediction task of drug-target and drug-disease associations. Binyu Wang, Hongyan Ye, Lianlian Wu, Xiaochen Bo, Zhongnan Zhang |
BIBM | 4 |
| 2023 | HSGCL-DTA: Hybrid-scale Graph Contrastive Learning based Drug-Target Binding Affinity PredictionabstractDrug-target binding affinity (DTA) is a critical criterion for drug screening. Accurate affinity prediction will significantly cut the cost of new drug development and accelerate the drug discovery process. However, most existing approaches frequently utilize sequence or structure information without incorporating any additional information. At the same time, they encode drugs and targets separately, ignoring the important existing drug-target relationships. In this study, we propose a novel DTA prediction approach, named HSGCL-DTA, which is based on hybrid-scale graph contrastive learning. To completely capture the global information and discriminative properties of the heterogeneous graphs, HSGCL-DTA divides the drug-target affinity graph into two subgraphs with stronger and weaker affinities respectively, and the node embeddings of the two subgraphs are obtained based on node-graph level contrastive learning. Afterwards, graph convolutional network (GCN) is used to encode the molecular graph of drugs and targets, and the node embeddings in the strong affinity subgraph are fused with molecule graph embeddings to fully utilize the distinct information in two different views. Another node-node level contrastive learning is performed between the affinity graph and molecular graphs, thereby filtering out task-independent noise that only appears in one graph. The final drug-target embeddings are put into a multilayer perceptron (MLP) for affinity prediction. Experiments on two widely-used datasets have shown that HSGCL-DTA achieves better prediction performance and generalization than the state-of-the-art DTA prediction methods. Hongyan Ye, Yingying Song, Binyu Wang, Lianlian Wu, Xiaochen Bo, Zhongnan Zhang |
ICTAI | 4 |
| 2023 | EDST: a decision stump based ensemble algorithm for synergistic drug combination predictionabstractINTRODUCTION: There are countless possibilities for drug combinations, which makes it expensive and time-consuming to rely solely on clinical trials to determine the effects of each possible drug combination. In order to screen out the most effective drug combinations more quickly, scholars began to apply machine learning to drug combination prediction. However, most of them are of low interpretability. Consequently, even though they can sometimes produce high prediction accuracy, experts in the medical and biological fields can still not fully rely on their judgments because of the lack of knowledge about the decision-making process. RELATED WORK: Decision trees and their ensemble algorithms are considered to be suitable methods for pharmaceutical applications due to their excellent performance and good interpretability. We review existing decision trees or decision tree ensemble algorithms in the medical field and point out their shortcomings. METHOD: This study proposes a decision stump (DS)-based solution to extract interpretable knowledge from data sets. In this method, a set of DSs is first generated to selectively form a decision tree (DST). Different from the traditional decision tree, our algorithm not only enables a partial exchange of information between base classifiers by introducing a stump exchange method but also uses a modified Gini index to evaluate stump performance so that the generation of each node is evaluated by a global view to maintain high generalization ability. Furthermore, these trees are combined to construct an ensemble of DST (EDST). EXPERIMENT: The two-drug combination data sets are collected from two cell lines with three classes (additive, antagonistic and synergistic effects) to test our method. Experimental results show that both our DST and EDST perform better than other methods. Besides, the rules generated by our methods are more compact and more accurate than other rule-based algorithms. Finally, we also analyze the extracted knowledge by the model in the field of bioinformatics. CONCLUSION: The novel decision tree ensemble model can effectively predict the effect of drug combination datasets and easily obtain the decision-making process. Lianlian Wu, Kunhong Liu 0001, Yong Xu 0009, Xiaochen Bo |
BMC Bioinform. | 2 |
| 2022 | An enhanced cascade-based deep forest model for drug combination predictionabstractCombination therapy has shown an obvious curative effect on complex diseases, whereas the search space of drug combinations is too large to be validated experimentally even with high-throughput screens. With the increase of the number of drugs, artificial intelligence techniques, especially machine learning methods, have become applicable for the discovery of synergistic drug combinations to significantly reduce the experimental workload. In this study, in order to predict novel synergistic drug combinations in various cancer cell lines, the cell line-specific drug-induced gene expression profile (GP) is added as a new feature type to capture the cellular response of drugs and reveal the biological mechanism of synergistic effect. Then, an enhanced cascade-based deep forest regressor (EC-DFR) is innovatively presented to apply the new small-scale drug combination dataset involving chemical, physical and biological (GP) properties of drugs and cells. Verified by the dataset, EC-DFR outperforms two state-of-the-art deep neural network-based methods and several advanced classical machine learning algorithms. Biological experimental validation performed subsequently on a set of previously untested drug combinations further confirms the performance of EC-DFR. What is more prominent is that EC-DFR can distinguish the most important features, making it more interpretable. By evaluating the contribution of each feature type, GP feature contributes 82.40%, showing the cellular responses of drugs may play crucial roles in synergism prediction. The analysis based on the top contributing genes in GP further demonstrates some potential relationships between the transcriptomic levels of key genes under drug regulation and the synergism of drug combinations. Weiping Lin, Lianlian Wu, Yuqi Wen, Bowei Yan, Chong Dai, Kunhong Liu 0001, Xiaochen Bo |
Briefings Bioinform. | 2 |
| 2022 | Computational methods, databases and tools for synthetic lethality predictionabstractSynthetic lethality (SL) occurs between two genes when the inactivation of either gene alone has no effect on cell survival but the inactivation of both genes results in cell death. SL-based therapy has become one of the most promising targeted cancer therapies in the last decade as PARP inhibitors achieve great success in the clinic. The key point to exploiting SL-based cancer therapy is the identification of robust SL pairs. Although many wet-lab-based methods have been developed to screen SL pairs, known SL pairs are less than 0.1% of all potential pairs due to large number of human gene combinations. Computational prediction methods complement wet-lab-based methods to effectively reduce the search space of SL pairs. In this paper, we review the recent applications of computational methods and commonly used databases for SL prediction. First, we introduce the concept of SL and its screening methods. Second, various SL-related data resources are summarized. Then, computational methods including statistical-based methods, network-based methods, classical machine learning methods and deep learning methods for SL prediction are summarized. In particular, we elaborate on the negative sampling methods applied in these models. Next, representative tools for SL prediction are introduced. Finally, the challenges and future work for SL prediction are discussed. Junshan Han, Yanpeng Zhao, Caiyun Zhao, Bowei Yan, Chong Dai, Lianlian Wu, Yuqi Wen, Dongjin Leng, Zhongming Wang, Xiaoxi Yang, Xiaochen Bo |
Briefings Bioinform. | 8 |
| 2022 | Machine learning methods, databases and tools for drug combination predictionabstractCombination therapy has shown an obvious efficacy on complex diseases and can greatly reduce the development of drug resistance. However, even with high-throughput screens, experimental methods are insufficient to explore novel drug combinations. In order to reduce the search space of drug combinations, there is an urgent need to develop more efficient computational methods to predict novel drug combinations. In recent decades, more and more machine learning (ML) algorithms have been applied to improve the predictive performance. The object of this study is to introduce and discuss the recent applications of ML methods and the widely used databases in drug combination prediction. In this study, we first describe the concept and controversy of synergism between drug combinations. Then, we investigate various publicly available data resources and tools for prediction tasks. Next, ML methods including classic ML and deep learning methods applied in drug combination prediction are introduced. Finally, we summarize the challenges to ML methods in prediction tasks and provide a discussion on future work. Lianlian Wu, Yuqi Wen, Dongjin Leng, Chong Dai, Zhongming Wang, Bowei Yan, Xiaochen Bo |
Briefings Bioinform. | 1 |
| 2022 | Systematic optimization of host-directed therapeutic targets and preclinical validation of repositioned antiviral drugsabstractInhibition of host protein functions using established drugs produces a promising antiviral effect with excellent safety profiles, decreased incidence of resistant variants and favorable balance of costs and risks. Genomic methods have produced a large number of robust host factors, providing candidates for identification of antiviral drug targets. However, there is a lack of global perspectives and systematic prioritization of known virus-targeted host proteins (VTHPs) and drug targets. There is also a need for host-directed repositioned antivirals. Here, we integrated 6140 VTHPs and grouped viral infection modes from a new perspective of enriched pathways of VTHPs. Clarifying the superiority of nonessential membrane and hub VTHPs as potential ideal targets for repositioned antivirals, we proposed 543 candidate VTHPs. We then presented a large-scale drug-virus network (DVN) based on matching these VTHPs and drug targets. We predicted possible indications for 703 approved drugs against 35 viruses and explored their potential as broad-spectrum antivirals. In vitro and in vivo tests validated the efficacy of bosutinib, maraviroc and dextromethorphan against human herpesvirus 1 (HHV-1), hepatitis B virus (HBV) and influenza A virus (IAV). Their drug synergy with clinically used antivirals was evaluated and confirmed. The results proved that low-dose dextromethorphan is better than high-dose in both single and combined treatments. This study provides a comprehensive landscape and optimization strategy for druggable VTHPs, constructing an innovative and potent pipeline to discover novel antiviral host proteins and repositioned drugs, which may facilitate their delivery to clinical application in translational medicine to combat fatal and spreading viral infections. Dafei Xie, Lianlian Wu, Pingkun Zhou, Xunlong Shi, Xiaochen Bo |
Briefings Bioinform. | 4 |
| 2021 | Multi-dimensional data integration algorithm based on random walk with restartabstractBACKGROUND: The accumulation of various multi-omics data and computational approaches for data integration can accelerate the development of precision medicine. However, the algorithm development for multi-omics data integration remains a pressing challenge. RESULTS: Here, we propose a multi-omics data integration algorithm based on random walk with restart (RWR) on multiplex network. We call the resulting methodology Random Walk with Restart for multi-dimensional data Fusion (RWRF). RWRF uses similarity network of samples as the basis for integration. It constructs the similarity network for each data type and then connects corresponding samples of multiple similarity networks to create a multiplex sample network. By applying RWR on the multiplex network, RWRF uses stationary probability distribution to fuse similarity networks. We applied RWRF to The Cancer Genome Atlas (TCGA) data to identify subtypes in different cancer data sets. Three types of data (mRNA expression, DNA methylation, and microRNA expression data) are integrated and network clustering is conducted. Experiment results show that RWRF performs better than single data type analysis and previous integrative methods. CONCLUSIONS: RWRF provides powerful support to users to decipher the cancer molecular subtypes, thus may benefit precision treatment of specific patients in clinical practice. Yuqi Wen, Xinyu Song 0002, Bowei Yan, Xiaoxi Yang, Lianlian Wu, Dongjin Leng, Xiaochen Bo |
BMC Bioinform. | 5 |
| 2021 | Synthetic Lethal Interactions Prediction Based on Multiple Similarity Measures Fusion
Lianlian Wu, Yuqi Wen, Xiaoxi Yang, Bowei Yan, Xiaochen Bo |
J. Comput. Sci. Technol. | 1 |
| 2021 | COMSUC: A web server for the identification of consensus molecular subtypes of cancer based on multiple methods and multi-omics dataabstractExtensive amounts of multi-omics data and multiple cancer subtyping methods have been developed rapidly, and generate discrepant clustering results, which poses challenges for cancer molecular subtype research. Thus, the development of methods for the identification of cancer consensus molecular subtypes is essential. The lack of intuitive and easy-to-use analytical tools has posed a barrier. Here, we report on the development of the COnsensus Molecular SUbtype of Cancer (COMSUC) web server. With COMSUC, users can explore consensus molecular subtypes of more than 30 cancers based on eight clustering methods, five types of omics data from public reference datasets or users' private data, and three consensus clustering methods. The web server provides interactive and modifiable visualization, and publishable output of analysis results. Researchers can also exchange consensus subtype results with collaborators via project IDs. COMSUC is now publicly and freely available with no login requirement at http://comsuc.bioinforai.tech/ (IP address: http://59.110.25.27/). For a video summary of this web server, see S1 Video and S1 File. Xinyu Song 0002, Xiaoxi Yang, Jijun Yu, Yuqi Wen, Lianlian Wu, Bowei Yan, Jiannan Feng, Xiaochen Bo |
PLoS Comput. Biol. | 6 |