VLDB 2026 Research / reviewers in the wild / expert
Huimin Luo
dblp:18/10303
· DBLP profile ↗
32ranked-venue papers
9as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 5 first-author · 15 since 2021Computer networks · 7 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tlaloc: A Generic Multipath Load Balancing for RoCE
Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
INFOCOM | 1 |
| 2026 | Euler: An Out-of-Order-Aware Load Balancing with Adaptive Granularity for AI Clusters
Jiafeng Jiang, Jiao Zhang 0002, Huimin Luo, Shuo Wang 0006, Tao Huang 0005 |
WCNC | 3 |
| 2026 | Weir: Scalable RDMA With Delay-Based RNIC Cache Control Software Middleware for Data Center NetworksabstractRemote Direct Memory Access (RDMA) is widely used in distributed services in Data Center Networks (DCNs) due to its high performance. As DCNs expand in scale, RDMA faces scalability issues. The reason is that the high concurrency Queue Pairs (QPs) lead to cache misses on RDMA Network Interface Card (RNIC) and frequent evictions, and the behaviour of fetching the cache via PCIe leads to performance degradation of RDMA. In this paper, we model the behaviour of Work Queue Element (WQE) on RNIC as a producer-consumer model and investigate that the root cause of WQE cache misses is the mismatch between the production rate of the CPU and the consumption rate of the RNIC. We design Weir from the perspective of WQE cache control to avoid cache misses and improve throughput under high concurrent QPs. Weir determines the cache occupancy on the RNIC by monitoring the number of active QPs and the increase/decrease in the life cycle of WQEs, and calculates the production rate and pacing by credit. The implementation of Weir exhibits minimal CPU overhead. Evaluation results show that Weir can maintain 97Gbps throughput without degradation even with up to 16K concurrent QPs, and effectively reduces various observable cache misses by$5\times $to$10\times $compared to commercial RNICs. Additionally, experiments show that Weir has better connection scalability than XRC and DCT. Jiao Zhang 0002, Yongchen Pan, Dexuan Liao, Huimin Luo, Tao Huang 0005, Haipeng Yao |
IEEE Trans. Netw. | 5 |
| 2025 | Multi-View Fusion with Knowledge Graph for Anticancer Drug Synergy PredictionabstractSynergistic drug combinations inhibit cancer cells through multiple mechanisms of action, significantly improving cancer treatment efficacy and alleviating drug resistance. This makes them an effective approach for cancer therapy. In recent years, the drug combination prediction methods based on machine learning and deep learning have been developed to pre-screen potential synergistic drug combinations. Among these, some knowledge graph (KG)-based methods predict drug combinations by leveraging the rich entity relationships in the KG. However, existing KG-based methods either fail to fully exploit the interaction between drugs and cell lines or overlook the implicit knowledge embedded in semantically related but topologically disconnected entities in the KG. Additionally, knowledge graphs often introduce noise into the representation learning process due to the imbalance in reality data distribution. To address these, we propose MFKGSynergy, a drug synergy prediction model based on knowledge graph and multi-view information fusion. The model constructs three views: (i) Drugcell Line interaction view for global interaction information between drugs and cell lines, using a Graph Transformer attention mechanism; (ii) Knowledge Graph View: a heterogeneous graph relation-aware attention network is utilized to learn high-order topological information between drugs/cell lines and entities along connected paths in KG; (iii) Latent Semantic View, which captures implicit knowledge in the KG by calculating the similarity between drugs. Then, contrastive learning is employed to align drug and cell line representations across different views. And a gated aggregation module is designed to weight and fuse interaction and KG information, effectively reducing noise caused by data distribution imbalance. It is shown that our model improves prediction performance compared to the state-of-theart model. Huimin Luo |
BIBM | 5 |
| 2025 | scPSNGCL: Single-Cell Graph Contrastive Clustering using a Pseudo-Siamese NetworkabstractSingle-cell clustering is a vital step in analyzing single-cell RNA sequencing (scRNA-seq) data, essential for revealing the complexity of the tissue and discovering cellular heterogeneity. Recently, deep learning-based methods have been successfully applied to single-cell sequencing analysis. However, existing models either fail to simultaneously account for both attribute information and structural information in single-cell data, or lack a mechanism to directly optimize the consistency between these two heterogeneous perspectives during training, thereby limiting the full exploitation of their complementary potential. In this study, we propose a single-cell graph contrastive clustering model based on a pseudo-siamese network, scPSNGCL. This framework utilizes the pseudo-siamese network to obtain fused cell representations and employs dual-level contrastive learning at both the feature and cluster levels to collaboratively guide the model in optimizing representations within the feature and semantic spaces. Finally, a self-supervised clustering module is applied to jointly optimize all loss functions, thereby comprehensively enhancing the clustering performance. Finally, the experimental results on 7 real-world datasets demonstrated that scPSNGCL outperforms other advanced methods. Maohua Qin, Huimin Luo |
BIBM | 4 |
| 2025 | Weir: Delay-based RNIC Cache Control Software Middleware for Scalable RDMA Networks
Yongchen Pan, Jiao Zhang 0002, Zirui Wan, Baohong Lin, Junliang Wang, Huimin Luo |
INFOCOM | 7 |
| 2025 | DeepHapNet: a haplotype assembly method based on RetNet and deep spectral clusteringabstractGene polymorphism originates from single-nucleotide polymorphisms (SNPs), and the analysis and study of SNPs are of great significance in the field of biogenetics. The haplotype, which consists of the sequence of SNP loci, carries more genetic information than a single SNP. Haplotype assembly plays a significant role in understanding gene function, diagnosing complex diseases, and pinpointing species genes. We propose a novel method, DeepHapNet, for haplotype assembly through the clustering of reads and learning correlations between read pairs. We employ a sequence model called Retentive Network (RetNet), which utilizes a multiscale retention mechanism to extract read features and learn the global relationships among them. Based on the feature representation of reads learned from the RetNet model, the clustering process of reads is implemented using the SpectralNet model, and, finally, haplotypes are constructed based on the read clusters. Experiments with simulated and real datasets show that the method performs well in the haplotype assembly problem of diploid and polyploid based on either long or short reads. The code implementation of DeepHapNet and the processing scripts for experimental data are publicly available at https://github.com/wjj6666/DeepHapNet. Jing-Jing Wei, Chaokun Yan, Huimin Luo |
Briefings Bioinform. | 5 |
| 2025 | deepTAD: an approach for identifying topologically associated domains based on convolutional neural network and transformer modelabstractMOTIVATION: Topologically associated domains (TADs) play a key role in the 3D organization and function of genomes, and accurate detection of TADs is essential for revealing the relationship between genomic structure and function. Most current methods are developed to extract features in Hi-C interaction matrix to identify TADs. However, due to complexities in Hi-C contact matrices, it is difficult to directly extract features associated with TADs, which prevents current methods from identifying accurate TADs. RESULTS: In this paper, a novel method is proposed, deepTAD, which is developed based on a convolutional neural network (CNN) and transformer model. First, based on Hi-C contact matrix, deepTAD utilizes CNN to directly extract features associated with TAD boundaries. Next, deepTAD takes advantage of the transformer model to analyze the variation features around TAD boundaries and determines the TAD boundaries. Second, deepTAD uses the Wilcoxon rank-sum test to further identify false-positive boundaries. Finally, deepTAD computes cosine similarity among identified TAD boundaries and assembles TAD boundaries to obtain hierarchical TADs. The experimental results show that TAD boundaries identified by deepTAD have a significant enrichment of biological features, including structural proteins, histone modifications, and transcription start site loci. Additionally, when evaluating the completeness and accuracy of identified TADs, deepTAD has a good performance compared with other methods. The source code of deepTAD is available at https://github.com/xiaoyan-wang99/deepTAD. Huimin Luo, Fei Guo 0001 |
Briefings Bioinform. | 4 |
| 2025 | DconnLoop: a deep learning model for predicting chromatin loops based on multi-source data integrationabstractBACKGROUND: Chromatin loops are critical for the three-dimensional organization of the genome and gene regulation. Accurate identification of chromatin loops is essential for understanding the regulatory mechanisms in disease. However, current mainstream detection methods rely primarily on single-source data, such as Hi-C, which limits these methods' ability to capture the diverse features of chromatin loop structures. In contrast, multi-source data integration and deep learning approaches, though not yet widely applied, hold significant potential. RESULTS: In this study, we developed a method called DconnLoop to integrate Hi-C, ChIP-seq, and ATAC-seq data to predict chromatin loops. This method achieves feature extraction and fusion of multi-source data by integrating residual mechanisms, directional connectivity excitation modules, and interactive feature space decoders. Finally, we apply density estimation and density clustering to the genome-wide prediction results to identify more representative loops. The code is available from https://github.com/kuikui-C/DconnLoop . CONCLUSIONS: The results demonstrate that DconnLoop outperforms existing methods in both precision and recall. In various experiments, including Aggregate Peak Analysis and peak enrichment comparisons, DconnLoop consistently shows advantages. Extensive ablation studies and validation across different sequencing depths further confirm DconnLoop's robustness and generalizability. Kuikui Cheng, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 4 |
| 2025 | SeqBalance: Congestion-Aware Load Balancing With No Reordering in Data Center NetworksabstractWith the rapid development of the Internet of Things (IoT), an increasing amount of sensor data generated by IoT applications has been transferred to data center networks for storage and data analysis. Remote Direct Memory Access (RDMA) is widely used in data center networks because of its high performance. However, due to the characteristics of RDMA’s retransmission strategy, current load balancing schemes for data center networks are unsuitable for RDMA. In this paper, we propose SeqBalance, a load balancing framework designed for RDMA. SeqBalance implements fine-grained load balancing for RDMA through a reasonable design and does not cause reordering problems. SeqBalance detects link congestion at the switch by sensing ECN signals and link utilization, and guides routing decisions accordingly. SeqBalance’s designs are all based on existing commercial RNICs and commercial programmable switches, so they are compatible with existing data center networks. We have implemented SeqBalance Shaper for fine-grained sub-flow splitting in Mellanox CX-6 RNIC and implemented routing decisions in Intel Tofino P4 programmable switch. The results of hardware testbed experiments and large-scale simulations show that compared with existing load balancing schemes, SeqBalance improves 24.7% and 15.9% on average FCT and 99th-percentile FCT. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Internet Things J. | 1 |
| 2025 | RoCELet: Host-Based Flowlet Load Balancing for RoCEabstractRemote Direct Memory Access (RDMA) is becoming a popular high-speed networking technology. It uses kernel bypass and zero copy to achieve high throughput and low latency with little CPU overhead. However, standard RoCE transmission uses Equal Cost Multipath (ECMP) for load balancing, which can result in lower transmission performance due to hash conflicts. Meanwhile, it has been verified that, unlike TCP, the unique retransmission mode and flow characteristics of RoCE make previous load balancing algorithms not well applied to RoCE. In this paper, we introduce RoCELet, a load balancing algorithm for RoCE. It achieves fine-grained RoCE load balancing by actively generating flowlets, effectively utilizing the rich end-to-end paths in the data center. We implement a prototype based on DPDK and evaluate it through small-scale testbed experiments and large-scale simulations. Our results show that compared to state-of-the-art load balancing algorithms, RoCELet optimizes 48.2% and 16.4% in average FCT and$99^{th}$-ile FCT, respectively. Huimin Luo, Jiao Zhang 0002, Mingxuan Yu, Jiafeng Jiang, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
IEEE Trans. Netw. | 1 |
| 2024 | ResDeepGS: A Deep Learning-Based Method for Crop Phenotype Prediction
Chaokun Yan, Huimin Luo |
ISBRA (2) | 5 |
| 2024 | Enhancing Privacy and Preserving Accuracy in Medical Image Classification with Limited Labeled Samples
Chaokun Yan, Menghan Yin, Wenjuan Liang, Haicao Yan, Huimin Luo |
ISBRA (1) | 5 |
| 2024 | Blaze: Delay-Aware Cloud-Edge Collaborative Service Function Chain Deployment with Network CalculusabstractWith the rapid development of Internet of the Things (IoT) technology, IoT services have higher and higher requirements for latency. In the IoT environment, virtual network functions (VNFs) are deployed on general-purpose hardware and are sequentially connected to form service function chain (SFC) to provide network services for IoT devices. However, the high latency of the link between the cloud center and the edge nodes and the resource capacity limitation of the edge nodes pose challenges to the deployment of SFCs in IoT devices. In this paper, we study the cloud-edge collaborative SFC deployment problem. We applied the network calculus theory to the cloud-edge collaborative SFC deployment for the first time, aiming to provide the end-to-end delay guarantee for the deployed SFC. We model the SFC deployment problem as Mixed Integer Nonlinear Programming (MINLP). Then we propose a heuristic algorithm (Blaze) to solve this problem. Blaze is proven to complete the deployment of SFCs in polynomial time. Finally, the algorithm is evaluated by experimental simulation. The experimental results show that compared with the existing state-of-the-art corresponding algorithms, the proposed algorithm achieves better performance in terms of the number of VNFs deployed in the cloud, resource consumption of edge nodes, and SFC request acceptance rate. Huimin Luo, Jiao Zhang 0002, Yongchen Pan, Tian Pan 0001, Tao Huang 0005 |
WCNC | 1 |
| 2023 | SLHSD: hybrid scaffolding method based on short and long readsabstractIn genome assembly, scaffolding can obtain more complete and continuous scaffolds. Current scaffolding methods usually adopt one type of read to construct a scaffold graph and then orient and order contigs. However, scaffolding with the strengths of two or more types of reads seems to be a better solution to some tricky problems. Combining the advantages of different types of data is significant for scaffolding. Here, a hybrid scaffolding method (SLHSD) is present that simultaneously leverages the precision of short reads and the length advantage of long reads. Building an optimal scaffold graph is an important foundation for getting scaffolds. SLHSD uses a new algorithm that combines long and short read alignment information to determine whether to add an edge and how to calculate the edge weight in a scaffold graph. In addition, SLHSD develops a strategy to ensure that edges with high confidence can be added to the graph with priority. Then, a linear programming model is used to detect and remove remaining false edges in the graph. We compared SLHSD with other scaffolding methods on five datasets. Experimental results show that SLHSD outperforms other methods. The open-source code of SLHSD is available at https://github.com/luojunwei/SLHSD. Ting Guan, Guolin Chen, Zhonghua Yu, Haixia Zhai, Chaokun Yan, Huimin Luo |
Briefings Bioinform. | 7 |
| 2023 | KGANSynergy: knowledge graph attention network for drug synergy predictionabstractCombination therapy is widely used to treat complex diseases, particularly in patients who respond poorly to monotherapy. For example, compared with the use of a single drug, drug combinations can reduce drug resistance and improve the efficacy of cancer treatment. Thus, it is vital for researchers and society to help develop effective combination therapies through clinical trials. However, high-throughput synergistic drug combination screening remains challenging and expensive in the large combinational space, where an array of compounds are used. To solve this problem, various computational approaches have been proposed to effectively identify drug combinations by utilizing drug-related biomedical information. In this study, considering the implications of various types of neighbor information of drug entities, we propose a novel end-to-end Knowledge Graph Attention Network to predict drug synergy (KGANSynergy), which utilizes neighbor information of known drugs/cell lines effectively. KGANSynergy uses knowledge graph (KG) hierarchical propagation to find multi-source neighbor nodes for drugs and cell lines. The knowledge graph attention network is designed to distinguish the importance of neighbors in a KG through a multi-attention mechanism and then aggregate the entity's neighbor node information to enrich the entity. Finally, the learned drug and cell line embeddings can be utilized to predict the synergy of drug combinations. Experiments demonstrated that our method outperformed several other competing methods, indicating that our method is effective in identifying drug combinations. Zhijie Gao, Chaokun Yan, Wenjuan Liang, Huimin Luo |
Briefings Bioinform. | 7 |
| 2023 | Detection of the Status of Diatom Blooms in the Tributaries of the Yangtze River Based on Sentinel-2 ImagesabstractDiatom blooms frequently observed in the river tributaries pose a threat to the regional water environment. Satellite remote sensing provides us with an effective tool to delineate the extent of large-scale harmful algal blooms (HABs) in both oceans and inland waters. However, the diatom bloom detection in river systems remains challenging, due to the high demand for a robust model to characterize the spatial and spectral features from the satellite images with high resolution and limited spectral bands. In this study, we developed a novel deep learning-based framework for the detection of the status of diatom blooms in river tributaries using Sentinel-2 MultiSpectral Imager (MSI) images. Distinct from the previous works for detection of bloom extent, the water pixels are categorized as ‘non-bloom’, ‘mild bloom’, and ‘severe bloom’, corresponding to the different bloom intensities. To achieve this, a simple convolutional neural network (CNN) is trained using the collected spectral samples from MSI images characterizing the different bloom statuses. The input features include the spectral variables highly correlated to the pivotal absorption and backscattering properties of diatoms, the chlorophyll-a (Chla), and water temperature. The classification results can then be obtained based on the spectral and environmental characteristics learned by the constructed model. The trained model was extensively tested in tributaries of Yangtze River, where both visual and quantitative validation demonstrated that the proposed model can achieve reliable detection results. Linwei Yue, Huimin Luo, Huanfeng Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Deep learning approach for cancer subtype classification using high-dimensional gene expression dataabstractMOTIVATION: Studies have shown that classifying cancer subtypes can provide valuable information for a range of cancer research, from aetiology and tumour biology to prognosis and personalized treatment. Current methods usually adopt gene expression data to perform cancer subtype classification. However, cancer samples are scarce, and the high-dimensional features of their gene expression data are too sparse to allow most methods to achieve desirable classification results. RESULTS: In this paper, we propose a deep learning approach by combining a convolutional neural network (CNN) and bidirectional gated recurrent unit (BiGRU): our approach, DCGN, aims to achieve nonlinear dimensionality reduction and learn features to eliminate irrelevant factors in gene expression data. Specifically, DCGN first uses the synthetic minority oversampling technique algorithm to equalize data. The CNN can handle high-dimensional data without stress and extract important local features, and the BiGRU can analyse deep features and retain their important information; the DCGN captures key features by combining both neural networks to overcome the challenges of small sample sizes and sparse, high-dimensional features. In the experiments, we compared the DCGN to seven other cancer subtype classification methods using breast and bladder cancer gene expression datasets. The experimental results show that the DCGN performs better than the other seven methods and can provide more satisfactory classification results. Jiquan Shen, Haixia Zhai, Zhengjiang Wu, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 8 |
| 2021 | Biomedical data and computational models for drug repositioning: a comprehensive reviewabstractDrug repositioning can drastically decrease the cost and duration taken by traditional drug research and development while avoiding the occurrence of unforeseen adverse events. With the rapid advancement of high-throughput technologies and the explosion of various biological data and medical data, computational drug repositioning methods have been appealing and powerful techniques to systematically identify potential drug-target interactions and drug-disease interactions. In this review, we first summarize the available biomedical data and public databases related to drugs, diseases and targets. Then, we discuss existing drug repositioning approaches and group them based on their underlying computational models consisting of classical machine learning, network propagation, matrix factorization and completion, and deep learning based models. We also comprehensively analyze common standard data sets and evaluation metrics used in drug repositioning, and give a brief comparison of various prediction methods on the gold standard data sets. Finally, we conclude our review with a brief discussion on challenges in computational drug repositioning, which includes the problem of reducing the noise and incompleteness of biomedical data, the ensemble of various computation drug repositioning methods, the importance of designing reliable negative samples selection methods, new techniques dealing with the data sparseness problem, the construction of large-scale and comprehensive benchmark data sets and the analysis and explanation of the underlying mechanisms of predicted interactions. Huimin Luo, Min Li 0007, Mengyun Yang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | A comprehensive review of scaffolding methods in genome assemblyabstractIn the field of genome assembly, scaffolding methods make it possible to obtain a more complete and contiguous reference genome, which is the cornerstone of genomic research. Scaffolding methods typically utilize the alignments between contigs and sequencing data (reads) to determine the orientation and order among contigs and to produce longer scaffolds, which are helpful for genomic downstream analysis. With the rapid development of high-throughput sequencing technologies, diverse types of reads have emerged over the past decade, especially in long-range sequencing, which have greatly enhanced the assembly quality of scaffolding methods. As the number of scaffolding methods increases, biology and bioinformatics researchers need to perform in-depth analyses of state-of-the-art scaffolding methods. In this article, we focus on the difficulties in scaffolding, the differences in characteristics among various kinds of reads, the methods by which current scaffolding methods address these difficulties, and future research opportunities. We hope this work will benefit the design of new scaffolding methods and the selection of appropriate scaffolding methods for specific biological studies. Yawei Wei, Mengna Lyu, Zhengjiang Wu, Huimin Luo, Chaokun Yan |
Briefings Bioinform. | 6 |
| 2021 | BreakNet: detecting deletions using long reads and a deep learning approachabstractBACKGROUND: Structural variations (SVs) occupy a prominent position in human genetic diversity, and deletions form an important type of SV that has been suggested to be associated with genetic diseases. Although various deletion calling methods based on long reads have been proposed, a new approach is still needed to mine features in long-read alignment information. Recently, deep learning has attracted much attention in genome analysis, and it is a promising technique for calling SVs. RESULTS: In this paper, we propose BreakNet, a deep learning method that detects deletions by using long reads. BreakNet first extracts feature matrices from long-read alignments. Second, it uses a time-distributed convolutional neural network (CNN) to integrate and map the feature matrices to feature vectors. Third, BreakNet employs a bidirectional long short-term memory (BLSTM) model to analyse the produced set of continuous feature vectors in both the forward and backward directions. Finally, a classification module determines whether a region refers to a deletion. On real long-read sequencing datasets, we demonstrate that BreakNet outperforms Sniffles, SVIM and cuteSV in terms of their F1 scores. The source code for the proposed method is available from GitHub at https://github.com/luojunwei/BreakNet . CONCLUSIONS: Our work shows that deep learning can be combined with long reads to call deletions more effectively than existing methods. Hongyu Ding, Jiquan Shen, Haixia Zhai, Zhengjiang Wu, Chaokun Yan, Huimin Luo |
BMC Bioinform. | 7 |
| 2021 | A Novel Drug Repositioning Approach Based on Collaborative Metric LearningabstractComputational drug repositioning, which is an efficient approach to find potential indications for drugs, has been used to increase the efficiency of drug development. The drug repositioning problem essentially is a top-K recommendation task that recommends most likely diseases to drugs based on drug and disease related information. Therefore, many recommendation methods can be adopted to drug repositioning. Collaborative metric learning (CML) algorithm can produce distance metrics that capture the important relationships among objects, and has been widely used in recommendation domains. By applying CML in drug repositioning, a joint metric space is learned to encode drug's relationships with different diseases. In this study, we propose a novel drug repositioning computational method using Collaborative Metric Learning to predict novel drug-disease associations based on known drug and disease related information. Specifically, the proposed method learns latent vectors of drugs and diseases by applying metric learning, and then predicts the association probability of one drug-disease pair based on the learned vectors. The comprehensive experimental results show that CMLDR outperforms the other state-of-the-art drug repositioning algorithms in terms of precision, recall, and AUPR. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Correction to: SLR: a scaffolding algorithm based on long reads and contig classificationabstractFollowing publication of the original article [1], the author reported that there is an error in the original article. Mengna Lyu, Ranran Chen, Huimin Luo, Chaokun Yan |
BMC Bioinform. | 5 |
| 2020 | NEDD: a network embedding based method for predicting drug-disease associationsabstractBACKGROUND: Drug discovery is known for the large amount of money and time it consumes and the high risk it takes. Drug repositioning has, therefore, become a popular approach to save time and cost by finding novel indications for approved drugs. In order to distinguish these novel indications accurately in a great many of latent associations between drugs and diseases, it is necessary to exploit abundant heterogeneous information about drugs and diseases. RESULTS: In this article, we propose a meta-path-based computational method called NEDD to predict novel associations between drugs and diseases using heterogeneous information. First, we construct a heterogeneous network as an undirected graph by integrating drug-drug similarity, disease-disease similarity, and known drug-disease associations. NEDD uses meta paths of different lengths to explicitly capture the indirect relationships, or high order proximity, within drugs and diseases, by which the low dimensional representation vectors of drugs and diseases are obtained. NEDD then uses a random forest classifier to predict novel associations between drugs and diseases. CONCLUSIONS: The experiments on a gold standard dataset which contains 1933 validated drug-disease associations show that NEDD produces superior prediction results compared with the state-of-the-art approaches. Renyi Zhou, Zhangli Lu, Huimin Luo, Ju Xiang, Min Zeng 0004, Min Li 0007 |
BMC Bioinform. | 3 |
| 2020 | GapReduce: A Gap Filling Algorithm Based on Partitioned Read SetsabstractWith the advances in technologies of sequencing and assembly, draft sequences of more and more genomes are available. However, there commonly exist gaps in these draft sequences which influence various downstream analysis of biological studies. Gap filling methods can shorten the length of gaps and improve the completion of these draft sequences of genomes. Although some gap filling tools have been developed, their effectiveness and accuracy need to be improved. In this study, we develop a novel tool, called GapReduce, which can fill the gaps using the paired reads. For a gap, GapReduce selects the reads whose mate reads are aligned on the left or the right flanking region, and partitions the reads to two sets. Then GapReduce adopts different $k$k values and $k$k-$mer$mer frequency thresholds to iteratively construct De Bruijn graphs, which are used for finding the correct path to fill the gap. For overcoming the branching problems caused by repetitive regions and sequencing errors in the procedure of path selection, GapReduce designs a novel approach that simultaneously considers $k$k-$mer$mer frequency and distribution of paired reads based on the partitioned read sets. We compare the performance of GapReduce with current popular gap filling tools. The experimental results demonstrate that GapReduce can produce satisfactory gap filling results, especially for long insert size datasets. GapReduce is publicly available for downloading at https://github.com/bioinfomaticsCSU/GapReduce. Jianxin Wang 0001, Juan Shang, Huimin Luo, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | Drug and disease similarity calculation platform for drug repositioningabstractDrug repositioning, aiming to infer potential indications for drugs efficiently, has achieved remarkable results in reducing the cycle, cost and risk of drug Research and Development (R&D), and mining new uses of known drugs. Currently, many computational drug repositioning strategies have been proposed. The similarity calculation, as one of the key steps of drug repositioning, has an important impact on the accuracy of computational drug repositioning. However, the biological data used for the similarity calculation come from a wide range of sources with different formats, and similarity calculation methods are developed in different programming languages, thus the similarity calculation methods are varying. To facilitate the similarity calculation for drug repositioning, we developed a computational platform consisting of various datasets and similarity measures for drugs and diseases, and four programming languages (Java, R, Python and MATLAB) are supported by our platform. Users can use relevant data and methods directly according to their needs and customize similarity calculation methods. The platform is available at: http://bioinformatics.csu.edu.cn/artemis/. Huimin Luo, Mengyun Yang, Fang-Xiang Wu, Jianxin Wang 0001 |
BIBM | 2 |
| 2019 | Drug repositioning based on bounded nuclear norm regularizationabstractMOTIVATION: Computational drug repositioning is a cost-effective strategy to identify novel indications for existing drugs. Drug repositioning is often modeled as a recommendation system problem. Taking advantage of the known drug-disease associations, the objective of the recommendation system is to identify new treatments by filling out the unknown entries in the drug-disease association matrix, which is known as matrix completion. Underpinned by the fact that common molecular pathways contribute to many different diseases, the recommendation system assumes that the underlying latent factors determining drug-disease associations are highly correlated. In other words, the drug-disease matrix to be completed is low-rank. Accordingly, matrix completion algorithms efficiently constructing low-rank drug-disease matrix approximations consistent with known associations can be of immense help in discovering the novel drug-disease associations. RESULTS: In this article, we propose to use a bounded nuclear norm regularization (BNNR) method to complete the drug-disease matrix under the low-rank assumption. Instead of strictly fitting the known elements, BNNR is designed to tolerate the noisy drug-drug and disease-disease similarities by incorporating a regularization term to balance the approximation error and the rank properties. Moreover, additional constraints are incorporated into BNNR to ensure that all predicted matrix entry values are within the specific interval. BNNR is carried out on an adjacency matrix of a heterogeneous drug-disease network, which integrates the drug-drug, drug-disease and disease-disease networks. It not only makes full use of available drugs, diseases and their association information, but also is capable of dealing with cold start naturally. Our computational results show that BNNR yields higher drug-disease association prediction accuracy than the current state-of-the-art methods. The most significant gain is in prediction precision measured as the fraction of the positive predictions that are truly positive, which is particularly useful in drug design practice. Cases studies also confirm the accuracy and reliability of BNNR. AVAILABILITY AND IMPLEMENTATION: The code of BNNR is freely available at https://github.com/BioinformaticsCSU/BNNR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mengyun Yang, Huimin Luo, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2019 | SLR: a scaffolding algorithm based on long reads and contig classificationabstractBACKGROUND: Scaffolding is an important step in genome assembly that orders and orients the contigs produced by assemblers. However, repetitive regions in contigs usually prevent scaffolding from producing accurate results. How to solve the problem of repetitive regions has received a great deal of attention. In the past few years, long reads sequenced by third-generation sequencing technologies (Pacific Biosciences and Oxford Nanopore) have been demonstrated to be useful for sequencing repetitive regions in genomes. Although some stand-alone scaffolding algorithms based on long reads have been presented, scaffolding still requires a new strategy to take full advantage of the characteristics of long reads. RESULTS: Here, we present a new scaffolding algorithm based on long reads and contig classification (SLR). Through the alignment information of long reads and contigs, SLR classifies the contigs into unique contigs and ambiguous contigs for addressing the problem of repetitive regions. Next, SLR uses only unique contigs to produce draft scaffolds. Then, SLR inserts the ambiguous contigs into the draft scaffolds and produces the final scaffolds. We compare SLR to three popular scaffolding tools by using long read datasets sequenced with Pacific Biosciences and Oxford Nanopore technologies. The experimental results show that SLR can produce better results in terms of accuracy and completeness. The open-source code of SLR is available at https://github.com/luojunwei/SLR. CONCLUSION: In this paper, we describes SLR, which is designed to scaffold contigs using long reads. We conclude that SLR can improve the completeness of genome assembly. Mengna Lyu, Ranran Chen, Huimin Luo, Chaokun Yan |
BMC Bioinform. | 5 |
| 2019 | Overlap matrix completion for predicting drug-associated indicationsabstractIdentification of potential drug-associated indications is critical for either approved or novel drugs in drug repositioning. Current computational methods based on drug similarity and disease similarity have been developed to predict drug-disease associations. When more reliable drug- or disease-related information becomes available and is integrated, the prediction precision can be continuously improved. However, it is a challenging problem to effectively incorporate multiple types of prior information, representing different characteristics of drugs and diseases, to identify promising drug-disease associations. In this study, we propose an overlap matrix completion (OMC) for bilayer networks (OMC2) and tri-layer networks (OMC3) to predict potential drug-associated indications, respectively. OMC is able to efficiently exploit the underlying low-rank structures of the drug-disease association matrices. In OMC2, first of all, we construct one bilayer network from drug-side aspect and one from disease-side aspect, and then obtain their corresponding block adjacency matrices. We then propose the OMC2 algorithm to fill out the values of the missing entries in these two adjacency matrices, and predict the scores of unknown drug-disease pairs. Moreover, we further extend OMC2 to OMC3 to handle tri-layer networks. Computational experiments on various datasets indicate that our OMC methods can effectively predict the potential drug-disease associations. Compared with the other state-of-the-art approaches, our methods yield higher prediction accuracy in 10-fold cross-validation and de novo experiments. In addition, case studies also confirm the effectiveness of our methods in identifying promising indications for existing drugs in practical applications. Mengyun Yang, Huimin Luo, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
PLoS Comput. Biol. | 2 |
| 2019 | Computational Drug Repositioning with Random Walk on a Heterogeneous NetworkabstractDrug repositioning is an efficient and promising strategy to identify new indications for existing drugs, which can improve the productivity of traditional drug discovery and development. Rapid advances in high-throughput technologies have generated various types of biomedical data over the past decades, which lay the foundations for furthering the development of computational drug repositioning approaches. Although many researches have tried to improve the repositioning accuracy by integrating information from multiple sources and different levels, it is still appealing to further investigate how to efficiently exploit valuable data for drug repositioning. In this study, we propose an efficient approach, Random Walk on a Heterogeneous Network for Drug Repositioning (RWHNDR), to prioritize candidate drugs for diseases. First, an integrated heterogeneous network is constructed by combining multiple sources including drugs, drug targets, diseases and disease genes data. Then, a random walk model is developed to capture the global information of the heterogeneous network. RWHNDR takes advantage of drug targets and disease genes data more comprehensively for drug repositioning. The experiment results show that our approach can achieve better performance, compared with other state-of-the-art approaches which prioritized candidate drugs based on multi-source data. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Kaijie Zhao, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | Computational drug repositioning using low-rank matrix approximation and randomized algorithmsabstractMotivation: Computational drug repositioning is an important and efficient approach towards identifying novel treatments for diseases in drug discovery. The emergence of large-scale, heterogeneous biological and biomedical datasets has provided an unprecedented opportunity for developing computational drug repositioning methods. The drug repositioning problem can be modeled as a recommendation system that recommends novel treatments based on known drug-disease associations. The formulation under this recommendation system is matrix completion, assuming that the hidden factors contributing to drug-disease associations are highly correlated and thus the corresponding data matrix is low-rank. Under this assumption, the matrix completion algorithm fills out the unknown entries in the drug-disease matrix by constructing a low-rank matrix approximation, where new drug-disease associations having not been validated can be screened. Results: In this work, we propose a drug repositioning recommendation system (DRRS) to predict novel drug indications by integrating related data sources and validated information of drugs and diseases. Firstly, we construct a heterogeneous drug-disease interaction network by integrating drug-drug, disease-disease and drug-disease networks. The heterogeneous network is represented by a large drug-disease adjacency matrix, whose entries include drug pairs, disease pairs, known drug-disease interaction pairs and unknown drug-disease pairs. Then, we adopt a fast Singular Value Thresholding (SVT) algorithm to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. The comprehensive experimental results show that DRRS improves the prediction accuracy compared with the other state-of-the-art approaches. In addition, case studies for several selected drugs further demonstrate the practical usefulness of the proposed method. Availability and implementation: http://bioinformatics.csu.edu.cn/resources/softs/DrugRepositioning/DRRS/index.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Huimin Luo, Min Li 0007, Shaokai Wang, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 1 |
| 2016 | Drug repositioning based on comprehensive similarity measures and Bi-Random walk algorithmabstractMOTIVATION: Drug repositioning, which aims to identify new indications for existing drugs, offers a promising alternative to reduce the total time and cost of traditional drug development. Many computational strategies for drug repositioning have been proposed, which are based on similarities among drugs and diseases. Current studies typically use either only drug-related properties (e.g. chemical structures) or only disease-related properties (e.g. phenotypes) to calculate drug or disease similarity, respectively, while not taking into account the influence of known drug-disease association information on the similarity measures. RESULTS: In this article, based on the assumption that similar drugs are normally associated with similar diseases and vice versa, we propose a novel computational method named MBiRW, which utilizes some comprehensive similarity measures and Bi-Random walk (BiRW) algorithm to identify potential novel indications for a given drug. By integrating drug or disease features information with known drug-disease associations, the comprehensive similarity measures are firstly developed to calculate similarity for drugs and diseases. Then drug similarity network and disease similarity network are constructed, and they are incorporated into a heterogeneous network with known drug-disease interactions. Based on the drug-disease heterogeneous network, BiRW algorithm is adopted to predict novel potential drug-disease associations. Computational experiment results from various datasets demonstrate that the proposed approach has reliable prediction performance and outperforms several recent computational drug repositioning approaches. Moreover, case studies of five selected drugs further confirm the superior performance of our method to discover potential indications for drugs practically. AVAILABILITY AND IMPLEMENTATION: http://github.com//bioinfomaticsCSU/MBiRW CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Huimin Luo, Jianxin Wang 0001, Min Li 0007, Xiaoqing Peng, Fang-Xiang Wu, Yi Pan 0001 |
Bioinform. | 1 |