EDBT 2026 Demo / reviewers in the wild / expert
Jialu Hu
dblp:57/7618
· DBLP profile ↗
23ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0002-3351-8020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 9 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Inference of gene coexpression networks from single-cell transcriptome data based on variance decomposition analysisabstractGene regulation varies across different cell types and developmental stages, leading to distinct cellular roles across cellular populations. Investigating cell type-specific gene coexpression is therefore crucial for understanding gene functions and disease pathology. However, reconstructing gene coexpression networks from single-cell transcriptome data is challenging due to artifacts, noise, and data sparsity. Here, we present an efficient method for inference of gene coexpression networks via variance decomposition analysis (GCNVDA) to explore the underlying gene regulatory mechanisms from single-cell transcriptome data. Our model incorporates multiple sources of variability, including a random effect term $G$ to capture gene-level variance and a random effect term $E$ to account for residual errors. We applied GCNVDA to three real-world single-cell datasets, demonstrating that our method outperforms existing state-of-the-art algorithms in both sensitivity and specificity for identifying tissue- or state-specific gene regulations. Furthermore, GCNVDA facilitates the discovery of functional modules that play critical roles in key biological processes such as embryonic development. These findings provide new insights into cell-specific regulatory mechanisms and have the potential to significantly advance research in developmental biology and disease pathology. Bin Lian, Haohui Zhang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, N. Ahmad Aziz, Jialu Hu |
Briefings Bioinform. | 7 |
| 2025 | cfMethylPre: deep transfer learning enhances cancer detection based on circulating cell-free DNA methylation profilingabstractCancer remains a significant global health burden, underscoring the need for innovative diagnostic tools to enable early detection and improve patient outcomes. While circulating cell-free DNA (cfDNA) methylation has emerged as a promising biomarker for noninvasive cancer diagnostics, existing methods often face limitations in handling the high-dimensionality of methylation data, small sample sizes, and a lack of biological interpretability. To address these challenges, we propose cfMethylPre, a novel deep transfer learning framework tailored for cancer detection using cfDNA methylation data. cfMethylPre leverages large language model pretrained embeddings from DNA sequence information and integrates them with methylation profiles to enhance feature representation. The deep transfer learning process involves pretraining on bulk DNA methylation data encompassing 2801 samples across 82 cancer types and normal controls, followed by fine-tuning with cfDNA methylation data. This approach ensures robust adaptation to cfDNA's unique characteristics while improving predictive accuracy. Our model achieved superior predictive accuracy compared with state-of-the-art methods, with a weighted Matthews Correlation Coefficient of 0.926 and a weighted F1-score of 0.942. Through model interpretation and biological experimental validation, we identified three novel breast cancer genes-PCDHA10, PRICKLE2, and PRTG-demonstrating their inhibitory effects on cell proliferation and migration in breast cancer cell lines. These findings establish cfMethylPre as a powerful and interpretable tool for cancer diagnostics and biological discovery, paving the way for its application in precision oncology. Xuchao Zhang, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Yanpu Wang, Tao Wang 0082 |
Briefings Bioinform. | 5 |
| 2025 | MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoderabstractSpatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST. Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082 |
Briefings Bioinform. | 6 |
| 2025 | DualMarker: A Multi-Source Fusion Identification Method for Prognostic Biomarkers of Breast Cancer Based on Dual-Layer Heterogeneous NetworkabstractBreast cancer is a complex disease that arises from multiple factors, including genetics, age, and environmental factors. Prognosis prediction for breast cancer is a challenging task that urgently needs to be addressed. Prognostic biomarkers can aid in predicting clinical outcomes for breast cancer patients, and network-based approaches are frequently employed to identify such biomarkers. However, the accuracy of these approaches based on single source biological network is poor due to incomplete interactions of single biological network. Some network-based approaches that integrate multiple biological networks have not considered network denoising, which may lead to the accuracy of these approaches to be improved. We propose a multi-source fusion identification method named DualMarker for prognostic biomarkers of breast cancer. This method constructs a dual-layer heterogeneous network by integrating multiple biological sources. To decrease the negative effects of incomplete interactions in biological networks, we denoise the constructed network. The ranking of features is obtained by the network propagation algorithm and the initial scoring strategy. Compared with six other network-based methods, DualMarker shows the best performance in six breast cancer datasets. Moreover, we have also demonstrated that the biomarkers identified by DualMarker are of interpretability biologically and closely associated with breast cancer patients' prognosis. Xingyi Li 0003, Gaoyuan Du, Zhelin Zhao, Ju Xiang, Jialu Hu, Xuequn Shang 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Accurately deciphering spatial domains for spatially resolved transcriptomics with stClusterabstractSpatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster. Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001 |
Briefings Bioinform. | 3 |
| 2024 | Scbean: a python library for single-cell multi-omics data analysisabstractSUMMARY: Single-cell multi-omics technologies provide a unique platform for characterizing cell states and reconstructing developmental process by simultaneously quantifying and integrating molecular signatures across various modalities, including genome, transcriptome, epigenome, and other omics layers. However, there is still an urgent unmet need for novel computational tools in this nascent field, which are critical for both effective and efficient interrogation of functionality across different omics modalities. Scbean represents a user-friendly Python library, designed to seamlessly incorporate a diverse array of models for the examination of single-cell data, encompassing both paired and unpaired multi-omics data. The library offers uniform and straightforward interfaces for tasks, such as dimensionality reduction, batch effect elimination, cell label transfer from well-annotated scRNA-seq data to scATAC-seq data, and the identification of spatially variable genes. Moreover, Scbean's models are engineered to harness the computational power of GPU acceleration through Tensorflow, rendering them capable of effortlessly handling datasets comprising millions of cells. AVAILABILITY AND IMPLEMENTATION: Scbean is released on the Python Package Index (PyPI) (https://pypi.org/project/scbean/) and GitHub (https://github.com/jhu99/scbean) under the MIT license. The documentation and example code can be found at https://scbean.readthedocs.io/en/latest/. Haohui Zhang, Bin Lian, Xingyi Li 0003, Tao Wang 0082, Xuequn Shang 0001, Ahmad Aziz, Jialu Hu |
Bioinform. | 10 |
| 2023 | Spherical Convolution-based Saliency Detection for FoV Prediction in 360-degree Video StreamingabstractField of view (FoV) prediction is a crucial issue in 360° video streaming, which is the basis for selectively transmitting panoramic videos to reduce bandwidth. The saliency feature is a very important part of FoV prediction. The saliency area identifies a user’s region of interest (RoI) and reflects the user’s viewing behavior preference. The regular convolutional neural network (CNN) cannot effectively extract the spatial representation of panoramic video content because significant geometric distortion will be introduced after panoramic video projection, especially in polar regions. In this paper, we propose a depth neural network model based on spherical convolution, which can learn the spatial features of the 360° videos by encoding the distortion invariance into the architecture of CNNs. A series of experiments on the public 360° video saliency dataset show the proposed model outperforms the existing saliency models. Finally, we embed the proposed saliency network into a popular FoV prediction framework and propose a complete FoV prediction framework for 360° video streaming. Shuai Peng, Jialu Hu, Changqiao Xu |
IWCMC | 2 |
| 2023 | A multi-view latent variable model reveals cellular heterogeneity in complex tissues for paired multimodal single-cell dataabstractMOTIVATION: Single-cell multimodal assays allow us to simultaneously measure two different molecular features of the same cell, enabling new insights into cellular heterogeneity, cell development and diseases. However, most existing methods suffer from inaccurate dimensionality reduction for the joint-modality data, hindering their discovery of novel or rare cell subpopulations. RESULTS: Here, we present VIMCCA, a computational framework based on variational-assisted multi-view canonical correlation analysis to integrate paired multimodal single-cell data. Our statistical model uses a common latent variable to interpret the common source of variances in two different data modalities. Our approach jointly learns an inference model and two modality-specific non-linear models by leveraging variational inference and deep learning. We perform VIMCCA and compare it with 10 existing state-of-the-art algorithms on four paired multi-modal datasets sequenced by different protocols. Results demonstrate that VIMCCA facilitates integrating various types of joint-modality data, thus leading to more reliable and accurate downstream analysis. VIMCCA improves our ability to identify novel or rare cell subtypes compared to existing widely used methods. Besides, it can also facilitate inferring cell lineage based on joint-modality profiles. AVAILABILITY AND IMPLEMENTATION: The VIMCCA algorithm has been implemented in our toolkit package scbean (≥0.5.0), and its code has been archived at https://github.com/jhu99/scbean under MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Lian, Haohui Zhang, Yuanke Zhong, Fashuai Wu, Knut Reinert, Xuequn Shang 0001, Jialu Hu |
Bioinform. | 10 |
| 2022 | A multi-source fusion method to identify biomarkers for breast cancer prognosis based on dual-layer heterogeneous networkabstractThe prognosis of breast cancer is challenging, which is an urgent problem to be solved. The prognostic biomarkers for breast cancer can help us predict the clinical outcomes of patients, and network-based methods are widely introduced to find prognostic biomarkers. According to the difference of input biological data, existing network-based biomarker prediction methods are mainly classified into two types: integrating single-source network or multi-source networks. However, the interactome of single-source network remains incomplete, and biological networks are noisy, which will hamper the network-based identification accuracy of biomarkers. In this study, we introduce a multi-source fusion method, DualMarker, which integrates multiple biological information sources and constructs a dual-layer heterogeneous network by fast network embedding. Next, we introduce a network enhancement method to denoise the constructed dual-layer heterogeneous network, and we implement network propagation algorithm on the constructed dual-layer heterogeneous network to rank the features. After comparing with competitive methods, we find that DualMarker substantially outperforms these methods. In addition, we verify that the biomarkers identified by DualMarker are closely related to the prognosis of breast cancer patients. Xingyi Li 0003, Zhelin Zhao, Ju Xiang, Jialu Hu, Xuequn Shang 0001 |
BIBM | 4 |
| 2022 | A versatile and scalable single-cell data integration algorithm based on domain-adversarial and variational approximationabstractSingle-cell technologies provide us new ways to profile transcriptomic landscape, chromatin accessibility, spatial expression patterns in heterogeneous tissues at the resolution of single cell. With enormous generated single-cell datasets, a key analytic challenge is to integrate these datasets to gain biological insights into cellular compositions. Here, we developed a domain-adversarial and variational approximation, DAVAE, which can integrate multiple single-cell datasets across samples, technologies and modalities with a single strategy. Besides, DAVAE can also integrate paired data of ATAC profile and transcriptome profile that are simultaneously measured from a same cell. With a mini-batch stochastic gradient descent strategy, it is scalable for large-scale data and can be accelerated by GPUs. Results on seven real data integration applications demonstrated the effectiveness and scalability of DAVAE in batch-effect removing, transfer learning and cell-type predictions for multiple single-cell datasets across samples, technologies and modalities. Availability: DAVAE has been implemented in a toolkit package "scbean" in the pypi repository, and the source code can be also freely accessible at https://github.com/jhu99/scbean. All our data and source code for reproducing the results of this paper can be accessible at https://github.com/jhu99/davae_paper. Jialu Hu, Yuanke Zhong, Xuequn Shang 0001 |
Briefings Bioinform. | 1 |
| 2021 | SCC: an accurate imputation method for scRNA-seq dropouts based on a mixture modelabstractBACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables the possibility of many in-depth transcriptomic analyses at a single-cell resolution. It's already widely used for exploring the dynamic development process of life, studying the gene regulation mechanism, and discovering new cell types. However, the low RNA capture rate, which cause highly sparse expression with dropout, makes it difficult to do downstream analyses. RESULTS: We propose a new method SCC to impute the dropouts of scRNA-seq data. Experiment results show that SCC gives competitive results compared to two existing methods while showing superiority in reducing the intra-class distance of cells and improving the clustering accuracy in both simulation and real data. CONCLUSIONS: SCC is an effective tool to resolve the dropout noise in scRNA-seq data. The code is freely accessible at https://github.com/nwpuzhengyan/SCC . Yuanke Zhong, Jialu Hu, Xuequn Shang 0001 |
BMC Bioinform. | 3 |
| 2020 | Twadn: an efficient alignment algorithm based on time warping for pairwise dynamic networksabstractBACKGROUND: Network alignment is an efficient computational framework in the prediction of protein function and phylogenetic relationships in systems biology. However, most of existing alignment methods focus on aligning PPIs based on static network model, which are actually dynamic in real-world systems. The dynamic characteristic of PPI networks is essential for understanding the evolution and regulation mechanism at the molecular level and there is still much room to improve the alignment quality in dynamic networks. RESULTS: In this paper, we proposed a novel alignment algorithm, Twadn, to align dynamic PPI networks based on a strategy of time warping. We compare Twadn with the existing dynamic network alignment algorithm DynaMAGNA++ and DynaWAVE and use area under the receiver operating characteristic curve and area under the precision-recall curve as evaluation indicators. The experimental results show that Twadn is superior to DynaMAGNA++ and DynaWAVE. In addition, we use protein interaction network of Drosophila to compare Twadn and the static network alignment algorithm NetCoffee2 and experimental results show that Twadn is able to capture timing information compared to NetCoffee2. CONCLUSIONS: Twadn is a versatile and efficient alignment tool that can be applied to dynamic network. Hopefully, its application can benefit the research community in the fields of molecular function and evolution. Yuanke Zhong, Junhao He, Yiqun Gao, Xuequn Shang 0001, Jialu Hu |
BMC Bioinform. | 8 |
| 2019 | Deep learning enables accurate alignment of single cell RNA-seq dataabstractAs more and more single-cell RNA-seq (scRNA-seq) datasets become available, carrying out compare between them is key. However, this task is challengeable due to differences caused by different experiment. We proposed a single cell alignment method using deep autoencoder followed by k-nearst-neighbor cells (scadKNN), which learns the feature representation of the data while eliminating batch effects and dropouts through deep autoencoder and uses the low-dimensional feature to align cell types, thereby reducing calculation effort and improving alignment accuracy. Experiments using different real datasets are employed to showcase the effectiveness of the proposed approach. Yuanke Zhong, Xuequn Shang 0001, Jialu Hu |
BIBM | 6 |
| 2019 | A novel algorithm based on bi-random walks to identify disease-related lncRNAsabstractBACKGROUNDS: There is evidence to suggest that lncRNAs are associated with distinct and diverse biological processes. The dysfunction or mutation of lncRNAs are implicated in a wide range of diseases. An accurate computational model can benefit the diagnosis of diseases and help us to gain a better understanding of the molecular mechanism. Although many related algorithms have been proposed, there is still much room to improve the accuracy of the algorithm. RESULTS: We developed a novel algorithm, BiWalkLDA, to predict disease-related lncRNAs in three real datasets, which have 528 lncRNAs, 545 diseases and 1216 interactions in total. To compare performance with other algorithms, the leave-one-out validation test was performed for BiWalkLDA and three other existing algorithms, SIMCLDA, LDAP and LRLSLDA. Additional tests were carefully designed to analyze the parameter effects such as α, β, l and r, which could help user to select the best choice of these parameters in their own application. In a case study of prostate cancer, eight out of the top-ten disease-related lncRNAs reported by BiWalkLDA were previously confirmed in literatures. CONCLUSIONS: In this paper, we develop an algorithm, BiWalkLDA, to predict lncRNA-disease association by using bi-random walks. It constructs a lncRNA-disease network by integrating interaction profile and gene ontology information. Solving cold-start problem by using neighbors' interaction profile information. Then, bi-random walks was applied to three real biological datasets. Results show that our method outperforms other algorithms in predicting lncRNA-disease association in terms of both accuracy and specificity. AVAILABILITY: https://github.com/screamer/BiwalkLDA. Jialu Hu, Yiqun Gao, Xuequn Shang 0001 |
BMC Bioinform. | 1 |
| 2019 | MD-SVM: a novel SVM-based algorithm for the motif discovery of transcription factor binding sitesabstractBACKGROUND: Transcription factors (TFs) play important roles in the regulation of gene expression. They can activate or block transcription of downstream genes in a manner of binding to specific genomic sequences. Therefore, motif discovery of these binding preference patterns is of central significance in the understanding of molecular regulation mechanism. Many algorithms have been proposed for the identification of transcription factor binding sites. However, it remains a challengeable problem. RESULTS: Here, we proposed a novel motif discovery algorithm based on support vector machine (MD-SVM) to learn a discriminative model for TF binding sites. MD-SVM firstly obtains position weight matrix (PWM) from a set of training datasets. Then it translates the MD problem into a computational framework of multiple instance learning (MIL). It was applied to several real biological datasets. Results show that our algorithm outperforms MI-SVM in terms of both accuracy and specificity. CONCLUSIONS: In this paper, we modeled the TF motif discovery problem as a MIL optimization problem. The SVM algorithm was adapted to discriminate positive and negative bags of instances. Compared to other svm-based algorithms, MD-SVM show its superiority over its competitors in term of ROC AUC. Hopefully, it could be of benefit to the research community in the understanding of molecular functions of DNA functional elements and transcription factors. Jialu Hu, Tianwei Liu, Yuanke Zhong, Yiqun Gao, Junhao He, Xuequn Shang 0001 |
BMC Bioinform. | 1 |
| 2018 | Identification of lncRNA-disease association using bi-random walks
Yiqun Gao, Jialu Hu, Xuequn Shang 0001 |
BIBM | 2 |
| 2018 | NetCoffee2: A Novel Global Alignment Algorithm for Multiple PPI Networks Based on Graph Feature Vectors
Jialu Hu, Junhao He, Yiqun Gao, Xuequn Shang 0001 |
ICIC (2) | 1 |
| 2018 | WebNetCoffee: a web-based application to identify functionally conserved proteins from Multiple PPI networksabstractBACKGROUND: The discovery of functionally conserved proteins is a tough and important task in system biology. Global network alignment provides a systematic framework to search for these proteins from multiple protein-protein interaction (PPI) networks. Although there exist many web servers for network alignment, no one allows to perform global multiple network alignment tasks on users' test datasets. RESULTS: Here, we developed a web server WebNetcoffee based on the algorithm of NetCoffee to search for a global network alignment from multiple networks. To build a series of online test datasets, we manually collected 218,339 proteins, 4,009,541 interactions and many other associated protein annotations from several public databases. All these datasets and alignment results are available for download, which can support users to perform algorithm comparison and downstream analyses. CONCLUSION: WebNetCoffee provides a versatile, interactive and user-friendly interface for easily running alignment tasks on both online datasets and users' test datasets, managing submitted jobs and visualizing the alignment results through a web browser. Additionally, our web server also facilitates graphical visualization of induced subnetworks for a given protein and its neighborhood. To the best of our knowledge, it is the first web server that facilitates the performing of global alignment for multiple PPI networks. AVAILABILITY: http://www.nwpu-bioinformatics.com/WebNetCoffee. Jialu Hu, Yiqun Gao, Junhao He, Xuequn Shang 0001 |
BMC Bioinform. | 1 |
| 2017 | MiteFinder: A fast approach to identify miniature inverted-repeat transposable elements on a genome-wide scaleabstractMiniature inverted-repeat transposable element (M ITE) is a type of class II non-autonomous transposable element playing a crucial role in the process of evolution in biology. Development of bioinformatics tools that are capable of effectively identifying MITEs can enable genome-wide studies of MITE patterns in eukaryotes. Here, we present a fast, accurate and memory-efficient tool, MiteFinder, for the identification of MITEs from genomics sequences. MiteFinder distinguishes itself from other existing methods by building k-mer indexes for large genomes, which can effectively identify all possible inverted repeats. To calculate the likelihood of how much a given sequence is a MITE, we introduce a log-ratio scoring model that uses the distribution of sequence pattern in two models, mite model (M) and null model (N). The results suggest that MiteFinder outperforms existing tools in both precision and recall. Besides, it is much faster and more memory-efficient than other tools in the detection. The source code is freely accessible at the website: https://github.com/screamer/miteFinder. Jialu Hu, Xuequn Shang 0001 |
BIBM | 1 |
| 2016 | Analyzing factors involved in the HPO-based semantic similarity calculationabstractAlthough disease diagnosis have greatly benefited from next generation sequencing technologies, it is still difficult to make the right diagnosis based on purely sequencing technologies for many diseases with complex phenotypes and high genetic heterogeneity. Recently, calculating Human Phenotype Ontology (HPO)-based phenotype semantic similarity has contributed a lot for completing disease diagnosis. However, factors which affect the accuracy of HPO-based semantic similarity have not been evaluated systematically. In this study, we propose a new framework called HPOFactor to evaluate these factors. Jiajie Peng, Jialu Hu, Xuequn Shang 0001 |
BIBM | 4 |
| 2015 | LocalAli: an evolutionary-based local alignment approach to identify functionally conserved modules in multiple networksabstractMOTIVATION: Sequences and protein interaction data are of significance to understand the underlying molecular mechanism of organisms. Local network alignment is one of key systematic ways for predicting protein functions, identifying functional modules and understanding the phylogeny from these data. Most of currently existing tools, however, encounter their limitations, which are mainly concerned with scoring scheme, speed and scalability. Therefore, there are growing demands for sophisticated network evolution models and efficient local alignment algorithms. RESULTS: We developed a fast and scalable local network alignment tool called LocalAli for the identification of functionally conserved modules in multiple networks. In this algorithm, we firstly proposed a new framework to reconstruct the evolution history of conserved modules based on a maximum-parsimony evolutionary model. By relying on this model, LocalAli facilitates interpretation of resulting local alignments in terms of conserved modules, which have been evolved from a common ancestral module through a series of evolutionary events. A meta-heuristic method simulated annealing was used to search for the optimal or near-optimal inner nodes (i.e. ancestral modules) of the evolutionary tree. To evaluate the performance and the statistical significance, LocalAli were tested on 26 real datasets and 1040 randomly generated datasets. The results suggest that LocalAli outperforms all existing algorithms in terms of coverage, consistency and scalability, meanwhile retains a high precision in the identification of functionally coherent subnetworks. AVAILABILITY: The source code and test datasets are freely available for download under the GNU GPL v3 license at https://code.google.com/p/localali/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jialu Hu, Knut Reinert |
Bioinform. | 1 |
| 2014 | NetCoffee: a fast and accurate global alignment approach to identify functionally conserved proteins in multiple networksabstractMOTIVATION: Owing to recent advancements in high-throughput technologies, protein-protein interaction networks of more and more species become available in public databases. The question of how to identify functionally conserved proteins across species attracts a lot of attention in computational biology. Network alignments provide a systematic way to solve this problem. However, most existing alignment tools encounter limitations in tackling this problem. Therefore, the demand for faster and more efficient alignment tools is growing. RESULTS: We present a fast and accurate algorithm, NetCoffee, which allows to find a global alignment of multiple protein-protein interaction networks. NetCoffee searches for a global alignment by maximizing a target function using simulated annealing on a set of weighted bipartite graphs that are constructed using a triplet approach similar to T-Coffee. To assess its performance, NetCoffee was applied to four real datasets. Our results suggest that NetCoffee remedies several limitations of previous algorithms, outperforms all existing alignment tools in terms of speed and nevertheless identifies biologically meaningful alignments. AVAILABILITY: The source code and data are freely available for download under the GNU GPL v3 license at https://code.google.com/p/netcoffee/. Jialu Hu, Birte Kehr, Knut Reinert |
Bioinform. | 1 |
| 2009 | Evaluation of subgraph searching algorithms detecting network motif in biological networks
Jialu Hu, Lin Gao 0006, Guimin Qin |
Frontiers Comput. Sci. China | 1 |