VLDB 2026 Research / reviewers in the wild / expert
Kang Ning 0001
dblp:n/KangNing
· DBLP profile ↗
28ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-3325-5387ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Techniques for learning and transferring knowledge for microbiome-based classification and prediction: review and assessmentabstractThe volume of microbiome data is growing at an exponential rate, and the current methodologies for big data mining are encountering substantial obstacles. Effectively managing and extracting valuable insights from these vast microbiome datasets has emerged as a significant challenge in the field of contemporary microbiome research. This comprehensive review delves into the utilization of foundation models and transfer learning techniques within the context of microbiome-based classification and prediction tasks, advocating for a transition away from traditional task-specific or scenario-specific models towards more adaptable, continuous learning models. The article underscores the practicality and benefits of initially constructing a robust foundation model, which can then be fine-tuned using transfer learning to tackle specific context tasks. In real-world scenarios, the application of transfer learning empowers models to leverage disease-related data from one geographical area and enhance diagnostic precision in different regions. This transition from relying on "good models" to embracing "adaptive models" resonates with the philosophy of "teaching a man to fish" thereby paving the way for advancements in personalized medicine and accurate diagnosis. Empirical research suggests that the integration of foundation models with transfer learning methodologies substantially boosts the performance of models when dealing with large-scale and diverse microbiome datasets, effectively mitigating the challenges posed by data heterogeneity. Haohong Zhang, Kang Ning 0001 |
Briefings Bioinform. | 3 |
| 2025 | Discovery and optimization of antimicrobial peptides from extreme environments on global scaleabstractAbstract Background Antimicrobial peptides (AMPs) inhibit microbial growth through membrane disruption or interference with intracellular processes, offering promising solutions to antimicrobial resistance (AMR). While AMPs have been extensively identified and verified from animal proteomes, reference microbial genomes and host environments, those from extreme habitats remain largely unexplored. Microbes in such extreme niches with low-level of nutrients evolve unique membrane modification and specialized metabolic pathway, representing a hidden reservoir for novel AMP discovery. Methods We developed Atlantis, a language model-driven framework for AMP detection and optimization from global extremophile (Fig. 1). It combined two modules: Atlantis-Prospect, which predict antimicrobial potency based on both sequence and structure information, and Atlantis-Search, which conducts iterative single-mutations with structural constraints to broaden peptides’ antimicrobial activity from narrow to broad spectrum. The pretrained structure-aware language model was fine-tuned on small proteins from extremophile to capture biome-specific evolutionary information. Results By mining 60,461 extremophile metagenomes from nine discrete habitats across multiple geographical scales, we identified 1662 non-redundant peptides with potential antimicrobial activity from 85 bacterial and archaeal phyla, none of which match existing databases. We synthesized fourteen peptides, and two of them showed stronger antimicrobial activity then commercial AMP, LL-37, in vitro against Escherichia coli, Pseudomonas aeruginosa and Staphylococcus aureus (Fig. 2). Zixin Kang, Haohong Zhang, Kang Ning 0001 |
Briefings Bioinform. | 3 |
| 2025 | Quantifying antibiotic resistome risks across environmental niches: the L-ARRAP for long-read metagenomic profilingabstractThe global dissemination of antibiotic resistance genes (ARGs) represents a critical challenge to One Health. Existing ARG risk assessment tools (e.g. MetaCompare, ARRI) are constrained by short-read sequencing data, limiting their utility for long-read platforms. To address this gap, we developed the Long-read based Antibiotic Resistome Risk Assessment Pipeline (L-ARRAP), which calculates the Long-read based Antibiotic Resistome Risk Index (L-ARRI) to quantify antibiotic resistome risks. Building upon our previous ARRI framework, L-ARRAP leverages long-read sequencing advantages to concurrently identify ARGs, mobile genetic elements, and human bacterial pathogens, integrating their interactions for risk scoring. Our results showed that L-ARRAP was not only able to accurately identify ARGs and evaluate the antibiotic resistance risk scores in samples of hospital wastewater (HWW), Chaohu lake, and human fecal samples, but also significantly distinguish the ARG risk in HWW samples between before and after disinfection groups, demonstrating the performance of L-ARRAP. Furthermore, L-ARRAP scores exhibited strong concordance with those generated by our laboratory-adapted MetaCompare variant (L-MetaCompare), corroborating its methodological reliability. Overall, to our knowledge, L-ARRAP is the first assessment pipeline of antibiotic resistome for long sequencing reads and has a great potential for monitoring the risk of ARGs in various environmental niches. Yujie Mao, Mingchao Wang, Yunyi Qin, Caili Zhang, Qingru Chen, Kang Ning 0001, Maozhen Han |
Briefings Bioinform. | 9 |
| 2025 | PREDAC-FluB: predicting antigenic clusters of seasonal influenza B viruses with protein language model embedding based convolutional neural networkabstractInfluenza poses a significant global public health threat, with vaccination being the most effective and economical preventive measure. However, these punctuated antigenic changes, particularly in HA, result in escape from the immunity that was induced by prior infection or vaccination. Accurately predicting antigenic variation and understanding the antigenic dynamics of influenza viruses are crucial for selecting appropriate vaccine strains, but no established methods exist for influenza B viruses. Therefore, we present PREDAC-FluB, a hybrid deep learning framework that integrates spatial feature extraction via CNN to model interactions in HA1 sequences, multimodal sequence representation combining ESM-2 embeddings with six physicochemical descriptors and continuous encoding (ESM2-7-features), and UMAP-guided clustering for antigenic cluster identification. Using data from 9036 B/Victoria-lineage and 4520 B/Yamagata-lineage influenza virus pair. PREDAC-FluB demonstrates superior performance over traditional machine learning methods in predicting antigenic variation in influenza viruses, successfully identifying major antigenic clusters. Specifically, PREDAC-FluB classified the B/Victoria lineage into nine antigenic clusters and the B/Yamagata lineage into three antigenic clusters. In five-fold cross-validation for B/Victoria viruses, PREDAC-FluB with ESM2-7-features encoding achieved AUROC values of 0.9961 on the validation set and 0.9856 on the independent test set. In retrospective testing for B/Victoria viruses, PREDAC-FluB achieved AUROC values ranging from 0.83 to 0.97, demonstrating high prediction accuracy and effectively capturing antigenic variation information. In conclusion, PREDAC-FluB is a robust tool for antigenic computation, capable of accurately predicting antigenic variation in influenza B viruses. Its high prediction accuracy makes it a promising auxiliary method for recommending future influenza vaccine strains. Wenping Xie, Jingze Liu, Jiangyuan Wang, Wenjie Han, Yousong Peng, Xiangjun Du, Kang Ning 0001, Taijiao Jiang |
Briefings Bioinform. | 9 |
| 2023 | Artificial intelligence-enabled microbiome-based diagnosis models for a broad spectrum of cancer typesabstractMicrobiome-based diagnosis of cancer is an increasingly important supplement for the genomics approach in cancer diagnosis, yet current models for microbiome-based diagnosis of cancer face difficulties in generality: not only diagnosis models could not be adapted from one cancer to another, but models built based on microbes from tissues could not be adapted for diagnosis based on microbes from blood. Therefore, a microbiome-based model suitable for a broad spectrum of cancer types is urgently needed. Here we have introduced DeepMicroCancer, a diagnosis model using artificial intelligence techniques for a broad spectrum of cancer types. Built based on the random forest models it has enabled superior performances on more than twenty types of cancers' tissue samples. And by using the transfer learning techniques, improved accuracies could be obtained, especially for cancer types with only a few samples, which could satisfy the requirement in clinical scenarios. Moreover, transfer learning techniques have enabled high diagnosis accuracy that could also be achieved for blood samples. These results indicated that certain sets of microbes could, if excavated using advanced artificial techniques, reveal the intricate differences among cancers and healthy individuals. Collectively, DeepMicroCancer has provided a new venue for accurate diagnosis of cancer based on tissue and blood materials, which could potentially be used in clinics. Haohong Zhang, Yuguo Zha, Lei Ji 0006, Yuwen Chu, Kang Ning 0001 |
Briefings Bioinform. | 8 |
| 2023 | Tracing human life trajectory using gut microbial communities by context-aware deep learningabstractThe gut microbial communities are highly plastic throughout life, and the human gut microbial communities show spatial-temporal dynamic patterns at different life stages. However, the underlying association between gut microbial communities and time-related factors remains unclear. The lack of context-awareness, insufficient data, and the existence of batch effect are the three major issues, making the life trajection of the host based on gut microbial communities problematic. Here, we used a novel computational approach (microDELTA, microbial-based deep life trajectory) to track longitudinal human gut microbial communities' alterations, which employs transfer learning for context-aware mining of gut microbial community dynamics at different life stages. Using an infant cohort, we demonstrated that microDELTA outperformed Neural Network for accurately predicting the age of infant with different delivery mode, especially for newborn infants of vaginal delivery with the area under the receiver operating characteristic curve of microDELTA and Neural Network at 0.811 and 0.436, respectively. In this context, we have discovered the influence of delivery mode on infant gut microbial communities. Along the human lifespan, we also applied microDELTA to a Chinese traveler cohort, a Hadza hunter-gatherer cohort and an elderly cohort. Results revealed the association between long-term dietary shifts during travel and adult gut microbial communities, the seasonal cycling of gut microbial communities for the Hadza hunter-gatherers, and the distinctive microbial pattern of elderly gut microbial communities. In summary, microDELTA can largely solve the issues in tracing the life trajectory of the human microbial communities and generate accurate and flexible models for a broad spectrum of microbial-based longitudinal researches. Haohong Zhang, Hui Chong, Qingyang Yu, Yuguo Zha, Mingyue Cheng 0003, Kang Ning 0001 |
Briefings Bioinform. | 6 |
| 2023 | ContactLib-ATT: A Structure-Based Search Engine for Homologous ProteinsabstractGeneral-purpose protein structure embedding can be used for many important protein biology tasks, such as protein design, drug design and binding affinity prediction. Recent researches have shown that attention-based encoder layers are more suitable to learn high-level features. Based on this key observation, we propose a two-level general-purpose protein structure embedding neural network, called ContactLib-ATT. On local embedding level, a biologically more meaningful contact context is introduced. On global embedding level, attention-based encoder layers are employed for better global representation learning. Our general-purpose protein structure embedding framework is trained and tested on the SCOP40 2.07 dataset. As a result, ContactLib-ATT achieves a SCOP superfamily classification accuracy of 82.4% (i.e., 6.7% higher than state-of-the-art method). On the same dataset, ContactLib-ATT is used to simulate a structure-based search engine for remote homologous proteins, and our top-10 candidate list contains at least one remote homolog with a probability of 91.9%. Cheng Chen 0031, Yuguo Zha, Daming Zhu, Kang Ning 0001, Xuefeng Cui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | EXPERT: transfer learning-enabled context-aware microbial community classificationabstractMicrobial community classification enables identification of putative type and source of the microbial community, thus facilitating a better understanding of how the taxonomic and functional structure were developed and maintained. However, previous classification models required a trade-off between speed and accuracy, and faced difficulties to be customized for a variety of contexts, especially less studied contexts. Here, we introduced EXPERT based on transfer learning that enabled the classification model to be adaptable in multiple contexts, with both high efficiency and accuracy. More importantly, we demonstrated that transfer learning can facilitate microbial community classification in diverse contexts, such as classification of microbial communities for multiple diseases with limited number of samples, as well as prediction of the changes in gut microbiome across successive stages of colorectal cancer. Broadly, EXPERT enables accurate and context-aware customized microbial community classification, and potentiates novel microbial knowledge discovery. Hui Chong, Yuguo Zha, Qingyang Yu, Mingyue Cheng 0003, Guangzhou Xiong, Xinhe Huang, Shijuan Huang, Chuqing Sun, Sicheng Wu, Wei-Hua Chen, Luís Pedro Coelho, Kang Ning 0001 |
Briefings Bioinform. | 13 |
| 2022 | Ontology-aware neural network: a general framework for pattern mining from microbiome dataabstractWith the rapid accumulation of microbiome data around the world, numerous computational bioinformatics methods have been developed for pattern mining from such paramount microbiome data. Current microbiome data mining methods, such as gene and species mining, rely heavily on sequence comparison. Most of these methods, however, have a clear trade-off, particularly, when it comes to big-data analytical efficiency and accuracy. Microbiome entities are usually organized in ontology structures, and pattern mining methods that have considered ontology structures could offer advantages in mining efficiency and accuracy. Here, we have summarized the ontology-aware neural network (ONN) as a novel framework for microbiome data mining. We have discussed the applications of ONN in multiple contexts, including gene mining, species mining and microbial community dynamic pattern mining. We have then highlighted one of the most important characteristics of ONN, namely, novel knowledge discovery, which makes ONN a standout among all microbiome data mining methods. Finally, we have provided several applications to showcase the advantage of ONN over other methods in microbiome data mining. In summary, ONN represents a paradigm shift for pattern mining from microbiome data: from traditional machine learning approach to ontology-aware and model-based approach, which has found its broad application scenarios in microbiome data mining. Yuguo Zha, Kang Ning 0001 |
Briefings Bioinform. | 2 |
| 2021 | Hydrogen bonds meet self-attention: all you need for protein structure embeddingabstractGeneral-purpose protein structure embedding can be used for many important protein biology tasks, such as protein design, drug design and binding affinity prediction. Recent researches have shown that attention-based encoder layers are more suitable to learn high-level features. Based on this key observation, we treat low-level representation learning and high-level representation learning separately, and propose a two-level general-purpose protein structure embedding neural network, called ContactLib-ATT. On the local embedding level, a simple yet meaningful hydrogen-bond representation is learned. On the global embedding level, attention-based encoder layers are employed for global representation learning. In our experiments, ContactLib-ATT achieves a SCOP superfamily classification accuracy of 82.4% (i.e., 6.7% higher than state-of-the-art method) on the SCOP 40 2.07 dataset. Moreover, ContactLib-ATT is demonstrated to successfully simulate a structure-based search engine for remote homologous proteins, and our top-10 candidate list contains at least one remote homolog with a probability of 91.9%. Cheng Chen 0031, Yuguo Zha, Daming Zhu, Kang Ning 0001, Xuefeng Cui |
BIBM | 4 |
| 2021 | Meta-Prism: Ultra-fast and highly accurate microbial community structure search utilizing dual indexing and parallel computationabstractMicrobiome samples are accumulating at an unprecedented speed. As a result, a massive amount of samples have become available for the mining of the intrinsic patterns among them. However, due to the lack of advanced computational tools, fast yet accurate comparisons and searches among thousands to millions of samples are still in urgent need. In this work, we proposed the Meta-Prism method for comparing and searching the microbial community structures amongst tens of thousands of samples. Meta-Prism is at least 10 times faster than contemporary methods serving the same purpose and can provide very accurate search results. The method is based on three computational techniques: dual-indexing approach for sample subgrouping, refined scoring function that could scrutinize the minute differences among samples, and parallel computation on CPU or GPU. The superiority of Meta-Prism on speed and accuracy for multiple sample searches is proven based on searching against ten thousand samples derived from both human and environments. Therefore, Meta-Prism could facilitate similarity search and in-depth understanding among massive number of heterogenous samples in the microbiome universe. The codes of Meta-Prism are available at: https://github.com/HUST-NingKang-Lab/metaPrism. Mo Zhu, Kang Ning 0001 |
Briefings Bioinform. | 3 |
| 2019 | Strain-GeMS: optimized subspecies identification from microbiome data based on accurate variant modelingabstractMOTIVATION: Subspecies identification is one of the most critical issues in microbiome studies, as it is directly related to their functions in response to the environmental stress and their feedbacks. However, identification of subspecies remains a challenge largely due to the small variation between different strains within the same species. Accurate identification of subspecies primarily relies on variant identification and categorization through microbiome data. However, current SNP calling and subspecies identification for microbiome data remain underdeveloped. RESULTS: In this work, we have proposed Strain-GeMS for subspecies identification from microbiome data, based on solid statistical model for SNP calling, as well as optimized procedure for subspecies identification. Results on simulated, ab initio and in vivo datasets have shown that Strain-GeMS could always generate more accurate results compared with other subspecies identification methods. AVAILABILITY AND IMPLEMENTATION: Strain-GeMS is available at: https://github.com/HUST-NingKang-Lab/straingems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chongyang Tan, Xinping Cui, Kang Ning 0001 |
Bioinform. | 4 |
| 2016 | MultiGeMS: detection of SNVs from multiple samples using model selection on high-throughput sequencing dataabstractMOTIVATION: Single nucleotide variant (SNV) detection procedures are being utilized as never before to analyze the recent abundance of high-throughput DNA sequencing data, both on single and multiple sample datasets. Building on previously published work with the single sample SNV caller genotype model selection (GeMS), a multiple sample version of GeMS (MultiGeMS) is introduced. Unlike other popular multiple sample SNV callers, the MultiGeMS statistical model accounts for enzymatic substitution sequencing errors. It also addresses the multiple testing problem endemic to multiple sample SNV calling and utilizes high performance computing (HPC) techniques. RESULTS: A simulation study demonstrates that MultiGeMS ranks highest in precision among a selection of popular multiple sample SNV callers, while showing exceptional recall in calling common SNVs. Further, both simulation studies and real data analyses indicate that MultiGeMS is robust to low-quality data. We also demonstrate that accounting for enzymatic substitution sequencing errors not only improves SNV call precision at low mapping quality regions, but also improves recall at reference allele-dominated sites with high mapping quality. AVAILABILITY AND IMPLEMENTATION: The MultiGeMS package can be downloaded from https://github.com/cui-lab/multigems CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gabriel H. Murillo, Na You, Xiaoquan Su, Muredach P. Reilly, Mingyao Li, Kang Ning 0001, Xinping Cui |
Bioinform. | 7 |
| 2016 | FALCON@home: a high-throughput protein structure prediction server based on remote homologue recognitionabstractSUMMARY: The protein structure prediction approaches can be categorized into template-based modeling (including homology modeling and threading) and free modeling. However, the existing threading tools perform poorly on remote homologous proteins. Thus, improving fold recognition for remote homologous proteins remains a challenge. Besides, the proteome-wide structure prediction poses another challenge of increasing prediction throughput. In this study, we presented FALCON@home as a protein structure prediction server focusing on remote homologue identification. The design of FALCON@home is based on the observation that a structural template, especially for remote homologous proteins, consists of conserved regions interweaved with highly variable regions. The highly variable regions lead to vague alignments in threading approaches. Thus, FALCON@home first extracts conserved regions from each template and then aligns a query protein with conserved regions only rather than the full-length template directly. This helps avoid the vague alignments rooted in highly variable regions, improving remote homologue identification. We implemented FALCON@home using the Berkeley Open Infrastructure of Network Computing (BOINC) volunteer computing protocol. With computation power donated from over 20,000 volunteer CPUs, FALCON@home shows a throughput as high as processing of over 1000 proteins per day. In the Critical Assessment of protein Structure Prediction (CASP11), the FALCON@home-based prediction was ranked the 12th in the template-based modeling category. As an application, the structures of 880 mouse mitochondria proteins were predicted, which revealed the significant correlation between protein half-lives and protein structural factors. AVAILABILITY AND IMPLEMENTATION: FALCON@home is freely available at http://protein.ict.ac.cn/FALCON/. CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haicang Zhang, Wei-Mou Zheng, Dong Xu 0002, Jianwei Zhu, Kang Ning 0001, Shiwei Sun, Shuaicheng Li 0001, Dongbo Bu |
Bioinform. | 7 |
| 2015 | Condensing Raman spectrum for single-cell phenotype analysisabstractBACKGROUND: In recent years, high throughput and non-invasive Raman spectrometry technique has matured as an effective approach to identification of individual cells by species, even in complex, mixed populations. Raman profiling is an appealing optical microscopic method to achieve this. To fully utilize Raman proling for single-cell analysis, an extensive understanding of Raman spectra is necessary to answer questions such as which filtering methodologies are effective for pre-processing of Raman spectra, what strains can be distinguished by Raman spectra, and what features serve best as Raman-based biomarkers for single-cells, etc. RESULTS: In this work, we have proposed an approach called rDisc to discretize the original Raman spectrum into only a few (usually less than 20) representative peaks (Raman shifts). The approach has advantages in removing noises, and condensing the original spectrum. In particular, effective signal processing procedures were designed to eliminate noise, utilising wavelet transform denoising, baseline correction, and signal normalization. In the discretizing process, representative peaks were selected to signicantly decrease the Raman data size. More importantly, the selected peaks are chosen as suitable to serve as key biological markers to differentiate species and other cellular features. Additionally, the classication performance of discretized spectra was found to be comparable to full spectrum having more than 1000 Raman shifts. Overall, the discretized spectrum needs about 5storage space of a full spectrum and the processing speed is considerably faster. This makes rDisc clearly superior to other methods for single-cell classication. Shiwei Sun, Xuetao Wang, Xin Gao 0001, Lihui Ren, Xiaoquan Su, Dongbo Bu, Kang Ning 0001 |
BMC Bioinform. | 7 |
| 2014 | GPU-Meta-Storms: computing the structure similarities among massive amount of microbial community samples using GPUabstractMOTIVATION: The number of microbial community samples is increasing with exponential speed. Data-mining among microbial community samples could facilitate the discovery of valuable biological information that is still hidden in the massive data. However, current methods for the comparison among microbial communities are limited by their ability to process large amount of samples each with complex community structure. SUMMARY: We have developed an optimized GPU-based software, GPU-Meta-Storms, to efficiently measure the quantitative phylogenetic similarity among massive amount of microbial community samples. Our results have shown that GPU-Meta-Storms would be able to compute the pair-wise similarity scores for 10 240 samples within 20 min, which gained a speed-up of >17 000 times compared with single-core CPU, and >2600 times compared with 16-core CPU. Therefore, the high-performance of GPU-Meta-Storms could facilitate in-depth data mining among massive microbial community samples, and make the real-time analysis and monitoring of temporal or conditional changes for microbial communities possible. AVAILABILITY AND IMPLEMENTATION: GPU-Meta-Storms is implemented by CUDA (Compute Unified Device Architecture) and C++. Source code is available at http://www.computationalbioenergy.org/meta-storms.html. Xiaoquan Su, Xuetao Wang, Gongchao Jing, Kang Ning 0001 |
Bioinform. | 4 |
| 2012 | Meta-Storms: efficient search for similar microbial communities based on a novel indexing scheme and similarity score for metagenomic dataabstractBACKGROUND: It has long been intriguing scientists to effectively compare different microbial communities (also referred as 'metagenomic samples' here) in a large scale: given a set of unknown samples, find similar metagenomic samples from a large repository and examine how similar these samples are. With the current metagenomic samples accumulated, it is possible to build a database of metagenomic samples of interests. Any metagenomic samples could then be searched against this database to find the most similar metagenomic sample(s). However, on one hand, current databases with a large number of metagenomic samples mostly serve as data repositories that offer few functionalities for analysis; and on the other hand, methods to measure the similarity of metagenomic data work well only for small set of samples by pairwise comparison. It is not yet clear, how to efficiently search for metagenomic samples against a large metagenomic database. RESULTS: In this study, we have proposed a novel method, Meta-Storms, that could systematically and efficiently organize and search metagenomic data. It includes the following components: (i) creating a database of metagenomic samples based on their taxonomical annotations, (ii) efficient indexing of samples in the database based on a hierarchical taxonomy indexing strategy, (iii) searching for a metagenomic sample against the database by a fast scoring function based on quantitative phylogeny and (iv) managing database by index export, index import, data insertion, data deletion and database merging. We have collected more than 1300 metagenomic data from the public domain and in-house facilities, and tested the Meta-Storms method on these datasets. Our experimental results show that Meta-Storms is capable of database creation and effective searching for a large number of metagenomic samples, and it could achieve similar accuracies compared with the current popular significance testing-based methods. CONCLUSION: Meta-Storms method would serve as a suitable database management and search system to quickly identify similar metagenomic samples from a large pool of samples. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoquan Su, Jian Xu 0020, Kang Ning 0001 |
Bioinform. | 3 |
| 2012 | SNP calling using genotype model selection on high-throughput sequencing dataabstractMOTIVATION: A review of the available single nucleotide polymorphism (SNP) calling procedures for Illumina high-throughput sequencing (HTS) platform data reveals that most rely mainly on base-calling and mapping qualities as sources of error when calling SNPs. Thus, errors not involved in base-calling or alignment, such as those in genomic sample preparation, are not accounted for. RESULTS: A novel method of consensus and SNP calling, Genotype Model Selection (GeMS), is given which accounts for the errors that occur during the preparation of the genomic sample. Simulations and real data analyses indicate that GeMS has the best performance balance of sensitivity and positive predictive value among the tested SNP callers. AVAILABILITY: The GeMS package can be downloaded from https://sites.google.com/a/bioinformatics.ucr.edu/xinping-cui/home/software or http://computationalbioenergy.org/software.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Na You, Gabriel H. Murillo, Xiaoquan Su, Xiaowei Zeng, Jian Xu 0020, Kang Ning 0001, Shoudong Zhang, Jiankang Zhu, Xinping Cui |
Bioinform. | 6 |
| 2011 | An Open-source Collaboration Environment for Metagenomics ResearchabstractBy analyzing metagenomic data from microbial communities, the taxonomical and functional component of hundreds of previously unknown microbial communities have been elucidated in the past few years. However, metagenomic data analyses are both data- and computation-intensive, which require extensive computational power. Most of the current metagenomic data analysis software were designed to be used on a single PC (Personal Computer), which could not match with the fast increasing number of large metagenomic projects' computational requirements. Therefore, advanced computational environment has to be developed to cope with such needs. In this paper, we proposed an open-source collaboration environment for metagenomic data analysis, which enabled the parallel analysis of multiple metagenomic datasets at the same time. By using this collaboration environment, researchers from different locations could submit their data, collaboratively configure the analysis pipeline, and perform data analysis efficiently. As of now, more than 30 metagenomic data analysis projects have already been conducted based on this environment. Xiaoquan Su, Yongzheng Ma, Xingzhi Chang, Kai Nan, Jian Xu 0020, Kang Ning 0001 |
eScience | 7 |
| 2010 | The utility of mass spectrometry-based proteomic data for validation of novel alternative splice forms reconstructed from RNA-Seq data: a preliminary assessmentabstractBACKGROUND: Most mass spectrometry (MS) based proteomic studies depend on searching acquired tandem mass (MS/MS) spectra against databases of known protein sequences. In these experiments, however, a large number of high quality spectra remain unassigned. These spectra may correspond to novel peptides not present in the database, especially those corresponding to novel alternative splice (AS) forms. Recently, fast and comprehensive profiling of mammalian genomes using deep sequencing (i.e. RNA-Seq) has become possible. MS-based proteomics can potentially be used as an aid for protein-level validation of novel AS events observed in RNA-Seq data. RESULTS: In this work, we have used publicly available mouse tissue proteomic and RNA-Seq datasets and have examined the feasibility of using MS data for the identification of novel AS forms by searching MS/MS spectra against translated mRNA sequences derived from RNA-Seq data. A significant correlation between the likelihood of identifying a peptide from MS/MS data and the number of reads in RNA-Seq data for the same gene was observed. Based on in silico experiments, it was also observed that only a fraction of novel AS forms identified from RNA-Seq had the corresponding junction peptide compatible with MS/MS sequencing. The number of novel peptides that were actually identified from MS/MS spectra was substantially lower than the number expected based on in silico analysis. CONCLUSIONS: The ability to confirm novel AS forms from MS/MS data in the dataset analyzed was found to be quite limited. This can be explained in part by low abundance of many novel transcripts, with the abundance of their corresponding protein products falling below the limit of detection by MS. Kang Ning 0001, Alexey I. Nesvizhskii |
BMC Bioinform. | 1 |
| 2010 | Examination of the relationship between essential genes in PPI network and hub proteins in reverse nearest neighbor topologyabstractBACKGROUND: In many protein-protein interaction (PPI) networks, densely connected hub proteins are more likely to be essential proteins. This is referred to as the "centrality-lethality rule", which indicates that the topological placement of a protein in PPI network is connected with its biological essentiality. Though such connections are observed in many PPI networks, the underlying topological properties for these connections are not yet clearly understood. Some suggested putative connections are the involvement of essential proteins in the maintenance of overall network connections, or that they play a role in essential protein clusters. In this work, we have attempted to examine the placement of essential proteins and the network topology from a different perspective by determining the correlation of protein essentiality and reverse nearest neighbor topology (RNN). RESULTS: The RNN topology is a weighted directed graph derived from PPI network, and it is a natural representation of the topological dependences between proteins within the PPI network. Similar to the original PPI network, we have observed that essential proteins tend to be hub proteins in RNN topology. Additionally, essential genes are enriched in clusters containing many hub proteins in RNN topology (RNN protein clusters). Based on these two properties of essential genes in RNN topology, we have proposed a new measure; the RNN cluster centrality. Results from a variety of PPI networks demonstrate that RNN cluster centrality outperforms other centrality measures with regard to the proportion of selected proteins that are essential proteins. We also investigated the biological importance of RNN clusters. CONCLUSIONS: This study reveals that RNN cluster centrality provides the best correlation of protein essentiality and placement of proteins in PPI network. Additionally, merged RNN clusters were found to be topologically important in that essential proteins are significantly enriched in RNN clusters, and biologically important because they play an important role in many Gene Ontology (GO) processes. Kang Ning 0001, Hoong Kee Ng, Sriganesh Srihari, Hon Wai Leong, Alexey I. Nesvizhskii |
BMC Bioinform. | 1 |
| 2010 | MCL-CAw: a refinement of MCL for detecting yeast complexes from weighted PPI networks by incorporating core-attachment structureabstractBACKGROUND: The reconstruction of protein complexes from the physical interactome of organisms serves as a building block towards understanding the higher level organization of the cell. Over the past few years, several independent high-throughput experiments have helped to catalogue enormous amount of physical protein interaction data from organisms such as yeast. However, these individual datasets show lack of correlation with each other and also contain substantial number of false positives (noise). Over these years, several affinity scoring schemes have also been devised to improve the qualities of these datasets. Therefore, the challenge now is to detect meaningful as well as novel complexes from protein interaction (PPI) networks derived by combining datasets from multiple sources and by making use of these affinity scoring schemes. In the attempt towards tackling this challenge, the Markov Clustering algorithm (MCL) has proved to be a popular and reasonably successful method, mainly due to its scalability, robustness, and ability to work on scored (weighted) networks. However, MCL produces many noisy clusters, which either do not match known complexes or have additional proteins that reduce the accuracies of correctly predicted complexes. RESULTS: Inspired by recent experimental observations by Gavin and colleagues on the modularity structure in yeast complexes and the distinctive properties of "core" and "attachment" proteins, we develop a core-attachment based refinement method coupled to MCL for reconstruction of yeast complexes from scored (weighted) PPI networks. We combine physical interactions from two recent "pull-down" experiments to generate an unscored PPI network. We then score this network using available affinity scoring schemes to generate multiple scored PPI networks. The evaluation of our method (called MCL-CAw) on these networks shows that: (i) MCL-CAw derives larger number of yeast complexes and with better accuracies than MCL, particularly in the presence of natural noise; (ii) Affinity scoring can effectively reduce the impact of noise on MCL-CAw and thereby improve the quality (precision and recall) of its predicted complexes; (iii) MCL-CAw responds well to most available scoring schemes. We discuss several instances where MCL-CAw was successful in deriving meaningful complexes, and where it missed a few proteins or whole complexes due to affinity scoring of the networks. We compare MCL-CAw with several recent complex detection algorithms on unscored and scored networks, and assess the relative performance of the algorithms on these networks. Further, we study the impact of augmenting physical datasets with computationally inferred interactions for complex detection. Finally, we analyse the essentiality of proteins within predicted complexes to understand a possible correlation between protein essentiality and their ability to form complexes. CONCLUSIONS: We demonstrate that core-attachment based refinement in MCL-CAw improves the predictions of MCL on yeast PPI networks. We show that affinity scoring improves the performance of MCL-CAw. Sriganesh Srihari, Kang Ning 0001, Hon Wai Leong |
BMC Bioinform. | 2 |
| 2008 | Detecting hubs and quasi cliques in scale-free networksabstractScale-free networks are believed to closely model most real-world networks. An interesting property of such networks is the existence of so-called hub and community structures. In this paper, we model hubs as high-degree nodes and communities as quasi cliques. We propose a new problem formulation called the ¿-list dominating set and show how this single problem is suited to model both the structures in real-world networks better than traditional problems like vertex cover and clique. Additionally, we provide a fixed-parameter tractable algorithm to this detect these structures and show experimental results on protein-protein interaction networks. Sriganesh Srihari, Hoong Kee Ng, Kang Ning 0001, Hon Wai Leong |
ICPR | 3 |
| 2007 | De Novo Peptide Sequencing for Mass Spectra Based on Multi-Charge Strong Tags
Kang Ning 0001, Ket Fah Chong, Hon Wai Leong |
APBC | 1 |
| 2007 | A New Approach for Similarity Queries of Biological Sequences in Databases
Hoong Kee Ng, Kang Ning 0001, Hon Wai Leong |
PAKDD | 2 |
| 2006 | Characterization of Multi-Charge Mass Spectra for Peptide Sequencing
Ket Fah Chong, Kang Ning 0001, Hon Wai Leong, Pavel A. Pevzner |
APBC | 2 |
| 2006 | Finding Patterns in Biological Sequences by Longest Common Subsequencesand Shortest Common SupersequencesabstractPatterns in biological sequences are important for revealing the relationship among biological sequences. Much research has been done on this problem, and the sensitivity and specificity of current algorithms are already quite satisfactory. However, in general, for problems on a set of sequences, the relationship among their patterns, their Longest Common Subsequences (LCS) and their Shortest Common Supersequences (SCS) are not examined carefully. Therefore, revealing the relationship between the patterns and LCS/SCS might provide us with a deeper view of the patterns of biological sequences, in turn leading to a better understanding of them. In this paper, we propose the PALS (PAtterns by Lcs and Scs) algorithms to discover patterns in a set of biological sequences by first generating the results for LCS and SCS of sequences by heuristic, and consequently derive the patterns from these results. Experiments show that the PALS algorithms perform well (both in efficiencies and in accuracies) on a variety of sequences. Kang Ning 0001, Hoong Kee Ng, Hon Wai Leong |
BIBE | 1 |
| 2006 | Towards a better solution to the shortest common supersequence problem: the deposition and reduction algorithmabstractBACKGROUND: The problem of finding a Shortest Common Supersequence (SCS) of a set of sequences is an important problem with applications in many areas. It is a key problem in biological sequences analysis. The SCS problem is well-known to be NP-complete. Many heuristic algorithms have been proposed. Some heuristics work well on a few long sequences (as in sequence comparison applications); others work well on many short sequences (as in oligo-array synthesis). Unfortunately, most do not work well on large SCS instances where there are many, long sequences. RESULTS: In this paper, we present a Deposition and Reduction (DR) algorithm for solving large SCS instances of biological sequences. There are two processes in our DR algorithm: deposition process, and reduction process. The deposition process is responsible for generating a small set of common supersequences; and the reduction process shortens these common supersequences by removing some characters while preserving the common supersequence property. Our evaluation on simulated data and real DNA and protein sequences show that our algorithm consistently produces the best results compared to many well-known heuristic algorithms, and especially on large instances. CONCLUSION: Our DR algorithm provides a partial answer to the open problem of designing efficient heuristic algorithm for SCS problem on many long sequences. Our algorithm has a bounded approximation ratio. The algorithm is efficient, both in running time and space complexity and our evaluation shows that it is practical even for SCS problems on many long sequences. Kang Ning 0001, Hon Wai Leong |
BMC Bioinform. | 1 |