Chongle Pan

dblp:34/5659 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0003-2860-0334ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Benchmarking interpretability of deep learning for predictive genomics: Recall, precision, and variability of feature attribution
abstract
Deep neural networks can model the nonlinear architecture of polygenic traits, yet the reliability of attribution methods to identify the genetic variants driving model predictions remains uncertain. We introduce a benchmarking framework that quantifies three aspects of interpretability: attribution recall, attribution precision, and stability, and apply it to deep learning models trained on UK Biobank genotypes for standing height prediction. After quality control, feed-forward neural networks were trained on more than half a million autosomal variants from approximately 300 thousand participants and evaluated using four attribution algorithms (Saliency, Gradient SHAP, DeepLIFT, Integrated Gradients) with and without SmoothGrad noise averaging. Attribution recall was assessed using synthetic spike-in variants with known additive, dominant, recessive, and epistatic effects, enabling direct measurement of sensitivity to diverse genetic architectures. Attribution precision estimated specificity using an equal number of null decoy variants that preserved allele structure while disrupting genotype-phenotype correspondence. Stability was measured by the consistency of variant-level attributions across an ensemble of independently trained models. SmoothGrad increased average recall across effect types by approximately 0.16 at the top 1% of the most highly attributed variants and improved average precision by about 0.06 at the same threshold, while stability remained comparable with median relative standard deviations of 0.4 to 0.5 across methods. Among the evaluated attribution methods, Saliency achieved the highest composite score, indicating that its simple gradient formulation provided the best overall balance of recall, precision, and stability.
Justin C. Reynolds, Chongle Pan
PLoS Comput. Biol.2
2024 Towards AI-Driven Healthcare: Systematic Optimization, Linguistic Analysis, and Clinicians' Evaluation of Large Language Models for Smoking Cessation Interventions
abstract
Creating intervention messages for smoking cessation is a labor-intensive process. Advances in Large Language Models (LLMs) offer a promising alternative for automated message generation. Two critical questions remain: 1) How to optimize LLMs to mimic human expert writing, and 2) Do LLM-generated messages meet clinical standards? We systematically examined the message generation and evaluation processes through three studies investigating prompt engineering (Study 1), decoding optimization (Study 2), and expert review (Study 3). We employed computational linguistic analysis in LLM assessment and established a comprehensive evaluation framework, incorporating automated metrics, linguistic attributes, and expert evaluations. Certified tobacco treatment specialists assessed the quality, accuracy, credibility, and persuasiveness of LLM-generated messages, using expert-written messages as the benchmark. Results indicate that larger LLMs, including ChatGPT, OPT-13B, and OPT-30B, can effectively emulate expert writing to generate well-written, accurate, and persuasive messages, thereby demonstrating the capability of LLMs in augmenting clinical practices of smoking cessation interventions.
Paul Calle, Ruosi Shao, Yunlong Liu 0004, Emily T. Hébert, Darla E. Kendzor, Jordan M. Neil, Michael S. Businelle, Chongle Pan
CHI8
2024 SEMQuant: Extending Sipros-Ensemble with Match-Between-Runs for Comprehensive Quantitative Metaproteomics
Bailu Zhang, Shichao Feng, Manushi Parajuli, Chongle Pan, Xuan Guo 0004
ISBRA (3)5
2023 Explainable multi-task learning improves the parallel estimation of polygenic risk scores for many diseases through shared genetic basis
abstract
Many complex diseases share common genetic determinants and are comorbid in a population. We hypothesized that the co-occurrences of diseases and their overlapping genetic etiology can be exploited to simultaneously improve multiple diseases' polygenic risk scores (PRS). This hypothesis was tested using a multi-task learning (MTL) approach based on an explainable neural network architecture. We found that parallel estimations of the PRS for 17 prevalent cancers in a pan-cancer MTL model were generally more accurate than independent estimations for individual cancers in comparable single-task learning (STL) models. Such performance improvement conferred by positive transfer learning was also observed consistently for 60 prevalent non-cancer diseases in a pan-disease MTL model. Interpretation of the MTL models revealed significant genetic correlations between the important sets of single nucleotide polymorphisms used by the neural network for PRS estimation. This suggested a well-connected network of diseases with shared genetic basis.
Adrien Badré, Chongle Pan
PLoS Comput. Biol.2
2022 IDIA: An Integrative Signal Extractor for Data-Independent Acquisition Proteomics
abstract
In proteomics, data-independent acquisition (DIA) has been shown to provide less biased and more reproducible results than data-dependent acquisition. Recently, many researchers have developed a series of methods to identify peptides and proteins by using spectrum libraries for DIA data. However, spectrum libraries are not always available for novel organisms or microbial communities. To detect peptides and proteins without a spectrum library, we developed IDIA, a library-free method using DIA data to generate pseudo-spectra that can be searched using conventional sequence database searching software. IDIA integrates two isotopic trace detection strategies and employs B-spline and Gaussian filters to help extract high-quality pseudo-spectra from the complex DIA data. The experimental results on human and yeast data demonstrated that our approach remarkably produced more peptide and protein identifications than the two state-of-the-art library-free methods, i.e., DIA-Umpire and Group-DIA. IDIA is freely available under the GNU GPL license at https://github.com/Biocomputing-Research-Group/IDIA.
Jiancheng Li, Chongle Pan, Xuan Guo 0004
BIBM2
2022 FineFDR: Fine-grained Taxonomy-specific False Discovery Rates Control in Metaproteomics
abstract
Microbial community proteomics, also termed metaproteomics, investigates all proteins expressed by a microbiota. Tandem mass spectrometry (MS/MS) is the typical method for identifying proteins in metaproteomics, which involves searching the mass spectra against a protein sequence database. A major post-analysis step is controlling the false discovery rate (FDR), i.e., the ratio of false positives to the total number of annotations. The current popular target-decoy FDR estimation method treats all the peptides and proteins equally and overlooks that they could have varied probabilities of being identified. In this study, we report FineFDR, a framework for FDR assessment at fine-grained levels with taxonomy information considered. FineFDR groups the identified peptide-spectrum matches, peptides, and proteins from different taxonomic units and estimates the FDR in each group separately. Empirical experiments on the simulated and real-world data sets demonstrate that our FineFDR achieved higher precision and more peptide and protein identifications when compared to the state-of-the-art methods, such as Comet, Percolator, TIDD, and Tailor. FineFDR is freely available under the GNU GPL license at https://github.com/Biocomputing-Research-Group/FDR.
Shengze Wang 0004, Shichao Feng, Chongle Pan, Xuan Guo 0004
BIBM3
2022 MetaLP: An integrative linear programming method for protein inference in metaproteomics
abstract
Metaproteomics based on high-throughput tandem mass spectrometry (MS/MS) plays a crucial role in characterizing microbiome functions. The acquired MS/MS data is searched against a protein sequence database to identify peptides, which are then used to infer a list of proteins present in a metaproteome sample. While the problem of protein inference has been well-studied for proteomics of single organisms, it remains a major challenge for metaproteomics of complex microbial communities because of the large number of degenerate peptides shared among homologous proteins in different organisms. This challenge calls for improved discrimination of true protein identifications from false protein identifications given a set of unique and degenerate peptides identified in metaproteomics. MetaLP was developed here for protein inference in metaproteomics using an integrative linear programming method. Taxonomic abundance information extracted from metagenomics shotgun sequencing or 16s rRNA gene amplicon sequencing, was incorporated as prior information in MetaLP. Benchmarking with mock, human gut, soil, and marine microbial communities demonstrated significantly higher numbers of protein identifications by MetaLP than ProteinLP, PeptideProphet, DeepPep, PIPQ, and Sipros Ensemble. In conclusion, MetaLP could substantially improve protein inference for complex metaproteomes by incorporating taxonomic abundance information in a linear programming model.
Shichao Feng, Hong-Long Ji, Huan Wang 0005, Bailu Zhang, Ryan Sterzenbach, Chongle Pan, Xuan Guo 0004
PLoS Comput. Biol.6
2018 SORA: Scalable Overlap-graph Reduction Algorithms for Genome Assembly using Apache Spark in the Cloud
Alexander J. Paul, Dylan Lawrence, Myoungkyu Song, Seung-Hwan Lim, Chongle Pan, Tae-Hyuk Ahn
BIBM5
2018 Sipros Ensemble improves database searching and filtering for complex metaproteomics
abstract
Motivation: Complex microbial communities can be characterized by metagenomics and metaproteomics. However, metagenome assemblies often generate enormous, and yet incomplete, protein databases, which undermines the identification of peptides and proteins in metaproteomics. This challenge calls for increased discrimination of true identifications from false identifications by database searching and filtering algorithms in metaproteomics. Results: Sipros Ensemble was developed here for metaproteomics using an ensemble approach. Three diverse scoring functions from MyriMatch, Comet and the original Sipros were incorporated within a single database searching engine. Supervised classification with logistic regression was used to filter database searching results. Benchmarking with soil and marine microbial communities demonstrated a higher number of peptide and protein identifications by Sipros Ensemble than MyriMatch/Percolator, Comet/Percolator, MS-GF+/Percolator, Comet & MyriMatch/iProphet and Comet & MyriMatch & MS-GF+/iProphet. Sipros Ensemble was computationally efficient and scalable on supercomputers. Availability and implementation: Freely available under the GNU GPL license at http://sipros.omicsbio.org. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Xuan Guo 0004, Qiuming Yao, Ryan S. Mueller, Jimmy K. Eng, David L. Tabb, IV William Judson Hervey, Chongle Pan
Bioinform.8
2015 Sigma: Strain-level inference of genomes from metagenomic analysis for biosurveillance
abstract
MOTIVATION: Metagenomic sequencing of clinical samples provides a promising technique for direct pathogen detection and characterization in biosurveillance. Taxonomic analysis at the strain level can be used to resolve serotypes of a pathogen in biosurveillance. Sigma was developed for strain-level identification and quantification of pathogens using their reference genomes based on metagenomic analysis. RESULTS: Sigma provides not only accurate strain-level inferences, but also three unique capabilities: (i) Sigma quantifies the statistical uncertainty of its inferences, which includes hypothesis testing of identified genomes and confidence interval estimation of their relative abundances; (ii) Sigma enables strain variant calling by assigning metagenomic reads to their most likely reference genomes; and (iii) Sigma supports parallel computing for fast analysis of large datasets. The algorithm performance was evaluated using simulated mock communities and fecal samples with spike-in pathogen strains. AVAILABILITY AND IMPLEMENTATION: Sigma was implemented in C++ with source codes and binaries freely available at http://sigma.omicsbio.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tae-Hyuk Ahn, Juanjuan Chai, Chongle Pan
Bioinform.3
2014 Omega: an Overlap-graph de novo Assembler for Metagenomics
abstract
MOTIVATION: Metagenomic sequencing allows reconstruction of microbial genomes directly from environmental samples. Omega (overlap-graph metagenome assembler) was developed for assembling and scaffolding Illumina sequencing data of microbial communities. RESULTS: Omega found overlaps between reads using a prefix/suffix hash table. The overlap graph of reads was simplified by removing transitive edges and trimming short branches. Unitigs were generated based on minimum cost flow analysis of the overlap graph and then merged to contigs and scaffolds using mate-pair information. In comparison with three de Bruijn graph assemblers (SOAPdenovo, IDBA-UD and MetaVelvet), Omega provided comparable overall performance on a HiSeq 100-bp dataset and superior performance on a MiSeq 300-bp dataset. In comparison with Celera on the MiSeq dataset, Omega provided more continuous assemblies overall using a fraction of the computing time of existing overlap-layout-consensus assemblers. This indicates Omega can more efficiently assemble longer Illumina reads, and at deeper coverage, for metagenomic datasets. AVAILABILITY AND IMPLEMENTATION: Implemented in C++ with source code and binaries freely available at http://omega.omicsbio.org.
Bahlul Haider, Tae-Hyuk Ahn, Brian Bushnell, Juanjuan Chai, Alex Copeland, Chongle Pan
Bioinform.6
2013 Sipros/ProRata: a versatile informatics system for quantitative community proteomics
abstract
SUMMARY: Sipros/ProRata is an open-source software package for end-to-end data analysis in a wide variety of community proteomics measurements. A database-searching program, Sipros 3.0, was developed for accurate general-purpose protein identification and broad-range post-translational modification searches. Hybrid Message Passing Interface/OpenMP parallelism of the new Sipros architecture allowed its computation to be scalable from desktops to supercomputers. The upgraded ProRata 3.0 performs label-free quantification and isobaric chemical labeling quantification in addition to metabolic labeling quantification. Sipros/ProRata is a versatile informatics system that enables identification and quantification of proteins and their variants in many types of community proteomics studies. AVAILABILITY: Both programs are freely available under the GNU GPL license at Sipros.omicsbio.org and ProRata.omicsbio.org.
Yingfeng Wang, Tae-Hyuk Ahn, Chongle Pan
Bioinform.4
2012 Exhaustive database searching for amino acid mutations in proteomes
abstract
MOTIVATION: Amino acid mutations in proteins can be found by searching tandem mass spectra acquired in shotgun proteomics experiments against protein sequences predicted from genomes. Traditionally, unconstrained searches for amino acid mutations have been accomplished by using a sequence tagging approach that combines de novo sequencing with database searching. However, this approach is limited by the performance of de novo sequencing. RESULTS: The Sipros algorithm v2.0 was developed to perform unconstrained database searching using high-resolution tandem mass spectra by exhaustively enumerating all single non-isobaric mutations for every residue in a protein database. The performance of Sipros for amino acid mutation identification exceeded that of an established sequence tagging algorithm, Inspect, based on benchmarking results from a Rhodopseudomonas palustris proteomics dataset. To demonstrate the viability of the algorithm for meta-proteomics, Sipros was used to identify amino acid mutations in a natural microbial community in acid mine drainage. AVAILABILITY: The Sipros algorithm is freely available at\newline http://code.google.com/p/sipros.
Doug Hyatt, Chongle Pan
Bioinform.2
2010 A high-throughput de novo sequencing approach for shotgun proteomics using high-resolution tandem mass spectrometry
abstract
BACKGROUND: High-resolution tandem mass spectra can now be readily acquired with hybrid instruments, such as LTQ-Orbitrap and LTQ-FT, in high-throughput shotgun proteomics workflows. The improved spectral quality enables more accurate de novo sequencing for identification of post-translational modifications and amino acid polymorphisms. RESULTS: In this study, a new de novo sequencing algorithm, called Vonode, has been developed specifically for analysis of such high-resolution tandem mass spectra. To fully exploit the high mass accuracy of these spectra, a unique scoring system is proposed to evaluate sequence tags based primarily on mass accuracy information of fragment ions. Consensus sequence tags were inferred for 11,422 spectra with an average peptide length of 5.5 residues from a total of 40,297 input spectra acquired in a 24-hour proteomics measurement of Rhodopseudomonas palustris. The accuracy of inferred consensus sequence tags was 84%. According to our comparison, the performance of Vonode was shown to be superior to the PepNovo v2.0 algorithm, in terms of the number of de novo sequenced spectra and the sequencing accuracy. CONCLUSIONS: Here, we improved de novo sequencing performance by developing a new algorithm specifically for high-resolution tandem mass spectral data. The Vonode algorithm is freely available for download at http://compbio.ornl.gov/Vonode.
Chongle Pan, W. Hayes McDonald, Patricia A. Carey, Jillian F. Banfield, Nathan Verberkmoes, Robert L. Hettich, Nagiza F. Samatova
BMC Bioinform.1
2005 A graph-theoretic approach for the separation of b and y ions in tandem mass spectra
abstract
MOTIVATION: Ion-type identification is a fundamental problem in computational proteomics. Methods for accurate identification of ion types provide the basis for many mass spectrometry data interpretation problems, including (a) de novo sequencing, (b) identification of post-translational modifications and mutations and (c) validation of database search results. RESULTS: Here, we present a novel graph-theoretic approach for solving the problem of separating b ions from y ions in a set of tandem mass spectra. We represent each spectral peak as a node and consider two types of edges: type-1 edge connecting two peaks probably of the same ion types and type-2 edge connecting two peaks probably of different ion types. The problem of ion-separation is formulated and solved as a graph partition problem, which is to partition the graph into three subgraphs, representing b, y and others ions, respectively, through maximizing the total weight of type-1 edges while minimizing the total weight of type-2 edges within each partitioned subgraph. We have developed a dynamic programming algorithm for rigorously solving this graph partition problem and implemented it as a computer program PRIME (PaRtition of Ion types in tandem Mass spEctra). The tests on a large amount of simulated mass spectra and 19 sets of high-quality experimental Fourier transform ion cyclotron resonance tandem mass spectra indicate that an accuracy level of approximately 90% for the separation of b and y ions was achieved. AVAILABILITY: The executable code of PRIME is available upon request. CONTACT: [email protected].
Chongle Pan, Victor Olman, Robert L. Hettich, Ying Xu 0001
Bioinform.2