Hunter N. B. Moseley

dblp:22/5162 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0003-3995-5368ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Predicting the pathway involvement of metabolites annotated in the MetaCyc knowledgebase
abstract
BACKGROUND: The associations of metabolites with biochemical pathways are highly useful information for interpreting molecular datasets generated in biological and biomedical research. However, such pathway annotations are sparse in most molecular datasets, limiting their utility for pathway level interpretation. To address these shortcomings, several past publications have presented machine learning models for predicting the pathway association of small biomolecule (metabolite and xenobiotic) using data from the Kyoto Encyclopedia of Genes and Genomes (KEGG). But other similar knowledgebases exist, for example MetaCyc, which has more compound entries and pathway definitions than KEGG. RESULTS: As a logical next step, we trained and evaluated multilayer perceptron models on compound entries and pathway annotations obtained from MetaCyc. From the models trained on this dataset, we observed a mean Matthews correlation coefficient (MCC) of 0.845 with 0.0101 standard deviation, compared to a mean MCC of 0.847 with 0.0098 standard deviation for the KEGG dataset. However, KEGG's 184 metabolic-only pathway predictions (out of 502 total pathways) have a mean MCC of 0.800 with 0.021 standard deviation. Since MetaCyc pathways are metabolic focused, the MetaCyc results represent over a 5.6% improvement in metabolic pathway prediction performance. CONCLUSIONS: These performance results are pragmatically the same, demonstrating that in aggregate, the 4055 MetaCyc pathways can be effectively predicted at the current state-of-the-art performance level.
Erik Huckvale, Hunter N. B. Moseley
BMC Bioinform.2
2023 kegg_pull: a software package for the RESTful access and pulling from the Kyoto Encyclopedia of Gene and Genomes
abstract
BACKGROUND: The Kyoto Encyclopedia of Genes and Genomes (KEGG) provides organized genomic, biomolecular, and metabolic information and knowledge that is reasonably current and highly useful for a wide range of analyses and modeling. KEGG follows the principles of data stewardship to be findable, accessible, interoperable, and reusable (FAIR) by providing RESTful access to their database entries via their web-accessible KEGG API. However, the overall FAIRness of KEGG is often limited by the library and software package support available in a given programming language. While R library support for KEGG is fairly strong, Python library support has been lacking. Moreover, there is no software that provides extensive command line level support for KEGG access and utilization. RESULTS: We present kegg_pull, a package implemented in the Python programming language that provides better KEGG access and utilization functionality than previous libraries and software packages. Not only does kegg_pull include an application programming interface (API) for Python programming, it also provides a command line interface (CLI) that enables utilization of KEGG for a wide range of shell scripting and data analysis pipeline use-cases. As kegg_pull's name implies, both the API and CLI provide versatile options for pulling (downloading and saving) an arbitrary (user defined) number of database entries from the KEGG API. Moreover, this functionality is implemented to efficiently utilize multiple central processing unit cores as demonstrated in several performance tests. Many options are provided to optimize fault-tolerant performance across a single or multiple processes, with recommendations provided based on extensive testing and practical network considerations. CONCLUSIONS: The new kegg_pull package enables new flexible KEGG retrieval use cases not available in previous software packages. The most notable new feature that kegg_pull provides is its ability to robustly pull an arbitrary number of KEGG entries with a single API method or CLI command, including pulling an entire KEGG database. We provide recommendations to users for the most effective use of kegg_pull according to their network and computational circumstances.
Erik Huckvale, Hunter N. B. Moseley
BMC Bioinform.2
2023 The metabolomics workbench file status website: a metadata repository promoting FAIR principles of metabolomics data
abstract
BACKGROUND: An updated version of the mwtab Python package for programmatic access to the Metabolomics Workbench (MetabolomicsWB) data repository was released at the beginning of 2021. Along with updating the package to match the changes to MetabolomicsWB's 'mwTab' file format specification and enhancing the package's functionality, the included validation facilities were used to detect and catalog file inconsistencies and errors across all publicly available datasets in MetabolomicsWB. RESULTS: The MetabolomicsWB File Status website was developed to provide continuous validation of MetabolomicsWB data files and a useful interface to all found inconsistencies and errors. This list of detectable issues/errors include format parsing errors, format compliance issues, access problems via MetabolomicsWB's REST interface, and other small inconsistencies that can hinder reusability. The website uses the mwtab Python package to pull down and validate each available analysis file and then generates an html report. The website is updated on a weekly basis. Moreover, the Python website design utilizes GitHub and GitHub.io, providing an easy to replicate template for implementing other metadata, virtual, and meta- repositories. CONCLUSIONS: The MetabolomicsWB File Status website provides a metadata repository of validation metadata to promote the FAIR use of existing metabolomics datasets from the MetabolomicsWB data repository.
Christian D. Powell, Hunter N. B. Moseley
BMC Bioinform.2
2021 MEScan: a powerful statistical framework for genome-scale mutual exclusivity analysis of cancer mutations
abstract
MOTIVATION: Cancer somatic driver mutations associated with genes within a pathway often show a mutually exclusive pattern across a cohort of patients. This mutually exclusive mutational signal has been frequently used to distinguish driver from passenger mutations and to investigate relationships among driver mutations. Current methods for de novo discovery of mutually exclusive mutational patterns are limited because the heterogeneity in background mutation rate can confound mutational patterns, and the presence of highly mutated genes can lead to spurious patterns. In addition, most methods only focus on a limited number of pre-selected genes and are unable to perform genome-wide analysis due to computational inefficiency. RESULTS: We introduce a statistical framework, MEScan, for accurate and efficient mutual exclusivity analysis at the genomic scale. Our framework contains a fast and powerful statistical test for mutual exclusivity with adjustment of the background mutation rate and impact of highly mutated genes, and a multi-step procedure for genome-wide screening with the control of false discovery rate. We demonstrate that MEScan more accurately identifies mutually exclusive gene sets than existing methods and is at least two orders of magnitude faster than most methods. By applying MEScan to data from four different cancer types and pan-cancer, we have identified several biologically meaningful mutually exclusive gene sets. AVAILABILITY AND IMPLEMENTATION: MEScan is available as an R package at https://github.com/MarkeyBBSRF/MEScan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sisheng Liu, Yanqi Xie, Tingting Zhai, Eugene W. Hinderer, Arnold J. Stromberg, Nathan L. Vanderford, Jill M. Kolesar, Hunter N. B. Moseley, Chunming Liu, Chi Wang 0003
Bioinform.9
2020 SSIF: Subsumption-based Sub-term Inference Framework to audit Gene Ontology
abstract
MOTIVATION: The Gene Ontology (GO) is the unifying biological vocabulary for codifying, managing and sharing biological knowledge. Quality issues in GO, if not addressed, can cause misleading results or missed biological discoveries. Manual identification of potential quality issues in GO is a challenging and arduous task, given its growing size. We introduce an automated auditing approach for suggesting potentially missing is-a relations, which may further reveal erroneous is-a relations. RESULTS: We developed a Subsumption-based Sub-term Inference Framework (SSIF) by leveraging a novel term-algebra on top of a sequence-based representation of GO concepts along with three conditional rules (monotonicity, intersection and sub-concept rules). Applying SSIF to the October 3, 2018 release of GO suggested 1938 unique potentially missing is-a relations. Domain experts evaluated a random sample of 210 potentially missing is-a relations. The results showed SSIF achieved a precision of 60.61, 60.49 and 46.03% for the monotonicity, intersection and sub-concept rules, respectively. AVAILABILITY AND IMPLEMENTATION: SSIF is implemented in Java. The source code is available at https://github.com/rashmie/SSIF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui
Bioinform.3
2019 Moiety modeling framework for deriving moiety abundances from mass spectrometry measured isotopologues
abstract
Abstract Background Stable isotope tracing can follow individual atoms through metabolic transformations through the detection of the incorporation of stable isotope within metabolites. This resulting data can be interpreted in terms related to metabolic flux. However, detection of a stable isotope in metabolites by mass spectrometry produces a profile of isotopologue peaks that requires deconvolution to ascertain the localization of isotope incorporation. Results To aid the interpretation of the mass spectroscopy isotopologue profile, we have developed a moiety modeling framework for deconvoluting metabolite isotopologue profiles involving single and multiple isotope tracers. This moiety modeling framework provides facilities for moiety model representation, moiety model optimization, and moiety model selection. The moiety_modeling package was developed from the idea of metabolite decomposition into moiety units based on metabolic transformations, i.e. a moiety model. The SAGA-optimize package, solving a boundary-value inverse problem through a combined simulated annealing and genetic algorithm, was developed for model optimization. Additional optimization methods from the Python scipy library are utilized as well. Several forms of the Akaike information criterion and Bayesian information criterion are provided for selecting between moiety models. Moiety models and associated isotopologue data are defined in a JSONized format. By testing the moiety modeling framework on the timecourses of 13 C isotopologue data for uridine diphosphate N-acetyl-D-glucosamine (UDP-GlcNAc) in human prostate cancer LnCaP-LN3 cells, we were able to confirm its robust performance in isotopologue deconvolution and moiety model selection. Conclusions SAGA-optimize is a useful Python package for solving boundary-value inverse problems, and the moiety_modeling package is an easy-to-use tool for mass spectroscopy isotopologue profile deconvolution involving single and multiple isotope tracers. Both packages are freely available on GitHub and via the Python Package Index.
Huan Jin, Hunter N. B. Moseley
BMC Bioinform.2
2018 A Lexical Approach to Identifying Subtype Inconsistencies in Biomedical Terminologies
Rashmie Abeysinghe, Fengbo Zheng, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui
BIBM4
2017 Auditing subtype inconsistencies among gene ontology concepts
abstract
Gene Ontology (GO) provides a controlled vocabulary for describing genes and related gene products. Quality assurance of Gene ontology (GO) is a vital aspect of the terminology management lifecycle. In this paper, we introduce a lexical-based inference approach to detecting subtype (or isa) inconsistencies among GO terms (i.e., biological concepts). We first model the name of each concept as a set of words. Then, we generate hierarchically linked and unlinked pairs of concepts (A, B), where A and B have the same number of words, and contain common words as well as a single different word. Each linked concept-pair infers a linked term-pair, and each unlinked concept-pair infers an unlinked term-pair. A term-pair appearing as both linked and unlinked is considered a potential inconsistency, which may represent a subtype inconsistency between the original linked and unlinked concept-pair. Applying this approach to the 03/28/2017 release of GO, a total of 3,715 potential subtype inconsistencies were obtained. Evaluation of a random sample of potential inconsistencies revealed two types of potential errors: missing subtype relations and incorrect subtype relations in GO, and achieved an accuracy of 56.33% for detecting such errors. This indicates that this lexical-based inference approach using the set-of-words model is a promising way to facilitate quality improvement of GO.
Rashmie Abeysinghe, Eugene W. Hinderer, Hunter N. B. Moseley, Licong Cui
BIBM3
2017 Proceedings of the 16th Annual UT-KBRIN Bioinformatics Summit 2016: bioinformatics: Burns, TN, USA. April 21-23, 2017
abstract
Memphis, Tennessee
Eric C. Rouchka, Julia H. Chariker, David Tieri, Juw Won Park, Shreedharkumar D. Rajurkar, Nishchal K. Verma, Yan Cui 0001, Mark L. Farman, Bradford Condon, Neil Moore, Jerzy W. Jaromczyk, Jolanta Jaromczyk, Daniel R. Harris, Patrick Calie, Eun Kyong Shin, Robert L. Davis, Arash Shaban-Nejad, Joshua M. Mitchell, Robert M. Flight, Qing Jun Wang, Richard M. Higashi, Teresa W.-M. Fan, Andrew N. Lane, Hunter N. B. Moseley, Liangqun Lu, Bernie J. Daigle, Andrey Smelter, Bailey K. Phan, Nathaniel J. Serpico, Ethan G. Toney, Caroline E. Melton, Jennifer R. Mandel, Bernie J. Daigle Jr., Kazi I. Zaman, Ramin Homayouni, Patrick J. Trainor, Samantha M. Carlisle, Andrew P. DeFilippis, Shesh N. Rai
BMC Bioinform.25
2017 A fast and efficient python library for interfacing with the Biological Magnetic Resonance Data Bank
abstract
BACKGROUND: The Biological Magnetic Resonance Data Bank (BMRB) is a public repository of Nuclear Magnetic Resonance (NMR) spectroscopic data of biological macromolecules. It is an important resource for many researchers using NMR to study structural, biophysical, and biochemical properties of biological macromolecules. It is primarily maintained and accessed in a flat file ASCII format known as NMR-STAR. While the format is human readable, the size of most BMRB entries makes computer readability and explicit representation a practical requirement for almost any rigorous systematic analysis. RESULTS: To aid in the use of this public resource, we have developed a package called nmrstarlib in the popular open-source programming language Python. The nmrstarlib's implementation is very efficient, both in design and execution. The library has facilities for reading and writing both NMR-STAR version 2.1 and 3.1 formatted files, parsing them into usable Python dictionary- and list-based data structures, making access and manipulation of the experimental data very natural within Python programs (i.e. "saveframe" and "loop" records represented as individual Python dictionary data structures). Another major advantage of this design is that data stored in original NMR-STAR can be easily converted into its equivalent JavaScript Object Notation (JSON) format, a lightweight data interchange format, facilitating data access and manipulation using Python and any other programming language that implements a JSON parser/generator (i.e., all popular programming languages). We have also developed tools to visualize assigned chemical shift values and to convert between NMR-STAR and JSONized NMR-STAR formatted files. Full API Reference Documentation, User Guide and Tutorial with code examples are also available. We have tested this new library on all current BMRB entries: 100% of all entries are parsed without any errors for both NMR-STAR version 2.1 and version 3.1 formatted files. We also compared our software to three currently available Python libraries for parsing NMR-STAR formatted files: PyStarLib, NMRPyStar, and PyNMRSTAR. CONCLUSIONS: The nmrstarlib package is a simple, fast, and efficient library for accessing data from the BMRB. The library provides an intuitive dictionary-based interface with which Python programs can read, edit, and write NMR-STAR formatted files and their equivalent JSONized NMR-STAR files. The nmrstarlib package can be used as a library for accessing and manipulating data stored in NMR-STAR files and as a command-line tool to convert from NMR-STAR file format into its equivalent JSON file format and vice versa, and to visualize chemical shift values. Furthermore, the nmrstarlib implementation provides a guide for effectively JSONizing other older scientific formats, improving the FAIRness of data in these formats.
Andrey Smelter, Morgan Astra, Hunter N. B. Moseley
BMC Bioinform.3
2014 Development of large-scale metabolite identification methods for metabolomics
abstract
Background Large-scale identification of metabolites is key to elucidating and modeling metabolism at the systems level. Advances in metabolomics technologies, particularly ultrahigh resolution mass spectrometry enable comprehensive and rapid analysis of metabolites, which is impractical to achieve by conventional methods. However, a significant barrier to meaningful data interpretation is the identification of a wide range of metabolites including unknowns and the determination of their role(s) in various metabolic networks. Our recent development of chemoselective (CS) probes to tag metabolite functional groups provides additional structural constraints for metabolite identification, but remains limited by the lack of functional groupresolved metabolite databases.
Joshua M. Mitchell, Teresa W.-M. Fan, Andrew N. Lane, Hunter N. B. Moseley
BMC Bioinform.4
2014 Coordination characterization of zinc metalloproteins
abstract
Background Zinc metalloproteins play a key role in all three domains of life. Prior studies have explored the use of zinc coordination characteristics for the prediction of functions in metalloproteins. However, prediction methods have been hampered by the inaccuracy in characterizing the zinc coordination environment. In order to minimize the impact of this factor, we improved our method to better characterize zinc coordination.
Sen Yao, Robert M. Flight, Hunter N. B. Moseley
BMC Bioinform.3
2012 Proceedings of the Eleventh Annual UT-ORNL-KBRIN Bioinformatics Summit 2012
abstract
The University of Tennessee (UT), the Oak Ridge National Laboratory (ORNL), and the Kentucky Biomedical Research Infrastructure Network (KBRIN), have collaborated over the past eleven years to share research and educational expertise in bioinformatics.One result of this collaboration is the joint sponsorship of an annual regional summit to bring together researchers, educators and students who are interested in bioinformatics from a variety of research and educational institutions.This summit provides unique opportunities for collaboration and forging links between members of the various institutions.This year, the Eleventh Annual UT-ORNL-KBRIN Bioinformatics Summit was held at the Seelbach Hilton Hotel in Louisville, Kentucky from March 30-April 1, 2012.A total of 232 participants pre-registered for the summit, with 126 from various Kentucky institutions and 80 from various Tennessee institutions.A number of additional participants came from universities and research institutions from other states and countries, e.g.
Eric C. Rouchka, Robert M. Flight, Hunter N. B. Moseley
BMC Bioinform.3
2010 Correcting for the effects of natural abundance in stable isotope resolved metabolomics experiments involving ultra-high resolution mass spectrometry
abstract
BACKGROUND: Stable isotope tracing with ultra-high resolution Fourier transform-ion cyclotron resonance-mass spectrometry (FT-ICR-MS) can provide simultaneous determination of hundreds to thousands of metabolite isotopologue species without the need for chromatographic separation. Therefore, this experimental metabolomics methodology may allow the tracing of metabolic pathways starting from stable-isotope-enriched precursors, which can improve our mechanistic understanding of cellular metabolism. However, contributions to the observed intensities arising from the stable isotope's natural abundance must be subtracted (deisotoped) from the raw isotopologue peaks before interpretation. Previously posed deisotoping problems are sidestepped due to the isotopic resolution and identification of individual isotopologue peaks. This peak resolution and identification come from the very high mass resolution and accuracy of FT-ICR-MS and present an analytically solvable deisotoping problem, even in the context of stable-isotope enrichment. RESULTS: We present both a computationally feasible analytical solution and an algorithm to this newly posed deisotoping problem, which both work with any amount of 13C or 15N stable-isotope enrichment. We demonstrate this algorithm and correct for the effects of 13C natural abundance on a set of raw isotopologue intensities for a specific phosphatidylcholine lipid metabolite derived from a 13C-tracing experiment. CONCLUSIONS: Correction for the effects of 13C natural abundance on a set of raw isotopologue intensities is computationally feasible when the raw isotopologues are isotopically resolved and identified. Such correction makes qualitative interpretation of stable isotope tracing easier and is required before attempting a more rigorous quantitative interpretation of the isotopologue data. The presented implementation is very robust with increasing metabolite size. Error analysis of the algorithm will be straightforward due to low relative error from the implementation itself. Furthermore, the algorithm may serve as an independent quality control measure for a set of observed isotopologue intensities.
Hunter N. B. Moseley
BMC Bioinform.1