Ron M. Stewart

dblp:29/7067 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 SKiM-GPT: combining biomedical literature-based discovery with large language model hypothesis evaluation
abstract
BACKGROUND: Generating and testing hypotheses is a critical aspect of biomedical science. Typically, researchers generate hypotheses by carefully analyzing available information and making logical connections, which are then tested. The accelerating growth of biomedical literature makes it increasingly difficult to keep pace with connections between biological entities emerging across biomedical research. Recently developed automated means of generating hypotheses can generate many more hypotheses than can be easily tested. One such approach involves literature‑based discovery (LBD) systems such as Serial KinderMiner (SKiM), which surfaces putative A‑B‑C links derived from term co‑occurrence. However, LBD systems leave three critical gaps: (i) they find statistical associations, not biological relationships; (ii) they can produce false‑positive leads; and (iii) they do not assess agreement with a hypothesis in question. As a result, LBD search results often require costly manual curation to be of practical utility to the researcher. Large language models (LLMs) have the potential to automate much of this curation step, but standalone LLMs are hampered by hallucinations, lack of transparency in information sources, and the inability to reference data not included in the training corpus. RESULTS: We introduce SKiM-GPT, a retrieval-augmented generation (RAG) system that combines SKiM’s co-occurrence search and retrieval with frontier LLMs to evaluate user-defined hypotheses. For every chosen A-B-C SKiM hit, SKiM-GPT retrieves appropriate PubMed abstract texts, filters out irrelevant abstracts with a fine-tuned relevance model, and prompts an LLM to evaluate the user’s hypothesis, given the relevant abstracts. Importantly, the SKiM-GPT system is transparent and human-verifiable: it displays the retrieved abstracts, the hypothesis score, and a justification for the score grounded in the texts and written in natural language. On a benchmark consisting of 14 disease-gene-drug hypotheses, SKiM-GPT achieves strong ordinal agreement with four expert biologists (Cohen’s κ = 0.84), demonstrating its ability to replicate expert judgment. CONCLUSIONS: SKiM-GPT is open-source ( https://github.com/stewart-lab/skimgpt ) and available through a web interface ( https://skim.morgridge.org ), enabling both wet-lab and computational researchers to systematically and efficiently evaluate biomedical hypotheses at scale.
Jack Freeman, Robert J. Millikin, Leo Xu, Ishaan Sharma, Bethany Moore, Cannon Lock, Kevin Shine George, Aviral Bal, Chitrasen Mohanty, Ron M. Stewart
BMC Bioinform.10
2023 Serial KinderMiner (SKiM) discovers and annotates biomedical knowledge using co-occurrence and transformer models
abstract
BACKGROUND: The PubMed archive contains more than 34 million articles; consequently, it is becoming increasingly difficult for a biomedical researcher to keep up-to-date with different knowledge domains. Computationally efficient and interpretable tools are needed to help researchers find and understand associations between biomedical concepts. The goal of literature-based discovery (LBD) is to connect concepts in isolated literature domains that would normally go undiscovered. This usually takes the form of an A-B-C relationship, where A and C terms are linked through a B term intermediate. Here we describe Serial KinderMiner (SKiM), an LBD algorithm for finding statistically significant links between an A term and one or more C terms through some B term intermediate(s). The development of SKiM is motivated by the observation that there are only a few LBD tools that provide a functional web interface, and that the available tools are limited in one or more of the following ways: (1) they identify a relationship but not the type of relationship, (2) they do not allow the user to provide their own lists of B or C terms, hindering flexibility, (3) they do not allow for querying thousands of C terms (which is crucial if, for instance, the user wants to query connections between a disease and the thousands of available drugs), or (4) they are specific for a particular biomedical domain (such as cancer). We provide an open-source tool and web interface that improves on all of these issues. RESULTS: We demonstrate SKiM's ability to discover useful A-B-C linkages in three control experiments: classic LBD discoveries, drug repurposing, and finding associations related to cancer. Furthermore, we supplement SKiM with a knowledge graph built with transformer machine-learning models to aid in interpreting the relationships between terms found by SKiM. Finally, we provide a simple and intuitive open-source web interface ( https://skim.morgridge.org ) with comprehensive lists of drugs, diseases, phenotypes, and symptoms so that anyone can easily perform SKiM searches. CONCLUSIONS: SKiM is a simple algorithm that can perform LBD searches to discover relationships between arbitrary user-defined concepts. SKiM is generalized for any domain, can perform searches with many thousands of C term concepts, and moves beyond the simple identification of an existence of a relationship; many relationships are given relationship type labels from our knowledge graph.
Robert J. Millikin, Kalpana Raja, John W. Steill, Cannon Lock, Xuancheng Tu, Lam C. Tsoi, Finn Kuusisto, Zijian Ni, Miron Livny, Brian Bockelman, James A. Thomson, Ron M. Stewart
BMC Bioinform.13
2021 CHARTS: a web application for characterizing and comparing tumor subpopulations in publicly available single-cell RNA-seq data sets
abstract
BACKGROUND: Single-cell RNA-seq (scRNA-seq) enables the profiling of genome-wide gene expression at the single-cell level and in so doing facilitates insight into and information about cellular heterogeneity within a tissue. This is especially important in cancer, where tumor and tumor microenvironment heterogeneity directly impact development, maintenance, and progression of disease. While publicly available scRNA-seq cancer data sets offer unprecedented opportunity to better understand the mechanisms underlying tumor progression, metastasis, drug resistance, and immune evasion, much of the available information has been underutilized, in part, due to the lack of tools available for aggregating and analysing these data. RESULTS: We present CHARacterizing Tumor Subpopulations (CHARTS), a web application for exploring publicly available scRNA-seq cancer data sets in the NCBI's Gene Expression Omnibus. More specifically, CHARTS enables the exploration of individual gene expression, cell type, malignancy-status, differentially expressed genes, and gene set enrichment results in subpopulations of cells across tumors and data sets. Along with the web application, we also make available the backend computational pipeline that was used to produce the analyses that are available for exploration in the web application. CONCLUSION: CHARTS is an easy to use, comprehensive platform for exploring single-cell subpopulations within tumors across the ever-growing collection of public scRNA-seq cancer data sets. CHARTS is freely available at charts.morgridge.org.
Matthew N. Bernstein, Zijian Ni, Mark E. Burkard, Christina Kendziorski, Ron M. Stewart
BMC Bioinform.6
2021 Interspecies chimeric conditions affect the developmental rate of human pluripotent stem cells
abstract
Human pluripotent stem cells hold significant promise for regenerative medicine. However, long differentiation protocols and immature characteristics of stem cell-derived cell types remain challenges to the development of many therapeutic applications. In contrast to the slow differentiation of human stem cells in vitro that mirrors a nine-month gestation period, mouse stem cells develop according to a much faster three-week gestation timeline. Here, we tested if co-differentiation with mouse pluripotent stem cells could accelerate the differentiation speed of human embryonic stem cells. Following a six-week RNA-sequencing time course of neural differentiation, we identified 929 human genes that were upregulated earlier and 535 genes that exhibited earlier peaked expression profiles in chimeric cell cultures than in human cell cultures alone. Genes with accelerated upregulation were significantly enriched in Gene Ontology terms associated with neurogenesis, neuron differentiation and maturation, and synapse signaling. Moreover, chimeric mixed samples correlated with in utero human embryonic samples earlier than human cells alone, and acceleration was dose-dependent on human-mouse co-culture ratios. The altered gene expression patterns and developmental rates described in this report have implications for accelerating human stem cell differentiation and the use of interspecies chimeric embryos in developing human organs for transplantation.
Jared Brown, Chris Barry, Matthew T. Schmitz, Cara Argus, Jennifer M. Bolin, Michael P. Schwartz, Amy Van Aartsen, John W. Steill, Scott Swanson, Ron M. Stewart, James A. Thomson, Christina Kendziorski
PLoS Comput. Biol.10
2019 Machine Learning to Predict Developmental Neurotoxicity with High-Throughput Data from 2D Bio-Engineered Tissues
abstract
animal studies, and assays of animal and human primary cell cultures, suffer from challenges related to time, cost, and applicability to human physiology. Prior work has demonstrated success employing machine learning to predict developmental neurotoxicity using gene expression data collected from human 3D tissue models exposed to various compounds. The 3D model is biologically similar to developing neural structures, but its complexity necessitates extensive expertise and effort to employ. By instead focusing solely on constructing an assay of developmental neurotoxicity, we propose that a simpler 2D tissue model may prove sufficient. We thus compare the accuracy of predictive models trained on data from a 2D tissue model with those trained on data from a 3D tissue model, and find the 2D model to be substantially more accurate. Furthermore, we find the 2D model to be more robust under stringent gene set selection, whereas the 3D model suffers substantial accuracy degradation. While both approaches have advantages and disadvantages, we propose that our described 2D approach could be a valuable tool for decision makers when prioritizing neurotoxicity screening.
Finn Kuusisto, Vítor Santos Costa, Zhonggang Hou, James A. Thomson, David Page, Ron M. Stewart
ICMLA6
2019 Automated minute scale RNA-seq of pluripotent stem cell differentiation reveals early divergence of human and mouse gene expression kinetics
abstract
Pluripotent stem cells retain the developmental timing of their species of origin in vitro, an observation that suggests the existence of a cell-intrinsic developmental clock, yet the nature and machinery of the clock remain a mystery. We hypothesize that one possible component may lie in species-specific differences in the kinetics of transcriptional responses to differentiation signals. Using a liquid-handling robot, mouse and human pluripotent stem cells were exposed to identical neural differentiation conditions and sampled for RNA-sequencing at high frequency, every 4 or 10 minutes, for the first 10 hours of differentiation to test for differences in transcriptomic response rates. The majority of initial transcriptional responses occurred within a rapid window in the first minutes of differentiation for both human and mouse stem cells. Despite similarly early onsets of gene expression changes, we observed shortened and condensed gene expression patterns in mouse pluripotent stem cells compared to protracted trends in human pluripotent stem cells. Moreover, the speed at which individual genes were upregulated, as measured by the slopes of gene expression changes over time, was significantly faster in mouse compared to human cells. These results suggest that downstream transcriptomic response kinetics to signaling cues are faster in mouse versus human cells, and may offer a partial account for the vast differences in developmental rates across species.
Chris Barry, Matthew T. Schmitz, Cara Argus, Jennifer M. Bolin, Mitchell D. Probasco, Ning Leng, Bret Duffin, John W. Steill, Scott Swanson, Brian E. McIntosh, Ron M. Stewart, Christina Kendziorski, James A. Thomson, Rhonda Bacher
PLoS Comput. Biol.11
2016 Baseline Regularization for Computational Drug Repositioning with Longitudinal Observational Data
Zhaobin Kuang, James A. Thomson, Michael Caldwell, Peggy L. Peissig, Ron M. Stewart, David Page
IJCAI5
2016 Computational Drug Repositioning Using Continuous Self-Controlled Case Series
abstract
Computational Drug Repositioning (CDR) is the task of discovering potential new indications for existing drugs by mining large-scale heterogeneous drug-related data sources. Leveraging the patient-level temporal ordering information between numeric physiological measurements and various drug prescriptions provided in Electronic Health Records (EHRs), we propose a Continuous Self-controlled Case Series (CSCCS) model for CDR. As an initial evaluation, we look for drugs that can control Fasting Blood Glucose (FBG) level in our experiments. Applying CSCCS to the Marshfield Clinic EHR, well-known drugs that are indicated for controlling blood glucose level are rediscovered. Furthermore, some drugs with recent literature support for the potential effect of blood glucose level control are also identified.
Zhaobin Kuang, James A. Thomson, Michael Caldwell, Peggy L. Peissig, Ron M. Stewart, David Page
KDD5
2016 Quality control of single-cell RNA-seq by SinQC
abstract
UNLABELLED: Single-cell RNA-seq (scRNA-seq) is emerging as a promising technology for profiling cell-to-cell variability in cell populations. However, the combination of technical noise and intrinsic biological variability makes detecting technical artifacts in scRNA-seq samples particularly challenging. Proper detection of technical artifacts is critical to prevent spurious results during downstream analysis. In this study, we present 'Single-cell RNA-seq Quality Control' (SinQC), a method and software tool to detect technical artifacts in scRNA-seq samples by integrating both gene expression patterns and data quality information. We apply SinQC to nine different scRNA-seq datasets, and show that SinQC is a useful tool for controlling scRNA-seq data quality. AVAILABILITY AND IMPLEMENTATION: SinQC software and documents are available at http://www.morgridge.net/SinQC.html CONTACTS: : [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Peng Jiang 0015, James A. Thomson, Ron M. Stewart
Bioinform.3
2016 OEFinder: a user interface to identify and visualize ordering effects in single-cell RNA-seq data
abstract
UNLABELLED: A recent article identified an artifact in multiple single-cell RNA-seq (scRNA-seq) datasets generated by the Fluidigm C1 platform. Specifically, Leng et al. showed significantly increased gene expression in cells captured from sites with small or large plate output IDs. We refer to this artifact as an ordering effect (OE). Including OE genes in downstream analyses could lead to biased results. To address this problem, we developed a statistical method and software called OEFinder to identify a sorted list of OE genes. OEFinder is available as an R package along with user-friendly graphical interface implementations which allows users to check for potential artifacts in scRNA-seq data generated by the Fluidigm C1 platform. AVAILABILITY AND IMPLEMENTATION: OEFinder is freely available at https://github.com/lengning/OEFinder CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ning Leng, Jeea Choi, Li-Fang Chu, James A. Thomson, Christina Kendziorski, Ron M. Stewart
Bioinform.6
2015 EBSeq-HMM: a Bayesian approach for identifying gene-expression changes in ordered RNA-seq experiments
abstract
MOTIVATION: With improvements in next-generation sequencing technologies and reductions in price, ordered RNA-seq experiments are becoming common. Of primary interest in these experiments is identifying genes that are changing over time or space, for example, and then characterizing the specific expression changes. A number of robust statistical methods are available to identify genes showing differential expression among multiple conditions, but most assume conditions are exchangeable and thereby sacrifice power and precision when applied to ordered data. RESULTS: We propose an empirical Bayes mixture modeling approach called EBSeq-HMM. In EBSeq-HMM, an auto-regressive hidden Markov model is implemented to accommodate dependence in gene expression across ordered conditions. As demonstrated in simulation and case studies, the output proves useful in identifying differentially expressed genes and in specifying gene-specific expression paths. EBSeq-HMM may also be used for inference regarding isoform expression. AVAILABILITY AND IMPLEMENTATION: An R package containing examples and sample datasets is available at Bioconductor. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ning Leng, Brian E. McIntosh, Bao Kim Nguyen, Bret Duffin, Shulan Tian, James A. Thomson, Colin N. Dewey, Ron M. Stewart, Christina Kendziorski
Bioinform.9
2014 MPBind: a Meta-motif-based statistical framework and pipeline to Predict Binding potential of SELEX-derived aptamers
abstract
UNLABELLED: Aptamers are 'synthetic antibodies' that can bind to target molecules with high affinity and specificity. Aptamers are chemically synthesized and their discovery can be performed completely in vitro, rather than relying on in vivo biological processes, making them well-suited for high-throughput discovery. However, a large fraction of the most enriched aptamers in Systematic Evolution of Ligands by EXponential enrichment (SELEX) rounds display poor binding activity. Here, we present MPBind, a Meta-motif-based statistical framework and pipeline to Predict the BIND: ing potential of SELEX-derived aptamers. Using human embryonic stem cell SELEX-Seq data, MPBind achieved high prediction accuracy for binding potential. Further analysis showed that MPBind is robust to both polymerase chain reaction amplification bias and incomplete sequencing of aptamer pools. These two biases usually confound aptamer analysis. AVAILABILITY AND IMPLEMENTATION: MPBind software and documents are available at http://www.morgridge.net/MPBind.html. The human embryonic stem cells whole-cell SELEX-Seq data are available at http://www.morgridge.net/Aptamer/.
Peng Jiang 0015, Susanne Meyer, Zhonggang Hou, Nicholas E. Propson, H. Tom Soh, James A. Thomson, Ron M. Stewart
Bioinform.7
2013 EBSeq: an empirical Bayes hierarchical model for inference in RNA-seq experiments
abstract
MOTIVATION: Messenger RNA expression is important in normal development and differentiation, as well as in manifestation of disease. RNA-seq experiments allow for the identification of differentially expressed (DE) genes and their corresponding isoforms on a genome-wide scale. However, statistical methods are required to ensure that accurate identifications are made. A number of methods exist for identifying DE genes, but far fewer are available for identifying DE isoforms. When isoform DE is of interest, investigators often apply gene-level (count-based) methods directly to estimates of isoform counts. Doing so is not recommended. In short, estimating isoform expression is relatively straightforward for some groups of isoforms, but more challenging for others. This results in estimation uncertainty that varies across isoform groups. Count-based methods were not designed to accommodate this varying uncertainty, and consequently, application of them for isoform inference results in reduced power for some classes of isoforms and increased false discoveries for others. RESULTS: Taking advantage of the merits of empirical Bayesian methods, we have developed EBSeq for identifying DE isoforms in an RNA-seq experiment comparing two or more biological conditions. Results demonstrate substantially improved power and performance of EBSeq for identifying DE isoforms. EBSeq also proves to be a robust approach for identifying DE genes. AVAILABILITY AND IMPLEMENTATION: An R package containing examples and sample datasets is available at http://www.biostat.wisc.edu/kendzior/EBSEQ/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ning Leng, John A. Dawson, James A. Thomson, Victor Ruotti, Anna I. Rissman, Bart M. G. Smits, Jill D. Haag, Michael N. Gould, Ron M. Stewart, Christina Kendziorski
Bioinform.9
2013 EBSeq: an empirical Bayes hierarchical model for inference in RNA-seq experiments
abstract
doi:10.1093/bioinformatics/btt087 Bioinformatics (2013) 29(8), 1035–1043 This article was published with the incorrect copyright information and should have appeared as below: This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/3.0/), which permits non-commercial re-use, distribution, and reproduction in any medium, provided the original work is properly cited. For commercial re-use, please contact [email protected]
Ning Leng, John A. Dawson, James A. Thomson, Victor Ruotti, Anna I. Rissman, Bart M. G. Smits, Jill D. Haag, Michael N. Gould, Ron M. Stewart, Christina Kendziorski
Bioinform.9
2013 Integrated Module and Gene-Specific Regulatory Inference Implicates Upstream Signaling Networks
abstract
Regulatory networks that control gene expression are important in diverse biological contexts including stress response and development. Each gene's regulatory program is determined by module-level regulation (e.g. co-regulation via the same signaling system), as well as gene-specific determinants that can fine-tune expression. We present a novel approach, Modular regulatory network learning with per gene information (MERLIN), that infers regulatory programs for individual genes while probabilistically constraining these programs to reveal module-level organization of regulatory networks. Using edge-, regulator- and module-based comparisons of simulated networks of known ground truth, we find MERLIN reconstructs regulatory programs of individual genes as well or better than existing approaches of network reconstruction, while additionally identifying modular organization of the regulatory networks. We use MERLIN to dissect global transcriptional behavior in two biological contexts: yeast stress response and human embryonic stem cell differentiation. Regulatory modules inferred by MERLIN capture co-regulatory relationships between signaling proteins and downstream transcription factors thereby revealing the upstream signaling systems controlling transcriptional responses. The inferred networks are enriched for regulators with genetic or physical interactions, supporting the inference, and identify modules of functionally related genes bound by the same transcriptional regulators. Our method combines the strengths of per-gene and per-module methods to reveal new insights into transcriptional regulation in stress and development.
Sushmita Roy, Stephen Lagree, Zhonggang Hou, James A. Thomson, Ron M. Stewart, Audrey P. Gasch
PLoS Comput. Biol.5
2013 Comparative RNA-seq Analysis in the Unsequenced Axolotl: The Oncogene Burst Highlights Early Gene Expression in the Blastema
abstract
The salamander has the remarkable ability to regenerate its limb after amputation. Cells at the site of amputation form a blastema and then proliferate and differentiate to regrow the limb. To better understand this process, we performed deep RNA sequencing of the blastema over a time course in the axolotl, a species whose genome has not been sequenced. Using a novel comparative approach to analyzing RNA-seq data, we characterized the transcriptional dynamics of the regenerating axolotl limb with respect to the human gene set. This approach involved de novo assembly of axolotl transcripts, RNA-seq transcript quantification without a reference genome, and transformation of abundances from axolotl contigs to human genes. We found a prominent burst in oncogene expression during the first day and blastemal/limb bud genes peaking at 7 to 14 days. In addition, we found that limb patterning genes, SALL genes, and genes involved in angiogenesis, wound healing, defense/immunity, and bone development are enriched during blastema formation and development. Finally, we identified a category of genes with no prior literature support for limb regeneration that are candidates for further evaluation based on their expression pattern during the regenerative process.
Ron M. Stewart, Cynthia Alexander Rascón, Shulan Tian, Jeff Nie, Chris Barry, Li-Fang Chu, Hamisha Ardalani, Ryan J. Wagner, Mitchell D. Probasco, Jennifer M. Bolin, Ning Leng, Srikumar Sengupta, Michael Volkmer, Bianca Habermann, Elly M. Tanaka, James A. Thomson, Colin N. Dewey
PLoS Comput. Biol.1
2010 RNA-Seq gene expression estimation with read mapping uncertainty
abstract
Abstract Motivation: RNA-Seq is a promising new technology for accurately measuring gene expression levels. Expression estimation with RNA-Seq requires the mapping of relatively short sequencing reads to a reference genome or transcript set. Because reads are generally shorter than transcripts from which they are derived, a single read may map to multiple genes and isoforms, complicating expression analyses. Previous computational methods either discard reads that map to multiple locations or allocate them to genes heuristically. Results: We present a generative statistical model and associated inference methods that handle read mapping uncertainty in a principled manner. Through simulations parameterized by real RNA-Seq data, we show that our method is more accurate than previous methods. Our improved accuracy is the result of handling read mapping uncertainty with a statistical model and the estimation of gene expression levels as the sum of isoform expression levels. Unlike previous methods, our method is capable of modeling non-uniform read distributions. Simulations with our method indicate that a read length of 20–25 bases is optimal for gene-level expression estimation from mouse and maize RNA-Seq data when sequencing throughput is fixed. Availability: An initial C++ implementation of our method that was used for the results presented in this article is available at http://deweylab.biostat.wisc.edu/rsem. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics on
Bo Li 0015, Victor Ruotti, Ron M. Stewart, James A. Thomson, Colin N. Dewey
Bioinform.3
2009 ProbeMatch: rapid alignment of oligonucleotides to genome allowing both gaps and mismatches
abstract
SUMMARY: We have developed a tool, called ProbeMatch, for matching a large set of oligonucleotide sequences against a genome database using gapped alignments. Unlike most of the existing tools such as ELAND which only perform ungapped alignments allowing at most two mismatches, ProbeMatch generates both ungapped and gapped alignments allowing up to three errors including insertion, deletion and mismatch. To speedup sequence alignment, ProbeMatch uses gapped q-grams and q-grams of various patterns to identify target hits to a query sequence. This approach results in fewer initial sequences to examine with no loss in sensitivity. ProbeMatch has been used to align 169,095 Illumina GAII reads against the human genome, which could not be mapped by ELAND, and found alignments for 28,625 reads of the 169,095 reads in less than 3 h. AVAILABILITY: Source code is freely available at (http://www.cs.wisc.edu/~jignesh/probematch/).
You Jung Kim, Nikhil Teletia, Victor Ruotti, Christopher A. Maher, Arul M. Chinnaiyan, Ron M. Stewart, James A. Thomson, Jignesh M. Patel
Bioinform.6