Nathan C. Sheffield

dblp:176/9924 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-5643-4068ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Automated biomedical hypothesis generation with time-aware hypergraph contrastive learning
abstract
Abstract Research in scientific domains now generates more than a million articles annually, overwhelming researchers and hindering discovery. This surge has sparked interest in biomedical hypothesis generation (HG), which aims to uncover implicit patterns among biomedical concepts. Most existing methods focus on pairwise link prediction, overlooking the complex, multi-concept relationships underlying many breakthroughs. We introduce HyHG , a temporal Hy pergraph contrastive learning framework for biomedical H ypothesis G eneration, which redefines hypotheses as hyperedges—sets of co-mentioned concepts in an article. By representing articles as hyperedges and organizing them into a temporal hypergraph, HyHG captures the evolution of scientific ideas over time. A transformer-based architecture learns from historical hyperedge sequences to predict future hyperedges—sets of concepts likely to co-occur in the future literature. To distinguish genuine hypotheses from misleading ones, HyHG employs a time-anchored contrastive loss and hard negative sampling based on minimal edits to real hyperedges. We demonstrate that HyHG achieves state-of-the-art performance on three biomedical datasets. Our code and data are available at: https://github.com/amir-hassan25/Temporal-Hypergraph-Contrastive-Learning.
Amir Hassan Shariatmadari, Sikun Guo, Nathan C. Sheffield, Aidong Zhang 0001, Kishlay Jha
Knowl. Inf. Syst.3
2025 HyHG: A Temporal Hypergraph Contrastive Learning Framework for Biomedical Hypothesis Generation
abstract
Biomedical research now generates more than a million articles annually, overwhelming researchers and hindering discovery. This surge has sparked interest in biomedical hypothesis generation (HG), which aims to uncover implicit patterns among biomedical concepts. Most existing methods focus on pairwise link prediction, overlooking the complex, multi-concept relationships underlying many breakthroughs. We introduce HyHG, a temporal Hypergraph contrastive learning framework for biomedical Hypothesis Generation, which redefines hypotheses as hyperedges–sets of co-mentioned concepts in an article. By representing articles as hyperedges and organizing them into a temporal hypergraph, HyHG captures the evolution of scientific ideas over time. A transformer-based architecture learns from historical hyperedge sequences to predict future hyperedges–sets of concepts likely to co-occur in future literature. To distinguish genuine hypotheses from misleading ones, HyHG employs a timeanchored contrastive loss and hard negative sampling based on minimal edits to real hyperedges. We demonstrate state-of-the-art performance on three biomedical datasets. Our code and data are available at: https://github.com/amirhassan25/Temporal-Hypergraph-Contrastive-Learning.
Amir Hassan Shariatmadari, Sikun Guo, Nathan C. Sheffield, Aidong Zhang 0001, Kishlay Jha
ICDM3
2025 ConceptDrift: leveraging spatial, temporal and semantic evolution of biomedical concepts for hypothesis generation
abstract
MOTIVATION: Hypothesis generation is a fundamental problem in biomedical text mining that aims to generate ideas that are new, interesting, and plausible by discovering unexplored links between biomedical concepts. Despite significant advances made by existing approaches, they do not fully leverage the evolutionary properties of biomedical concepts. This is limiting because scientific knowledge continually evolves over time, with new facts being added and old ones becoming obsolete. Thus, it is crucial to capture the evolutionary properties of biomedical concepts from multiple perspectives (e.g. spatial, temporal, and semantic) to generate hypotheses that reflect the up-to-date information landscape of the biomedical domain. RESULTS: We introduce a novel framework, ConceptDrift, that models the hypothesis generation task as a sequence of temporal graphlets and simultaneously encodes spatial, temporal, and semantic change. Unlike existing approaches that treat these dimensions independently, ConceptDrift is the first to provide a holistic understanding of concept evolution by integrating them into a unified framework. Grounded in the theories of the Distributional Hypothesis and Conceptual Change, our method adapts these principles to the unique challenges of large-scale biomedical literature. We conduct extensive experiments across multiple datasets and demonstrate that ConceptDrift consistently outperforms state-of-the-art baselines in generating accurate and meaningful hypotheses. Our framework shows immediate practical benefits for web-based literature mining tools in life sciences and biomedicine, offering more robust and predictive feature representations. AVAILABILITY AND IMPLEMENTATION: https://github.com/amir-hassan25/ConceptDrift (DOI: 10.6084/m9.figshare.29975476).
Amir Hassan Shariatmadari, Alireza Jafari, Sikun Guo, Sneha Srinivasan, Nathan C. Sheffield, Aidong Zhang 0001, Kishlay Jha
Bioinform.5
2024 DeepGSEA: explainable deep gene set enrichment analysis for single-cell transcriptomic data
abstract
MOTIVATION: Gene set enrichment (GSE) analysis allows for an interpretation of gene expression through pre-defined gene set databases and is a critical step in understanding different phenotypes. With the rapid development of single-cell RNA sequencing (scRNA-seq) technology, GSE analysis can be performed on fine-grained gene expression data to gain a nuanced understanding of phenotypes of interest. However, with the cellular heterogeneity in single-cell gene profiles, current statistical GSE analysis methods sometimes fail to identify enriched gene sets. Meanwhile, deep learning has gained traction in applications like clustering and trajectory inference in single-cell studies due to its prowess in capturing complex data patterns. However, its use in GSE analysis remains limited, due to interpretability challenges. RESULTS: In this paper, we present DeepGSEA, an explainable deep gene set enrichment analysis approach which leverages the expressiveness of interpretable, prototype-based neural networks to provide an in-depth analysis of GSE. DeepGSEA learns the ability to capture GSE information through our designed classification tasks, and significance tests can be performed on each gene set, enabling the identification of enriched sets. The underlying distribution of a gene set learned by DeepGSEA can be explicitly visualized using the encoded cell and cellular prototype embeddings. We demonstrate the performance of DeepGSEA over commonly used GSE analysis methods by examining their sensitivity and specificity with four simulation studies. In addition, we test our model on three real scRNA-seq datasets and illustrate the interpretability of DeepGSEA by showing how its results can be explained. AVAILABILITY AND IMPLEMENTATION: https://github.com/Teddy-XiongGZ/DeepGSEA.
Guangzhi Xiong, Nathan Leroy, Stefan Bekiranov, Nathan C. Sheffield, Aidong Zhang 0001
Bioinform.4
2023 GEOfetch: a command-line tool for downloading data and standardized metadata from GEO and SRA
abstract
MOTIVATION: The Gene Expression Omnibus has become an important source of biological data for secondary analysis. However, there is no simple, programmatic way to download data and metadata from Gene Expression Omnibus (GEO) in a standardized annotation format. RESULTS: To address this, we present GEOfetch-a command-line tool that downloads and organizes data and metadata from GEO and SRA. GEOfetch formats the downloaded metadata as a Portable Encapsulated Project, providing universal format for the reanalysis of public data. AVAILABILITY AND IMPLEMENTATION: GEOfetch is available on Bioconda and the Python Package Index (PyPI).
Oleksandr Khoroshevskyi, Nathan Leroy, Vincent P. Reuter, Nathan C. Sheffield
Bioinform.4
2023 excluderanges: exclusion sets for T2T-CHM13, GRCm39, and other genome assemblies
abstract
SUMMARY: Exclusion regions are sections of reference genomes with abnormal pileups of short sequencing reads. Removing reads overlapping them improves biological signal, and these benefits are most pronounced in differential analysis settings. Several labs created exclusion region sets, available primarily through ENCODE and Github. However, the variety of exclusion sets creates uncertainty which sets to use. Furthermore, gap regions (e.g. centromeres, telomeres, short arms) create additional considerations in generating exclusion sets. We generated exclusion sets for the latest human T2T-CHM13 and mouse GRCm39 genomes and systematically assembled and annotated these and other sets in the excluderanges R/Bioconductor data package, also accessible via the BEDbase.org API. The package provides unified access to 82 GenomicRanges objects covering six organisms, multiple genome assemblies, and types of exclusion regions. For human hg38 genome assembly, we recommend hg38.Kundaje.GRCh38_unified_blacklist as the most well-curated and annotated, and sets generated by the Blacklist tool for other organisms. AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/excluderanges/. Package website: https://dozmorovlab.github.io/excluderanges/.
Jonathan D. Ogata, Wancen Mu, Eric S. Davis, Bingjie Xue, J. Chuck Harrell, Nathan C. Sheffield, Douglas H. Phanstiel, Michael I. Love, Mikhail G. Dozmorov
Bioinform.6
2021 IGD: high-performance search for large-scale genomic interval datasets
abstract
SUMMARY: Databases of large-scale genome projects now contain thousands of genomic interval datasets. These data are a critical resource for understanding the function of DNA. However, our ability to examine and integrate interval data of this scale is limited. Here, we introduce the integrated genome database (IGD), a method and tool for searching genome interval datasets more than three orders of magnitude faster than existing approaches, while using only one hundredth of the memory. IGD uses a novel linear binning method that allows us to scale analysis to billions of genomic regions. AVAILABILITYAND IMPLEMENTATION: https://github.com/databio/IGD. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianglin Feng, Nathan C. Sheffield
Bioinform.2
2021 Embeddings of genomic region sets capture rich biological associations in lower dimensions
abstract
MOTIVATION: Genomic region sets summarize functional genomics data and define locations of interest in the genome such as regulatory regions or transcription factor binding sites. The number of publicly available region sets has increased dramatically, leading to challenges in data analysis. RESULTS: We propose a new method to represent genomic region sets as vectors, or embeddings, using an adapted word2vec approach. We compared our approach to two simpler methods based on interval unions or term frequency-inverse document frequency and evaluated the methods in three ways: First, by classifying the cell line, antibody or tissue type of the region set; second, by assessing whether similarity among embeddings can reflect simulated random perturbations of genomic regions; and third, by testing robustness of the proposed representations to different signal thresholds for calling peaks. Our word2vec-based region set embeddings reduce dimensionality from more than a hundred thousand to 100 without significant loss in classification performance. The vector representation could identify cell line, antibody and tissue type with over 90% accuracy. We also found that the vectors could quantitatively summarize simulated random perturbations to region sets and are more robust to subsampling the data derived from different peak calling thresholds. Our evaluations demonstrate that the vectors retain useful biological information in relatively lower-dimensional spaces. We propose that vector representation of region sets is a promising approach for efficient analysis of genomic region data. AVAILABILITY AND IMPLEMENTATION: https://github.com/databio/regionset-embedding. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Erfaneh Gharavi, Aaron Gu, Guangtao Zheng, Jason P. Smith, Hyun Jae Cho, Aidong Zhang 0001, Donald E. Brown, Nathan C. Sheffield
Bioinform.8
2021 Refget: standardized access to reference sequences
abstract
MOTIVATION: Reference sequences are essential in creating a baseline of knowledge for many common bioinformatics methods, especially those using genomic sequencing. RESULTS: We have created refget, a Global Alliance for Genomics and Health API specification to access reference sequences and sub-sequences using an identifier derived from the sequence itself. We present four reference implementations across in-house and cloud infrastructure, a compliance suite and a web report used to ensure specification conformity across implementations. AVAILABILITY AND IMPLEMENTATION: The refget specification can be found at: https://w3id.org/ga4gh/refget. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andy Yates, Jeremy Adams, Somesh Chaturvedi, Robert Davies, Matthew R. Laird, Rasko Leinonen, Rishi Nag, Nathan C. Sheffield, Oliver Hofmann 0001, Thomas M. Keane
Bioinform.8
2019 Augmented Interval List: a novel data structure for efficient genomic interval search
abstract
MOTIVATION: Genomic data is frequently stored as segments or intervals. Because this data type is so common, interval-based comparisons are fundamental to genomic analysis. As the volume of available genomic data grows, developing efficient and scalable methods for searching interval data is necessary. RESULTS: We present a new data structure, the Augmented Interval List (AIList), to enumerate intersections between a query interval q and an interval set R. An AIList is constructed by first sorting R as a list by the interval start coordinate, then decomposing it into a few approximately flattened components (sublists), and then augmenting each sublist with the running maximum interval end. The query time for AIList is O(log2N+n+m), where n is the number of overlaps between R and q, N is the number of intervals in the set R and m is the average number of extra comparisons required to find the n overlaps. Tested on real genomic interval datasets, AIList code runs 5-18 times faster than standard high-performance code based on augmented interval-trees, nested containment lists or R-trees (BEDTools). For large datasets, the memory-usage for AIList is 4-60% of other methods. The AIList data structure, therefore, provides a significantly improved fundamental operation for highly scalable genomic data analysis. AVAILABILITY AND IMPLEMENTATION: An implementation of the AIList data structure with both construction and search algorithms is available at http://ailist.databio.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jianglin Feng, Aakrosh Ratan, Nathan C. Sheffield
Bioinform.3
2018 MIRA: an R package for DNA methylation-based inference of regulatory activity
abstract
Summary: DNA methylation contains information about the regulatory state of the cell. MIRA aggregates genome-scale DNA methylation data into a DNA methylation profile for a given region set with shared biological annotation. Using this profile, MIRA infers and scores the collective regulatory activity for the region set. MIRA facilitates regulatory analysis in situations where classical regulatory assays would be difficult and allows public sources of region sets to be leveraged for novel insight into the regulatory state of DNA methylation datasets. Availability and implementation: http://bioconductor.org/packages/MIRA.
John T. Lawson, Eleni M. Tomazou, Christoph Bock, Nathan C. Sheffield
Bioinform.4
2018 BART: a transcription factor prediction tool with query gene sets or epigenomic profiles
abstract
Summary: Identification of functional transcription factors that regulate a given gene set is an important problem in gene regulation studies. Conventional approaches for identifying transcription factors, such as DNA sequence motif analysis, are unable to predict functional binding of specific factors and not sensitive enough to detect factors binding at distal enhancers. Here, we present binding analysis for regulation of transcription (BART), a novel computational method and software package for predicting functional transcription factors that regulate a query gene set or associate with a query genomic profile, based on more than 6000 existing ChIP-seq datasets for over 400 factors in human or mouse. This method demonstrates the advantage of utilizing publicly available data for functional genomics research. Availability and implementation: BART is implemented in Python and available at http://faculty.virginia.edu/zanglab/bart. Supplementary information: Supplementary data are available at Bioinformatics online.
Zhenjia Wang, Mete Civelek, Clint L. Miller, Nathan C. Sheffield, Michael J. Guertin, Chongzhi Zang
Bioinform.4
2016 LOLA: enrichment analysis for genomic region sets and regulatory elements in R and Bioconductor
abstract
UNLABELLED: Genomic datasets are often interpreted in the context of large-scale reference databases. One approach is to identify significantly overlapping gene sets, which works well for gene-centric data. However, many types of high-throughput data are based on genomic regions. Locus Overlap Analysis (LOLA) provides easy and automatable enrichment analysis for genomic region sets, thus facilitating the interpretation of functional genomics and epigenomics data. AVAILABILITY AND IMPLEMENTATION: R package available in Bioconductor and on the following website: http://lola.computational-epigenetics.org.
Nathan C. Sheffield, Christoph Bock
Bioinform.1