S. Blair Hedges

dblp:12/6411 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
1since 2021 · last 2023
0000-0002-0652-2411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
phylogenetics
0.722023
Discovering research articles containing evolutionary timetrees by machine learning · Bioinform. 2023
TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006
Data mining › text mining
text classification
0.212023
Discovering research articles containing evolutionary timetrees by machine learning · Bioinform. 2023
Data mining
text mining
0.212023
Discovering research articles containing evolutionary timetrees by machine learning · Bioinform. 2023
Bioinformatics and computational biology › phylogenetics › evolutionary history reconstruction
divergence time estimation
0.222011
TimeTree2: species divergence times on the iPhone · Bioinform. 2011
TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006
Bioinformatics and computational biology
molecular evolution
0.112011
TimeTree2: species divergence times on the iPhone · Bioinform. 2011
Bioinformatics and computational biology
knowledge base
0.012006
TimeTree: a public knowledge-base of divergence times among organisms · Bioinform. 2006

Methods — techniques the papers use, named apart from their topics

text mining · 1.3machine learning · 1.3BERT · 1.3molecular clock analysis · 0.1hierarchical data interpretation · 0.1hierarchical tree of life structure · 0.1
YearPublicationVenuePosition
2023 Discovering research articles containing evolutionary timetrees by machine learning
abstract
MOTIVATION: Timetrees depict evolutionary relationships between species and the geological times of their divergence. Hundreds of research articles containing timetrees are published in scientific journals every year. The TimeTree (TT) project has been manually locating, curating and synthesizing timetrees from these articles for almost two decades into a TimeTree of Life, delivered through a unique, user-friendly web interface (timetree.org). The manual process of finding articles containing timetrees is becoming increasingly expensive and time-consuming. So, we have explored the effectiveness of text-mining approaches and developed optimizations to find research articles containing timetrees automatically. RESULTS: We have developed an optimized machine learning system to determine if a research article contains an evolutionary timetree appropriate for inclusion in the TT resource. We found that BERT classification fine-tuned on whole-text articles achieved an F1 score of 0.67, which we increased to 0.88 by text-mining article excerpts surrounding the mentioning of figures. The new method is implemented in the TimeTreeFinder (TTF) tool, which automatically processes millions of articles to discover timetree-containing articles. We estimate that the TTF tool would produce twice as many timetree-containing articles as those discovered manually, whose inclusion in the TT database would potentially double the knowledge accessible to a wider community. Manual inspection showed that the precision on out-of-distribution recently published articles is 87%. This automation will speed up the collection and curation of timetrees with much lower human and time costs. AVAILABILITY AND IMPLEMENTATION: https://github.com/marija-stanojevic/time-tree-classification. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marija Stanojevic, Jovan Andjelkovic, Adrienne Kasprowicz, Louise A. Huuki, Jennifer Chao, S. Blair Hedges, Sudhir Kumar 0001, Zoran Obradovic
Bioinform.6
2011 TimeTree2: species divergence times on the iPhone
abstract
SUMMARY: Scientists, educators and the general public often need to know times of divergence between species. But they rarely can locate that information because it is buried in the scientific literature, usually in a format that is inaccessible to text search engines. We have developed a public knowledgebase that enables data-driven access to the collection of peer-reviewed publications in molecular evolution and phylogenetics that have reported estimates of time of divergence between species. Users can query the TimeTree resource by providing two names of organisms (common or scientific) that can correspond to species or groups of species. The current TimeTree web resource (TimeTree2) contains timetrees reported from molecular clock analyses in 910 published studies and 17 341 species that span the diversity of life. TimeTree2 interprets complex and hierarchical data from these studies for each user query, which can be launched using an iPhone application, in addition to the website. Published time estimates are now readily accessible to the scientific community, K-12 and college educators, and the general public, without requiring knowledge of evolutionary nomenclature. AVAILABILITY: TimeTree2 is accessible from the URL http://www.timetree.org, with an iPhone app available from iTunes (http://itunes.apple.com/us/app/timetree/id372842500?mt=8) and a YouTube tutorial (http://www.youtube.com/watch?v=CxmshZQciwo).
Sudhir Kumar 0001, S. Blair Hedges
Bioinform.2
2006 TimeTree: a public knowledge-base of divergence times among organisms
abstract
UNLABELLED: Biologists and other scientists routinely need to know times of divergence between species and to construct phylogenies calibrated to time (timetrees). Published studies reporting time estimates from molecular data have been increasing rapidly, but the data have been largely inaccessible to the greater community of scientists because of their complexity. TimeTree brings these data together in a consistent format and uses a hierarchical structure, corresponding to the tree of life, to maximize their utility. Results are presented and summarized, allowing users to quickly determine the range and robustness of time estimates and the degree of consensus from the published literature. AVAILABILITY: TimeTree is available at http://www.timetree.net
S. Blair Hedges, Joel Dudley, Sudhir Kumar 0001
Bioinform.1
2005 Evolutionary sequence analysis of complete eukaryote genomes
abstract
BACKGROUND: Gene duplication and gene loss during the evolution of eukaryotes have hindered attempts to estimate phylogenies and divergence times of species. Although current methods that identify clusters of orthologous genes in complete genomes have helped to investigate gene function and gene content, they have not been optimized for evolutionary sequence analyses requiring strict orthology and complete gene matrices. Here we adopt a relatively simple and fast genome comparison approach designed to assemble orthologs for evolutionary analysis. Our approach identifies single-copy genes representing only species divergences (panorthologs) in order to minimize potential errors caused by gene duplication. We apply this approach to complete sets of proteins from published eukaryote genomes specifically for phylogeny and time estimation. RESULTS: Despite the conservative criterion used, 753 panorthologs (proteins) were identified for evolutionary analysis with four genomes, resulting in a single alignment of 287,000 amino acids. With this data set, we estimate that the divergence between deuterostomes and arthropods took place in the Precambrian, approximately 400 million years before the first appearance of animals in the fossil record. Additional analyses were performed with seven, 12, and 15 eukaryote genomes resulting in similar divergence time estimates and phylogenies. CONCLUSION: Our results with available eukaryote genomes agree with previous results using conventional methods of sequence data assembly from genomes. They show that large sequence data sets can be generated relatively quickly and efficiently for evolutionary analyses of complete genomes.
Jaime E. Blair, Prachi Shah, S. Blair Hedges
BMC Bioinform.3
2003 Comparison of mode estimation methods and application in molecular clock analysis
abstract
BACKGROUND: Distributions of time estimates in molecular clock studies are sometimes skewed or contain outliers. In those cases, the mode is a better estimator of the overall time of divergence than the mean or median. However, different methods are available for estimating the mode. We compared these methods in simulations to determine their strengths and weaknesses and further assessed their performance when applied to real data sets from a molecular clock study. RESULTS: We found that the half-range mode and robust parametric mode methods have a lower bias than other mode methods under a diversity of conditions. However, the half-range mode suffers from a relatively high variance and the robust parametric mode is more susceptible to bias by outliers. We determined that bootstrapping reduces the variance of both mode estimators. Application of the different methods to real data sets yielded results that were concordant with the simulations. CONCLUSION: Because the half-range mode is a simple and fast method, and produced less bias overall in our simulations, we recommend the bootstrapped version of it as a general-purpose mode estimator and suggest a bootstrap method for obtaining the standard error and 95% confidence interval of the mode.
S. Blair Hedges, Prachi Shah
BMC Bioinform.1