Aimee R. Taylor

dblp:254/3311 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
1since 2021 · last 2026
0000-0002-2337-8992ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
statistical genetics
1.012026
Pv3Rs: Plasmodium vivax relapse, recrudescence, and reinfection statistical genetic inference · Bioinform. 2026
Bioinformatics and computational biology
drug discovery
0.412019
A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery · Bioinform. 2019
Bioinformatics and computational biology › drug discovery › computational drug discovery
machine learning for drug discovery
0.412019
A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery · Bioinform. 2019
Bioinformatics and computational biology › molecular informatics › cheminformatics
quantitative structure-activity relationship
0.412019
A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery · Bioinform. 2019

Methods — techniques the papers use, named apart from their topics

posterior probability computation · 1.0bayesian inference · 1.0support vector machine · 0.4ridge regression · 0.4rank-based loss function · 0.4random forest · 0.4quantile-activity bootstrap · 0.4neural nets · 0.4
YearPublicationVenuePosition
2026 Pv3Rs: Plasmodium vivax relapse, recrudescence, and reinfection statistical genetic inference
abstract
SUMMARY: A Plasmodium vivax recurrence can be caused by activation of dormant liver-stage hypnozoites (relapse), failure to successfully treat a blood-stage infection (recrudescence), or a new mosquito inoculation (reinfection) (Pv3Rs). Inference of the cause of recurrence is important, especially in clinical trials, where molecular correction is used to improve estimates of antimalarial treatment efficacy. Pv3Rs implements statistical inference of P. vivax recurrence states (recrudescence, relapse, reinfection) from P. vivax genetic data. Under various simplifying assumptions, we model genetic relationships between parasites within and across infections to compute the posterior probabilities of recurrence states. Pv3Rs is applicable to a variety of genetic data types, including length polymorphic loci and amplicon sequencing data. AVAILABILITY AND IMPLEMENTATION: Pv3Rs is freely available on CRAN under the GNU license. Its CRAN DOI is 10.32614/CRAN.package.Pv3Rs. Source code and installation instructions can be found at https://cran.r-project.org/web/packages/Pv3Rs/index.html.
Yong See Foo, Aimee R. Taylor
Bioinform.3
2019 A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
abstract
MOTIVATION: Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs. RESULTS: The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and 'memorize' the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand. AVAILABILITY AND IMPLEMENTATION: All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Oliver P. Watson, Isidro Cortes-Ciriano, Aimee R. Taylor, James A. Watson
Bioinform.3