Simon Rogers

dblp:69/2554 · DBLP profile ↗
← Back
40ranked-venue papers
11as first author
7since 2021 · last 2023
0000-0003-3578-4477ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 1 first-authorHuman-computer interaction and ubiquitous computing · 8 · 2 first-authorSystems, architecture and hardware · 3Databases, data management, data science and information retrieval · 2Computer networks · 1
YearPublicationVenuePosition
2023 TopNEXt: automatic DDA exclusion framework for multi-sample mass spectrometry experiments
abstract
MOTIVATION: Liquid Chromatography Tandem Mass Spectrometry experiments aim to produce high-quality fragmentation spectra, which can be used to annotate metabolites. However, current Data-Dependent Acquisition approaches may fail to collect spectra of sufficient quality and quantity for experimental outcomes, and extend poorly across multiple samples by failing to share information across samples or by requiring manual expert input. RESULTS: We present TopNEXt, a real-time scan prioritization framework that improves data acquisition in multi-sample Liquid Chromatography Tandem Mass Spectrometry metabolomics experiments. TopNEXt extends traditional Data-Dependent Acquisition exclusion methods across multiple samples by using a Region of Interest and intensity-based scoring system. Through both simulated and lab experiments, we show that methods incorporating these novel concepts acquire fragmentation spectra for an additional 10% of our set of target peaks and with an additional 20% of acquisition intensity. By increasing the quality and quantity of fragmentation spectra, TopNEXt can help improve metabolite identification with a potential impact across a variety of experimental contexts. AVAILABILITY AND IMPLEMENTATION: TopNEXt is implemented as part of the ViMMS framework and the latest version can be found at https://github.com/glasgowcompbio/vimms. A stable version used to produce our results can be found at 10.5281/zenodo.7468914.
Ross McBride, Joe Wandy, Stefan Weidt, Simon Rogers, Vinny Davies, Rónán Daly, Kevin Bryson 0001
Bioinform.4
2022 Using topic modeling to detect cellular crosstalk in scRNA-seq
abstract
Cell-cell interactions are vital for numerous biological processes including development, differentiation, and response to inflammation. Currently, most methods for studying interactions on scRNA-seq level are based on curated databases of ligands and receptors. While those methods are useful, they are limited to our current biological knowledge. Recent advances in single cell protocols have allowed for physically interacting cells to be captured, and as such we have the potential to study interactions in a complemantary way without relying on prior knowledge. We introduce a new method based on Latent Dirichlet Allocation (LDA) for detecting genes that change as a result of interaction. We apply our method to synthetic datasets to demonstrate its ability to detect genes that change in an interacting population compared to a reference population. Next, we apply our approach to two datasets of physically interacting cells to identify the genes that change as a result of interaction, examples include adhesion and co-stimulatory molecules which confirm physical interaction between cells. For each dataset we produce a ranking of genes that are changing in subpopulations of the interacting cells. In addition to the genes discussed in the original publications, we highlight further candidates for interaction in the top 100 and 300 ranked genes. Lastly, we apply our method to a dataset generated by a standard droplet-based protocol not designed to capture interacting cells, and discuss its suitability for analysing interactions. We present a method that streamlines detection of interactions and does not require prior clustering and generation of synthetic reference profiles to detect changes in expression.
Alexandrina Pancheva, Helen Wheadon, Simon Rogers, Thomas D. Otto
PLoS Comput. Biol.3
2021 Probabilistic framework for integration of mass spectrum and retention time information in small molecule identification
abstract
MOTIVATION: Identification of small molecules in a biological sample remains a major bottleneck in molecular biology, despite a decade of rapid development of computational approaches for predicting molecular structures using mass spectrometry (MS) data. Recently, there has been increasing interest in utilizing other information sources, such as liquid chromatography (LC) retention time (RT), to improve identifications solely based on MS information, such as precursor mass-per-charge and tandem mass spectrometry (MS2). RESULTS: We put forward a probabilistic modelling framework to integrate MS and RT data of multiple features in an LC-MS experiment. We model the MS measurements and all pairwise retention order information as a Markov random field and use efficient approximate inference for scoring and ranking potential molecular structures. Our experiments show improved identification accuracy by combining MS2 data and retention orders using our approach, thereby outperforming state-of-the-art methods. Furthermore, we demonstrate the benefit of our model when only a subset of LC-MS features has MS2 measurements available besides MS1. AVAILABILITY AND IMPLEMENTATION: Software and data are freely available at https://github.com/aalto-ics-kepaco/msms_rt_score_integration. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Eric Bach 0002, Simon Rogers, John Williamson 0001, Juho Rousu
Bioinform.2
2021 Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions
abstract
Specialised metabolites from microbial sources are well-known for their wide range of biomedical applications, particularly as antibiotics. When mining paired genomic and metabolomic data sets for novel specialised metabolites, establishing links between Biosynthetic Gene Clusters (BGCs) and metabolites represents a promising way of finding such novel chemistry. However, due to the lack of detailed biosynthetic knowledge for the majority of predicted BGCs, and the large number of possible combinations, this is not a simple task. This problem is becoming ever more pressing with the increased availability of paired omics data sets. Current tools are not effective at identifying valid links automatically, and manual verification is a considerable bottleneck in natural product research. We demonstrate that using multiple link-scoring functions together makes it easier to prioritise true links relative to others. Based on standardising a commonly used score, we introduce a new, more effective score, and introduce a novel score using an Input-Output Kernel Regression approach. Finally, we present NPLinker, a software framework to link genomic and metabolomic data. Results are verified using publicly available data sets that include validated links.
Grímur Hjörleifsson Eldjárn, Andrew Ramsay, Justin J. J. van der Hooft, Katherine R. Duncan, Sylvia Soldatou, Juho Rousu, Rónán Daly, Joe Wandy, Simon Rogers
PLoS Comput. Biol.9
2021 Spec2Vec: Improved mass spectral similarity scoring through learning of structural relationships
abstract
Spectral similarity is used as a proxy for structural similarity in many tandem mass spectrometry (MS/MS) based metabolomics analyses such as library matching and molecular networking. Although weaknesses in the relationship between spectral similarity scores and the true structural similarities have been described, little development of alternative scores has been undertaken. Here, we introduce Spec2Vec, a novel spectral similarity score inspired by a natural language processing algorithm-Word2Vec. Spec2Vec learns fragmental relationships within a large set of spectral data to derive abstract spectral embeddings that can be used to assess spectral similarities. Using data derived from GNPS MS/MS libraries including spectra for nearly 13,000 unique molecules, we show how Spec2Vec scores correlate better with structural similarity than cosine-based scores. We demonstrate the advantages of Spec2Vec in library matching and molecular networking. Spec2Vec is computationally more scalable allowing structural analogue searches in large databases within seconds.
Florian Huber 0001, Lars Ridder, Stefan Verhoeven, Jurriaan H. Spaaks, Faruk Diblen, Simon Rogers, Justin J. J. van der Hooft
PLoS Comput. Biol.6
2021 PIEZO1 and the mechanism of the long circulatory longevity of human red blood cells
abstract
Human red blood cells (RBCs) have a circulatory lifespan of about four months. Under constant oxidative and mechanical stress, but devoid of organelles and deprived of biosynthetic capacity for protein renewal, RBCs undergo substantial homeostatic changes, progressive densification followed by late density reversal among others, changes assumed to have been harnessed by evolution to sustain the rheological competence of the RBCs for as long as possible. The unknown mechanisms by which this is achieved are the subject of this investigation. Each RBC traverses capillaries between 1000 and 2000 times per day, roughly one transit per minute. A dedicated Lifespan model of RBC homeostasis was developed as an extension of the RCM introduced in the previous paper to explore the cumulative patterns predicted for repetitive capillary transits over a standardized lifespan period of 120 days, using experimental data to constrain the range of acceptable model outcomes. Capillary transits were simulated by periods of elevated cell/medium volume ratios and by transient deformation-induced permeability changes attributed to PIEZO1 channel mediation as outlined in the previous paper. The first unexpected finding was that quantal density changes generated during single capillary transits cease accumulating after a few days and cannot account for the observed progressive densification of RBCs on their own, thus ruling out the quantal hypothesis. The second unexpected finding was that the documented patterns of RBC densification and late reversal could only be emulated by the implementation of a strict time-course of decay in the activities of the calcium and Na/K pumps, suggestive of a selective mechanism enabling the extended longevity of RBCs. The densification pattern over most of the circulatory lifespan was determined by calcium pump decay whereas late density reversal was shaped by the pattern of Na/K pump decay. A third finding was that both quantal changes and pump-decay regimes were necessary to account for the documented lifespan pattern, neither sufficient on their own. A fourth new finding revealed that RBCs exposed to levels of PIEZO1-medited calcium permeation above certain thresholds in the circulation could develop a pattern of early or late hyperdense collapse followed by delayed density reversal. When tested over much reduced lifespan periods the results reproduced the known circulatory fate of irreversible sickle cells, the cell subpopulation responsible for vaso-occlusion and for most of the clinical manifestations of sickle cell disease. Analysis of the results provided an insightful new understanding of the mechanisms driving the changes in RBC homeostasis during circulatory aging in health and disease.
Simon Rogers, Virgilio L. Lew
PLoS Comput. Biol.1
2021 Up-down biphasic volume response of human red blood cells to PIEZO1 activation during capillary transits
abstract
In this paper we apply a novel JAVA version of a model on the homeostasis of human red blood cells (RBCs) to investigate the changes RBCs experience during single capillary transits. In the companion paper we apply a model extension to investigate the changes in RBC homeostasis over the approximately 200000 capillary transits during the ~120 days lifespan of the cells. These are topics inaccessible to direct experimentation but rendered mature for a computational modelling approach by the large body of recent and early experimental results which robustly constrain the range of parameter values and model outcomes, offering a unique opportunity for an in depth study of the mechanisms involved. Capillary transit times vary between 0.5 and 1.5s during which the red blood cells squeeze and deform in the capillary stream transiently opening stress-gated PIEZO1 channels allowing ion gradient dissipation and creating minuscule quantal changes in RBC ion contents and volume. Widely accepted views, based on the effects of experimental shear stress on human RBCs, suggested that quantal changes generated during capillary transits add up over time to develop the documented changes in RBC density and composition during their long circulatory lifespan, the quantal hypothesis. Applying the new red cell model (RCM) we investigated here the changes in homeostatic variables that may be expected during single capillary transits resulting from transient PIEZO1 channel activation. The predicted quantal volume changes were infinitesimal in magnitude, biphasic in nature, and essentially irreversible within inter-transit periods. A sub-second transient PIEZO1 activation triggered a sharp swelling peak followed by a much slower recovery period towards lower-than-baseline volumes. The peak response was caused by net CaCl2 and fluid gain via PIEZO1 channels driven by the steep electrochemical inward Ca2+ gradient. The ensuing dehydration followed a complex time-course with sequential, but partially overlapping contributions by KCl loss via Ca2+-activated Gardos channels, restorative Ca2+ extrusion by the plasma membrane calcium pump, and chloride efflux by the Jacobs-Steward mechanism. The change in relative cell volume predicted for single capillary transits was around 10-5, an infinitesimal volume change incompatible with a functional role in capillary flow. The biphasic response predicted by the RCM appears to conform to the quantal hypothesis, but whether its cumulative effects could account for the documented changes in density during RBC senescence required an investigation of the effects of myriad transits over the full four months circulatory lifespan of the cells, the subject of the next paper.
Simon Rogers, Virgilio L. Lew
PLoS Comput. Biol.1
2020 Predicting host taxonomic information from viral genomes: A comparison of feature representations
abstract
The rise in metagenomics has led to an exponential growth in virus discovery. However, the majority of these new virus sequences have no assigned host. Current machine learning approaches to predicting virus host interactions have a tendency to focus on nucleotide features, ignoring other representations of genomic information. Here we investigate the predictive potential of features generated from four different 'levels' of viral genome representation: nucleotide, amino acid, amino acid properties and protein domains. This more fully exploits the biological information present in the virus genomes. Over a hundred and eighty binary datasets for infecting versus non-infecting viruses at all taxonomic ranks of both eukaryote and prokaryote hosts were compiled. The viral genomes were converted into the four different levels of genome representation and twenty feature sets were generated by extracting k-mer compositions and predicted protein domains. We trained and tested Support Vector Machine, SVM, classifiers to compare the predictive capacity of each of these feature sets for each dataset. Our results show that all levels of genome representation are consistently predictive of host taxonomy and that prediction k-mer composition improves with increasing k-mer length for all k-mer based features. Using a phylogenetically aware holdout method, we demonstrate that the predictive feature sets contain signals reflecting both the evolutionary relationship between the viruses infecting related hosts, and host-mimicry. Our results demonstrate that incorporating a range of complementary features, generated purely from virus genome sequences, leads to improved accuracy for a range of virus host prediction tasks enabling computational assignment of host taxonomic information.
Francesca Young, Simon Rogers, David L. Robertson
PLoS Comput. Biol.2
2020 Per-Host DDoS Mitigation by Direct-Control Reinforcement Learning
abstract
DDoS attacks plague the availability of online services today, yet like many cybersecurity problems are evolving and non-stationary. Normal and attack patterns shift as new protocols and applications are introduced, further compounded by burstiness and seasonal variation. Accordingly, it is difficult to apply machine learning-based techniques and defences in practice. Reinforcement learning (RL) may overcome this detection problem for DDoS attacks by managing and monitoring consequences; an agent's role is to learn to optimise performance criteria (which are always available) in an online manner. We advance the state-of-the-art in RL-based DDoS mitigation by introducing two agent classes designed to act on a per-flow basis, in a protocol-agnostic manner for any network topology. This is supported by an in-depth investigation of feature suitability and empirical evaluation. Our results show the existence of flow features with high predictive power for different traffic classes, when used as a basis for feedback-loop-like control. We show that the new RL agent models can offer a significant increase in goodput of legitimate TCP traffic for many choices of host density.
Kyle A. Simpson, Simon Rogers, Dimitrios P. Pezaros
IEEE Trans. Netw. Serv. Manag.2
2018 ShinyKGode: an interactive application for ODE parameter inference using gradient matching
abstract
Motivation: Mathematical modelling based on ordinary differential equations (ODEs) is widely used to describe the dynamics of biological systems, particularly in systems and pathway biology. Often the kinetic parameters of these ODE systems are unknown and have to be inferred from the data. Approximate parameter inference methods based on gradient matching (which do not require performing computationally expensive numerical integration of the ODEs) have been getting popular in recent years, but many implementations are difficult to run without expert knowledge. Here, we introduce ShinyKGode, an interactive web application to perform fast parameter inference on ODEs using gradient matching. Results: ShinyKGode can be used to infer ODE parameters on simulated and observed data using gradient matching. Users can easily load their own models in Systems Biology Markup Language format, and a set of pre-defined ODE benchmark models are provided in the application. Inferred parameters are visualized alongside diagnostic plots to assess convergence. Availability and implementation: The R package for ShinyKGode can be installed through the Comprehensive R Archive Network (CRAN). Installation instructions, as well as tutorial videos and source code are available at https://joewandy.github.io/shinyKGode. Supplementary information: Supplementary data are available at Bioinformatics online.
Joe Wandy, Mu Niu, Diana Giurghita, Rónán Daly, Simon Rogers, Dirk Husmeier
Bioinform.5
2018 Ms2lda.org: web-based topic modelling for substructure discovery in mass spectrometry
abstract
MOTIVATION: We recently published MS2LDA, a method for the decomposition of sets of molecular fragment data derived from large metabolomics experiments. To make the method more widely available to the community, here we present ms2lda.org, a web application that allows users to upload their data, run MS2LDA analyses and explore the results through interactive visualizations. RESULTS: Ms2lda.org takes tandem mass spectrometry data in many standard formats and allows the user to infer the sets of fragment and neutral loss features that co-occur together (Mass2Motifs). As an alternative workflow, the user can also decompose a data set onto predefined Mass2Motifs. This is accomplished through the web interface or programmatically from our web service. AVAILABILITY AND IMPLEMENTATION: The website can be found at http://ms2lda.org, while the source code is available at https://github.com/sdrogers/ms2ldaviz under the MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Joe Wandy, Yunfeng Zhu, Justin J. J. van der Hooft, Rónán Daly, Michael P. Barrett, Simon Rogers
Bioinform.6
2017 Single-shot clothing category recognition in free-configurations with application to autonomous clothes sorting
abstract
This paper proposes a single-shot approach for recognising clothing categories from 2.5D features. We propose two visual features, BSP (B-Spline Patch) and TSD (Topology Spatial Distances) for this task. The local BSP features are encoded by LLC (Locality-constrained Linear Coding) and fused with three different global features. Our visual feature is robust to deformable shapes and our approach is able to recognise the category of unknown clothing in unconstrained and random configurations. We integrated the category recognition pipeline with a stereo vision system, clothing instance detection, and dual-arm manipulators to achieve an autonomous sorting system. To verify the performance of our proposed method, we build a high-resolution RGBD clothing dataset of 50 clothing items of 5 categories sampled in random configurations (a total of 2,100 clothing samples). Experimental results show that our approach is able to reach 83.2% accuracy while classifying clothing items which were previously unseen during training. This advances beyond the previous state-of-the-art by 36.2%. Finally, we evaluate the proposed approach in an autonomous robot sorting system, in which the robot recognises a clothing item from an unconstrained pile, grasps it, and sorts it into a box according to its category. Our proposed sorting system achieves reasonable sorting success rates with single-shot perception.
Li Sun 0005, Gerardo Aragon-Camarasa, Simon Rogers, Rustam Stolkin, J. Paul Siebert
IROS3
2017 PiMP my metabolome: an integrated, web-based tool for LC-MS metabolomics data
abstract
SUMMARY: The Polyomics integrated Metabolomics Pipeline (PiMP) fulfils an unmet need in metabolomics data analysis. PiMP offers automated and user-friendly analysis from mass spectrometry data acquisition to biological interpretation. Our key innovations are the Summary Page, which provides a simple overview of the experiment in the format of a scientific paper, containing the key findings of the experiment along with associated metadata; and the Metabolite Page, which provides a list of each metabolite accompanied by 'evidence cards', which provide a variety of criteria behind metabolite annotation including peak shapes, intensities in different sample groups and database information. AVAILABILITY AND IMPLEMENTATION: PiMP is available at http://polyomics.mvls.gla.ac.uk, and access is freely available on request. 50 GB of space is allocated for data storage, with unrestricted number of samples and analyses per user. Source code is available at https://github.com/RonanDaly/pimp and licensed under the GPL. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yoann Gloaguen, Fraser R. Morton, Rónán Daly, Ross Gurden, Simon Rogers, Joe Wandy, Michael P. Barrett, Karl E. V. Burgess
Bioinform.5
2016 Detecting Swipe Errors on Touchscreens using Grip Modulation
abstract
We show that when users make errors on mobile devices they make immediate and distinct physical responses that can be observed with standard sensors. We used three standard cognitive tasks (Flanker, Stroop and SART) to induce errors from 20 participants. Using simple low-resolution capacitive touch sensors placed around a standard mobile device and the built-in accelerometer, we demonstrate that errors can be predicted at low error rates from micro-adjustments to hand grip and movement in the period shortly after swiping the touchscreen. Specifically, when combining features derived from hand grip and movement we obtain a mean AUC of 0.96 (with false accept and reject rates both below 10%). Our results demonstrate that hand grip and movement provide strong and low latency evidence for mistakes. The ability to detect user errors in this way could be a valuable component in future interaction systems, allowing interfaces to make it easier for users to correct erroneous inputs.
Mohammad Faizuddin Mohd Noor, Simon Rogers, John Williamson 0001
CHI2
2016 Fast Parameter Inference in Nonlinear Dynamical Systems using Iterative Gradient Matching
abstract
Parameter inference in mechanistic models of coupled differential equations is a topical and challenging problem. We propose a new method based on kernel ridge regression and gradient matching, and an objective function that simultaneously encourages goodness of fit and penalises inconsistencies with the differential equations. Fast minimisation is achieved by exploiting partial convexity inherent in this function, and setting up an iterative algorithm in the vein of the EM algorithm. An evaluation of the proposed method on various benchmark data suggests that it compares favourably with state-of-the-art alternatives.
Mu Niu, Simon Rogers, Maurizio Filippone, Dirk Husmeier
ICML2
2016 Recognising the clothing categories from free-configuration using Gaussian-Process-based interactive perception
abstract
In this paper, we propose a Gaussian Process-based interactive perception approach for recognising highly-wrinkled clothes. We have integrated this recognition method within a clothes sorting pipeline for the pre-washing stage of an autonomous laundering process. Our approach differs from reported clothing manipulation approaches by allowing the robot to update its perception confidence via numerous interactions with the garments. The classifiers predominantly reported in clothing perception (e.g. SVM, Random Forest) studies do not provide true classification probabilities, due to their inherent structure. In contrast, probabilistic classifiers (of which the Gaussian Process is a popular example) are able to provide predictive probabilities. In our approach, we employ a multi-class Gaussian Process classification using the Laplace approximation for posterior inference and optimising hyper-parameters via marginal likelihood maximisation. Our experimental results show that our approach is able to recognise unknown garments from highly-occluded and wrinkled configurations and demonstrates a substantial improvement over non-interactive perception approaches.
Li Sun 0005, Simon Rogers, Gerardo Aragon-Camarasa, J. Paul Siebert
ICRA2
2015 Accurate garment surface analysis using an active stereo robot head with application to dual-arm flattening
abstract
We present a visually guided, dual-arm, industrial robot system that is capable of autonomously flattening garments by means of a novel visual perception pipeline that fully interprets high-quality RGB-D images of a clothing scene based on an active stereo robot head. A segmented clothing range map is B-Spline smoothed prior to being parsed by means of shape and topology analysis into ‘wrinkle’ structures. The length, width and height of each wrinkle is used to quantify the topology of each wrinkle and thereby rank wrinkles by size such that a greedy algorithm can identify the largest wrinkle present. A flattening plan optimised for the largest detected wrinkle is formulated based on dual-arm manipulation. We report the validation of our autonomous flattening behaviour and observe that dual-arm flattening requires significantly fewer manipulation iterations than single-arm flattening. Our experimental results also reveal that the flattening process is heavily influenced by the quality of the RGB-D sensor: use of a custom off-the-shelf high-resolution stereo-based sensor system outperformed a commercial low-resolution kinect-like camera in terms of required flattening iterations.
Li Sun 0005, Gerardo Aragon-Camarasa, Simon Rogers, J. Paul Siebert
ICRA3
2015 Incorporating peak grouping information for alignment of multiple liquid chromatography-mass spectrometry datasets
abstract
MOTIVATION: The combination of liquid chromatography and mass spectrometry (LC/MS) has been widely used for large-scale comparative studies in systems biology, including proteomics, glycomics and metabolomics. In almost all experimental design, it is necessary to compare chromatograms across biological or technical replicates and across sample groups. Central to this is the peak alignment step, which is one of the most important but challenging preprocessing steps. Existing alignment tools do not take into account the structural dependencies between related peaks that coelute and are derived from the same metabolite or peptide. We propose a direct matching peak alignment method for LC/MS data that incorporates related peaks information (within each LC/MS run) and investigate its effect on alignment performance (across runs). The groupings of related peaks necessary for our method can be obtained from any peak clustering method and are built into a pair-wise peak similarity score function. The similarity score matrix produced is used by an approximation algorithm for the weighted matching problem to produce the actual alignment result. RESULTS: We demonstrate that related peak information can improve alignment performance. The performance is evaluated on a set of benchmark datasets, where our method performs competitively compared to other popular alignment tools. AVAILABILITY: The proposed alignment method has been implemented as a stand-alone application in Python, available for download at http://github.com/joewandy/peak-grouping-alignment.
Joe Wandy, Rónán Daly, Rainer Breitling, Simon Rogers
Bioinform.4
2014 28 frames later: predicting screen touches from back-of-device grip changes
abstract
We demonstrate that front-of-screen targeting on mobile phones can be predicted from back-of-device grip manipulations. Using simple, low-resolution capacitive touch sensors placed around a standard phone, we outline a machine learning approach to modelling the grip modulation and inferring front-of-screen touch targets. We experimentally demonstrate that grip is a remarkably good predictor of touch, and we can predict touch position 200ms before contact with an accuracy of 18mm.
Mohammad Faizuddin Mohd Noor, Andrew Ramsay, Stephen Hughes, Simon Rogers, John Williamson 0001, Roderick Murray-Smith
CHI4
2014 Uncertain text entry on mobile devices
abstract
Users often struggle to enter text accurately on touchscreen keyboards. To address this, we present a flexible decoder for touchscreen text entry that combines probabilistic touch models with a language model. We investigate two different touch models. The first touch model is based on a Gaussian Process regression approach and implicitly models the inherent uncertainty of the touching process. The second touch model allows users to explicitly control the uncertainty via touch pressure. Using the first model we show that the character error rate can be reduced by up to 7% over a baseline method, and by up to 1.3% over a leading commercial keyboard. Using the second model we demonstrate that providing users with control over input certainty reduces the amount of text users have to correct manually and increases the text entry rate.
Daryl Weir, Henning Pohl, Simon Rogers, Keith Vertanen, Per Ola Kristensson
CHI3
2014 Detecting Missing Content Queries in an SMS-Based HIV/AIDS FAQ Retrieval System
Edwin Thuma, Simon Rogers, Iadh Ounis
ECIR2
2014 MetAssign: probabilistic annotation of metabolites from LC-MS data using a Bayesian clustering approach
abstract
MOTIVATION: The use of liquid chromatography coupled to mass spectrometry has enabled the high-throughput profiling of the metabolite composition of biological samples. However, the large amount of data obtained can be difficult to analyse and often requires computational processing to understand which metabolites are present in a sample. This article looks at the dual problem of annotating peaks in a sample with a metabolite, together with putatively annotating whether a metabolite is present in the sample. The starting point of the approach is a Bayesian clustering of peaks into groups, each corresponding to putative adducts and isotopes of a single metabolite. RESULTS: The Bayesian modelling introduced here combines information from the mass-to-charge ratio, retention time and intensity of each peak, together with a model of the inter-peak dependency structure, to increase the accuracy of peak annotation. The results inherently contain a quantitative estimate of confidence in the peak annotations and allow an accurate trade-off between precision and recall. Extensive validation experiments using authentic chemical standards show that this system is able to produce more accurate putative identifications than other state-of-the-art systems, while at the same time giving a probabilistic measure of confidence in the annotations. AVAILABILITY AND IMPLEMENTATION: The software has been implemented as part of the mzMatch metabolomics analysis pipeline, which is available for download at http://mzmatch.sourceforge.net/.
Rónán Daly, Simon Rogers, Joe Wandy, Andris Jankevics, Karl E. V. Burgess, Rainer Breitling
Bioinform.2
2014 Stronger findings for metabolomics through Bayesian modeling of multiple peaks and compound correlations
abstract
MOTIVATION: Data analysis for metabolomics suffers from uncertainty because of the noisy measurement technology and the small sample size of experiments. Noise and the small sample size lead to a high probability of false findings. Further, individual compounds have natural variation between samples, which in many cases renders them unreliable as biomarkers. However, the levels of similar compounds are typically highly correlated, which is a phenomenon that we model in this work. RESULTS: We propose a hierarchical Bayesian model for inferring differences between groups of samples more accurately in metabolomic studies, where the observed compounds are collinear. We discover that the method decreases the error of weak and non-existent covariate effects, and thereby reduces false-positive findings. To achieve this, the method makes use of the mass spectral peak data by clustering similar peaks into latent compounds, and by further clustering latent compounds into groups that respond in a coherent way to the experimental covariates. We demonstrate the method with three simulated studies and validate it with a metabolomic benchmark dataset. AVAILABILITY AND IMPLEMENTATION: An implementation in R is available at http://research.ics.aalto.fi/mi/software/peakANOVA/.
Tommi Suvitaival, Simon Rogers, Samuel Kaski
Bioinform.2
2014 Stronger findings from mass spectral data through multi-peak modeling
abstract
BACKGROUND: Mass spectrometry-based metabolomic analysis depends upon the identification of spectral peaks by their mass and retention time. Statistical analysis that follows the identification currently relies on one main peak of each compound. However, a compound present in the sample typically produces several spectral peaks due to its isotopic properties and the ionization process of the mass spectrometer device. In this work, we investigate the extent to which these additional peaks can be used to increase the statistical strength of differential analysis. RESULTS: We present a Bayesian approach for integrating data of multiple detected peaks that come from one compound. We demonstrate the approach through a simulated experiment and validate it on ultra performance liquid chromatography-mass spectrometry (UPLC-MS) experiments for metabolomics and lipidomics. Peaks that are likely to be associated with one compound can be clustered by the similarity of their chromatographic shape. Changes of concentration between sample groups can be inferred more accurately when multiple peaks are available. CONCLUSIONS: When the sample-size is limited, the proposed multi-peak approach improves the accuracy at inferring covariate effects. An R implementation and data are available at http://research.ics.aalto.fi/mi/software/peakANOVA/.
Tommi Suvitaival, Simon Rogers, Samuel Kaski
BMC Bioinform.2
2013 ODE parameter inference using adaptive gradient matching with Gaussian processes
abstract
Parameter inference in mechanistic models based on systems of coupled differential equations is a topical yet computationally challenging problem, due to the need to follow each parameter adaptation with a numerical integration of the differential equations. Techniques based on gradient matching, which aim to minimize the discrepancy between the slope of a data interpolant and the derivatives predicted from the differential equations, offer a computationally appealing shortcut to the inference problem. The present paper discusses a method based on nonparametric Bayesian statistics with Gaussian processes due to Calderhead et al. (2008), and shows how inference in this model can be substantially improved by consistently sampling from the joint distribution of the ODE parameters and GP hyperparameters. We demonstrate the efficiency of our adaptive gradient matching technique on three benchmark systems, and perform a detailed comparison with the method in Calderhead et al. (2008) and the explicit ODE integration approach, both in terms of parameter inference accuracy and in terms of computational efficiency.
Frank Dondelinger, Dirk Husmeier, Simon Rogers, Maurizio Filippone
AISTATS3
2013 User-specific touch models in a cross-device context
abstract
We present a machine learning approach to train user-specific offset models, which map actual to intended touch locations to improve accuracy. We propose a flexible framework to adapt and apply models trained on touch data from one device and user to others. This paper presents a study of the first published experimental data from multiple devices per user, and indicates that models not only improve accuracy between repeated sessions for the same user, but across devices and users, too. Device-specific models outperform unadapted user-specific models from different devices. However, with both user- and device-specific data, we demonstrate that our approach allows to combine this information to adapt models to the targeted device resulting in significant improvement. On average, adapted models improved accuracy by over 8%. We show that models can be obtained from a small number of touches (≈60). We also apply models to predict input-styles and identify users.
Daniel Buschek, Simon Rogers, Roderick Murray-Smith
Mobile HCI2
2013 Sparse selection of training data for touch correction systems
abstract
Touch offset models which improve input accuracy on mobile touch screen devices typically require the use of a large number of training points. In this paper, we describe a method for selecting training points such that high performance can be attained with fewer data. We use the Relevance Vector Machine (RVM) algorithm, and show that performance improvements can be obtained with fewer than 10 training examples. We show that the distribution of training points is conserved across users and contains interesting structure, and compare the RVM to two other offset prediction models for small training set sizes.
Daryl Weir, Daniel Buschek, Simon Rogers
Mobile HCI3
2013 Exploiting Query Logs and Field-Based Models to Address Term Mismatch in an HIV/AIDS FAQ Retrieval System
Edwin Thuma, Simon Rogers, Iadh Ounis
NLDB2
2013 Investigating the Disagreement Between Clinicians' Ratings of Patients in ICUs
abstract
We present a Bayesian analysis of ordinal annotations made by clinicians of patients in intensive care. In particular, we investigate the different ways in which clinicians can disagree and how their disagreement is reduced once they take part in a recently proposed procedure (INSIGHT) that aims at improving consistency. The model combines a nonparametric function (loosely interpretable as the health of the patient) with clinician-specific generative procedures for producing the observed ordinal values. Our analysis provides valuable details of the rating behavior of the individual clinicians and shows that the INSIGHT procedure is particularly effective at removing (some) clinician-specific inconsistencies and biases.
Simon Rogers, Derek H. Sleeman, John Kinsella
IEEE J. Biomed. Health Informatics1
2012 A user-specific machine learning approach for improving touch accuracy on mobile devices
abstract
We present a flexible Machine Learning approach for learning user-specific touch input models to increase touch accuracy on mobile devices. The model is based on flexible, non-parametric Gaussian Process regression and is learned using recorded touch inputs. We demonstrate that significant touch accuracy improvements can be obtained when either raw sensor data is used as an input or when the device's reported touch location is used as an input, with the latter marginally outperforming the former. We show that learned offset functions are highly nonlinear and user-specific and that user-specific models outperform models trained on data pooled from several users. Crucially, significant performance improvements can be obtained with a small (≈200) number of training examples, easily obtained for a particular user through a calibration game or from keyboard entry data.
Daryl Weir, Simon Rogers, Roderick Murray-Smith, Markus Löchtefeld
UIST2
2011 AnglePose: robust, precise capacitive touch tracking via 3d orientation estimation
abstract
We present a finger-tracking system for touch-based interaction which can track 3D finger angle in addition to position, using low-resolution conventional capacitive sensors, therefore compensating for the inaccuracy due to pose variation in conventional touch systems. Probabilistic inference about the pose of the finger is carried out in real-time using a particle filter; this results in an efficient and robust pose estimator which also gives appropriate uncertainty estimates. We show empirically that tracking the full pose of the finger results in greater accuracy in pointing tasks with small targets than competitive techniques. Our model can detect and cope with different finger sizes and the use of either fingers or thumbs, bringing a significant potential for improvement in one-handed interaction with touch devices. In addition to the gain in accuracy we also give examples of how this technique could open up the space of novel interactions.
Simon Rogers, John Williamson 0001, Craig D. Stewart, Roderick Murray-Smith
CHI1
2010 FingerCloud: uncertainty and autonomy handover incapacitive sensing
abstract
We describe a particle filtering approach to inferring finger movements on capacitive sensing arrays. This technique allows the efficient combination of human movement models with accurate sensing models, and gives high-fidelity results with low-resolution sensor grids and tracks finger height. Our model provides uncertainty estimates, which can be linked to the interaction to provide appropriately smoothed responses as sensing perfomance degrades; system autonomy is increased as estimates of user behaviour become less certain. We demonstrate the particle filter approach with a map browser running with a very small sensor board, where finger position uncertainty is linked to autonomy handover.
Simon Rogers, John Williamson 0001, Craig D. Stewart, Roderick Murray-Smith
CHI1
2010 Infinite factorization of multiple non-parametric views
Simon Rogers, Arto Klami, Janne Sinkkonen, Mark A. Girolami, Samuel Kaski
Mach. Learn.1
2009 Probabilistic assignment of formulas to mass peaks in metabolomics experiments
abstract
MOTIVATION: High-accuracy mass spectrometry is a popular technology for high-throughput measurements of cellular metabolites (metabolomics). One of the major challenges is the correct identification of the observed mass peaks, including the assignment of their empirical formula, based on the measured mass. RESULTS: We propose a novel probabilistic method for the assignment of empirical formulas to mass peaks in high-throughput metabolomics mass spectrometry measurements. The method incorporates information about possible biochemical transformations between the empirical formulas to assign higher probability to formulas that could be created from other metabolites in the sample. In a series of experiments, we show that the method performs well and provides greater insight than assignments based on mass alone. In addition, we extend the model to incorporate isotope information to achieve even more reliable formula identification. AVAILABILITY: A supplementary document, Matlab code, data and further information are available from http://www.dcs.gla.ac.uk/inference/metsamp.
Simon Rogers, Richard A. Scheltema, Mark A. Girolami, Rainer Breitling
Bioinform.1
2008 Investigating the correspondence between transcriptomic and proteomic expression profiles using coupled cluster models
abstract
MOTIVATION: Modern transcriptomics and proteomics enable us to survey the expression of RNAs and proteins at large scales. While these data are usually generated and analyzed separately, there is an increasing interest in comparing and co-analyzing transcriptome and proteome expression data. A major open question is whether transcriptome and proteome expression is linked and how it is coordinated. RESULTS: Here we have developed a probabilistic clustering model that permits analysis of the links between transcriptomic and proteomic profiles in a sensible and flexible manner. Our coupled mixture model defines a prior probability distribution over the component to which a protein profile should be assigned conditioned on which component the associated mRNA profile belongs to. We apply this approach to a large dataset of quantitative transcriptomic and proteomic expression data obtained from a human breast epithelial cell line (HMEC). The results reveal a complex relationship between transcriptome and proteome with most mRNA clusters linked to at least two protein clusters, and vice versa. A more detailed analysis incorporating information on gene function from the Gene Ontology database shows that a high correlation of mRNA and protein expression is limited to the components of some molecular machines, such as the ribosome, cell adhesion complexes and the TCP-1 chaperonin involved in protein folding. AVAILABILITY: Matlab code is available from the authors on request.
Simon Rogers, Mark A. Girolami, Walter Kolch, Katrina M. Waters, Tao Liu 0003, Brian Thrall, H. Steven Wiley
Bioinform.1
2007 Bayesian model-based inference of transcription factor activity
abstract
BACKGROUND: In many approaches to the inference and modeling of regulatory interactions using microarray data, the expression of the gene coding for the transcription factor is considered to be an accurate surrogate for the true activity of the protein it produces. There are many instances where this is inaccurate due to post-translational modifications of the transcription factor protein. Inference of the activity of the transcription factor from the expression of its targets has predominantly involved linear models that do not reflect the nonlinear nature of transcription. We extend a recent approach to inferring the transcription factor activity based on nonlinear Michaelis-Menten kinetics of transcription from maximum likelihood to fully Bayesian inference and give an example of how the model can be further developed. RESULTS: We present results on synthetic and real microarray data. Additionally, we illustrate how gene and replicate specific delays can be incorporated into the model. CONCLUSION: We demonstrate that full Bayesian inference is appropriate in this application and has several benefits over the maximum likelihood approach, especially when the volume of data is limited. We also show the benefits of using a non-linear model over a linear model, particularly in the case of repression.
Simon Rogers, Raya Khanin, Mark A. Girolami
BMC Bioinform.1
2006 Variational Bayesian Multinomial Probit Regression with Gaussian Process Priors
abstract
It is well known in the statistics literature that augmenting binary and polychotomous response models with gaussian latent variables enables exact Bayesian analysis via Gibbs sampling from the parameter posterior. By adopting such a data augmentation strategy, dispensing with priors over regression coefficients in favor of gaussian process (GP) priors over functions, and employing variational approximations to the full posterior, we obtain efficient computational methods for GP classification in the multiclass setting.1 The model augmentation with additional latent variables ensures full a posteriori class coupling while retaining the simple a priori independent GP covariance structure from which sparse approximations, such as multiclass informative vector machines (IVM), emerge in a natural and straightforward manner. This is the first time that a fully variational Bayesian treatment for multiclass GP classification has been developed without having to resort to additional explicit approximations to the nongaussian likelihood term. Empirical comparisons with exact analysis use Markov Chain Monte Carlo (MCMC) and Laplace approximations illustrate the utility of the variational approximation as a computationally economic alternative to full MCMC and it is shown to be more accurate than the Laplace approximation.
Mark A. Girolami, Simon Rogers
Neural Comput.2
2005 Hierarchic Bayesian models for kernel learning
abstract
The integration of diverse forms of informative data by learning an optimal combination of base kernels in classification or regression problems can provide enhanced performance when compared to that obtained from any single data source. We present a Bayesian hierarchical model which enables kernel learning and present effective variational Bayes estimators for regression and classification. Illustrative experiments demonstrate the utility of the proposed method. Matlab code replicating results reported is available at http://www.dcs.gla.ac.uk/~srogers/kernel_comb.html.
Mark A. Girolami, Simon Rogers
ICML2
2005 A Bayesian regression approach to the inference of regulatory networks from gene expression data
abstract
Motivation: There is currently much interest in reverse-engineering regulatory relationships between genes from microarray expression data. We propose a new algorithmic method for inferring such interactions between genes using data from gene knockout experiments. The algorithm we use is the Sparse Bayesian regression algorithm of Tipping and Faul. This method is highly suited to this problem as it does not require the data to be discretized, overcomes the need for an explicit topology search and, most importantly, requires no heuristic thresholding of the discovered connections. Results: Using simulated expression data, we are able to show that this algorithm outperforms a recently published correlation-based approach. Crucially, it does this without the need to set any ad hoc threshold on possible connections. Availability: Matlab code which allows all experimental results to be reproduced is available at http://www.dcs.gla.ac.uk/~srogers/reg_nets.html Contact: [email protected] Supplementary information: Appendices and supplementary figures mentioned in the text can be found at http://www.dcs.gla.ac.uk/~srogers/reg_nets.html
Simon Rogers, Mark A. Girolami
Bioinform.1
2005 The Latent Process Decomposition of cDNA Microarray Data Sets
abstract
We present a new computational technique (a software implementation, data sets, and supplementary information are available at http://www.enm.bris.ac.uk/lpd/) which enables the probabilistic analysis of cDNA microarray data and we demonstrate its effectiveness in identifying features of biomedical importance. A hierarchical Bayesian model, called Latent Process Decomposition (LPD), is introduced in which each sample in the data set is represented as a combinatorial mixture over a finite set of latent processes, which are expected to correspond to biological processes. Parameters in the model are estimated using efficient variational methods. This type of probabilistic model is most appropriate for the interpretation of measurement data generated by cDNA microarray technology. For determining informative substructure in such data sets, the proposed model has several important advantages over the standard use of dendrograms. First, the ability to objectively assess the optimal number of sample clusters. Second, the ability to represent samples and gene expression levels using a common set of latent variables (dendrograms cluster samples and gene expression values separately which amounts to two distinct reduced space representations). Third, in constrast to standard cluster models, observations are not assigned to a single cluster and, thus, for example, gene expression levels are modeled via combinations of the latent processes identified by the algorithm. We show this new method compares favorably with alternative cluster analysis methods. To illustrate its potential, we apply the proposed technique to several microarray data sets for cancer. For these data sets it successfully decomposes the data into known subtypes and indicates possible further taxonomic subdivision in addition to highlighting, in a wholly unsupervised manner, the importance of certain genes which are known to be medically significant. To illustrate its wider applicability, we also illustrate its performance on a microarray data set for yeast.
Simon Rogers, Mark A. Girolami, Colin Campbell, Rainer Breitling
IEEE ACM Trans. Comput. Biol. Bioinform.1