Paul R. Burton

dblp:166/8799 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0001-5799-9634ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 81% Medical and health informatics · 19%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › statistical genetics › genetic association study
genetic association study design
0.212015
ESPRESSO: taking into account assessment errors on outcome and exposures in power analysis for association studies · Bioinform. 2015
Bioinformatics and computational biology
power analysis
0.212015
ESPRESSO: taking into account assessment errors on outcome and exposures in power analysis for association studies · Bioinform. 2015
Bioinformatics and computational biology › biostatistics
sample size determination
0.212015
ESPRESSO: taking into account assessment errors on outcome and exposures in power analysis for association studies · Bioinform. 2015
Medical and health informatics › clinical informatics
clinical data management
0.112015
Data Safe Havens in health research and healthcare · Bioinform. 2015

Methods — techniques the papers use, named apart from their topics

multi-task learning · 0.6federated learning · 0.6DataSHIELD · 0.6simulation-based power analysis · 0.2
YearPublicationVenuePosition
2024 Federated privacy-protected meta- and mega-omics data analysis in multi-center studies with a fully open-source analytic platform
abstract
The importance of maintaining data privacy and complying with regulatory requirements is highlighted especially when sharing omic data between different research centers. This challenge is even more pronounced in the scenario where a multi-center effort for collaborative omics studies is necessary. OmicSHIELD is introduced as an open-source tool aimed at overcoming these challenges by enabling privacy-protected federated analysis of sensitive omic data. In order to ensure this, multiple security mechanisms have been included in the software. This innovative tool is capable of managing a wide range of omic data analyses specifically tailored to biomedical research. These include genome and epigenome wide association studies and differential gene expression analyses. OmicSHIELD is designed to support both meta- and mega-analysis, so that it offers a wide range of capabilities for different analysis designs. We present a series of use cases illustrating some examples of how the software addresses real-world analyses of omic data.
Xavier Escriba-Montagut, Yannick Marcon, Augusto Anguita-Ruiz, Demetris Avraam, Jose Urquiza, Andrei S. Morgan, Rebecca C. Wilson, Paul R. Burton, Juan R. González
PLoS Comput. Biol.8
2022 dsMTL: a computational framework for privacy-preserving, distributed multi-task machine learning
abstract
MOTIVATION: In multi-cohort machine learning studies, it is critical to differentiate between effects that are reproducible across cohorts and those that are cohort-specific. Multi-task learning (MTL) is a machine learning approach that facilitates this differentiation through the simultaneous learning of prediction tasks across cohorts. Since multi-cohort data can often not be combined into a single storage solution, there would be the substantial utility of an MTL application for geographically distributed data sources. RESULTS: Here, we describe the development of 'dsMTL', a computational framework for privacy-preserving, distributed multi-task machine learning that includes three supervised and one unsupervised algorithms. First, we derive the theoretical properties of these methods and the relevant machine learning workflows to ensure the validity of the software implementation. Second, we implement dsMTL as a library for the R programming language, building on the DataSHIELD platform that supports the federated analysis of sensitive individual-level data. Third, we demonstrate the applicability of dsMTL for comorbidity modeling in distributed data. We show that comorbidity modeling using dsMTL outperformed conventional, federated machine learning, as well as the aggregation of multiple models built on the distributed datasets individually. The application of dsMTL was computationally efficient and highly scalable when applied to moderate-size (n < 500), real expression data given the actual network latency. AVAILABILITY AND IMPLEMENTATION: dsMTL is freely available at https://github.com/transbioZI/dsMTLBase (server-side package) and https://github.com/transbioZI/dsMTLClient (client-side package). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Han Cao 0001, Youcheng Zhang, Jan Baumbach, Paul R. Burton, Dominic B. Dwyer, Nikolaos Koutsouleris, Julian O. Matschinske, Yannick Marcon, Sivanesan Rajan, Thilo Rieg, Patricia Ryser-Welch, Julian Späth, Carl Herrmann, Emanuel Schwarz
Bioinform.4
2021 Orchestrating privacy-protected big data analyses of data from different resources with R and DataSHIELD
abstract
Combined analysis of multiple, large datasets is a common objective in the health- and biosciences. Existing methods tend to require researchers to physically bring data together in one place or follow an analysis plan and share results. Developed over the last 10 years, the DataSHIELD platform is a collection of R packages that reduce the challenges of these methods. These include ethico-legal constraints which limit researchers' ability to physically bring data together and the analytical inflexibility associated with conventional approaches to sharing results. The key feature of DataSHIELD is that data from research studies stay on a server at each of the institutions that are responsible for the data. Each institution has control over who can access their data. The platform allows an analyst to pass commands to each server and the analyst receives results that do not disclose the individual-level data of any study participants. DataSHIELD uses Opal which is a data integration system used by epidemiological studies and developed by the OBiBa open source project in the domain of bioinformatics. However, until now the analysis of big data with DataSHIELD has been limited by the storage formats available in Opal and the analysis capabilities available in the DataSHIELD R packages. We present a new architecture ("resources") for DataSHIELD and Opal to allow large, complex datasets to be used at their original location, in their original format and with external computing facilities. We provide some real big data analysis examples in genomics and geospatial projects. For genomic data analyses, we also illustrate how to extend the resources concept to address specific big data infrastructures such as GA4GH or EGA, and make use of shell commands. Our new infrastructure will help researchers to perform data analyses in a privacy-protected way from existing data sharing initiatives or projects. To help researchers use this framework, we describe selected packages and present an online book (https://isglobal-brge.github.io/resource_bookdown).
Yannick Marcon, Tom Bishop, Demetris Avraam, Xavier Escriba-Montagut, Patricia Ryser-Welch, Stuart Wheater, Paul R. Burton, Juan R. González
PLoS Comput. Biol.7
2015 Data Safe Havens in health research and healthcare
abstract
MOTIVATION: The data that put the 'evidence' into 'evidence-based medicine' are central to developments in public health, primary and hospital care. A fundamental challenge is to site such data in repositories that can easily be accessed under appropriate technical and governance controls which are effectively audited and are viewed as trustworthy by diverse stakeholders. This demands socio-technical solutions that may easily become enmeshed in protracted debate and controversy as they encounter the norms, values, expectations and concerns of diverse stakeholders. In this context, the development of what are called 'Data Safe Havens' has been crucial. Unfortunately, the origins and evolution of the term have led to a range of different definitions being assumed by different groups. There is, however, an intuitively meaningful interpretation that is often assumed by those who have not previously encountered the term: a repository in which useful but potentially sensitive data may be kept securely under governance and informatics systems that are fit-for-purpose and appropriately tailored to the nature of the data being maintained, and may be accessed and utilized by legitimate users undertaking work and research contributing to biomedicine, health and/or to ongoing development of healthcare systems. RESULTS: This review explores a fundamental question: 'what are the specific criteria that ought reasonably to be met by a data repository if it is to be seen as consistent with this interpretation and viewed as worthy of being accorded the status of 'Data Safe Haven' by key stakeholders'? We propose 12 such criteria. CONTACT: [email protected].
Paul R. Burton, Madeleine J. Murtagh, Andy Boyd, James Bryan Williams, Edward S. Dove, Susan E. Wallace, Anne-Marie Tassé, Julian Little, Rex L. Chisholm, Amadou Gaye, Kristian Hveem, Anthony J. Brookes, Pat Goodwin, Jon Fistein, Martin Bobrow, Bartha M. Knoppers
Bioinform.1
2015 ESPRESSO: taking into account assessment errors on outcome and exposures in power analysis for association studies
abstract
MOTIVATION: Very large studies are required to provide sufficiently big sample sizes for adequately powered association analyses. This can be an expensive undertaking and it is important that an accurate sample size is identified. For more realistic sample size calculation and power analysis, the impact of unmeasured aetiological determinants and the quality of measurement of both outcome and explanatory variables should be taken into account. Conventional methods to analyse power use closed-form solutions that are not flexible enough to cater for all of these elements easily. They often result in a potentially substantial overestimation of the actual power. RESULTS: In this article, we describe the Estimating Sample-size and Power in R by Exploring Simulated Study Outcomes tool that allows assessment errors in power calculation under various biomedical scenarios to be incorporated. We also report a real world analysis where we used this tool to answer an important strategic question for an existing cohort. AVAILABILITY AND IMPLEMENTATION: The software is available for online calculation and downloads at http://espresso-research.org. The code is freely available at https://github.com/ESPRESSO-research. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Amadou Gaye, Thomas W. Y. Burton, Paul R. Burton
Bioinform.3