Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marie-Pier Scott-Boyer

dblp:136/2552 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0003-2626-4584ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
multi-omics data integration
1.122022
timeOmics: an R package for longitudinal multi-omics data integration · Bioinform. 2022
KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021
Bioinformatics and computational biology › bioinformatics infrastructure
biological data management
0.512021
KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021
Bioinformatics and computational biology › statistical genetics › quantitative trait locus mapping
expression quantitative trait loci analysis
0.212013
iBMQ: a R/Bioconductor package for integrated Bayesian modeling of eQTL data · Bioinform. 2013
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP
0.012013
iBMQ: a R/Bioconductor package for integrated Bayesian modeling of eQTL data · Bioinform. 2013
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.012013
iBMQ: a R/Bioconductor package for integrated Bayesian modeling of eQTL data · Bioinform. 2013

Methods — techniques the papers use, named apart from their topics

preprocessing · 0.6modeling · 0.6clustering · 0.6uniform data exchange model · 0.5elasticsearch-based storage · 0.5markov chain monte carlo · 0.3hierarchical bayesian modeling · 0.3
YearPublicationVenuePosition
2023 Use of Elasticsearch-based business intelligence tools for integration and visualization of biological data
abstract
The emergence of massive datasets exploring the multiple levels of molecular biology has made their analysis and knowledge transfer more complex. Flexible tools to manage big biological datasets could be of great help for standardizing the usage of developed data visualizations and integration methods. Business intelligence (BI) tools have been used in many fields as exploratory tools. They have numerous connectors to link numerous data repositories with a unified graphic interface, offering an overview of data and facilitating interpretation for decision makers. BI tools could be a flexible and user-friendly way of handling molecular biological data with interactive visualizations. However, it is rather uncommon to see such tools used for the exploration of massive and complex datasets in biological fields. We believe that two main obstacles could be the reason. Firstly, we posit that the way to import data into BI tools are not compatible with biological databases. Secondly, BI tools may not be adapted to certain particularities of complex biological data, namely, the size, the variability of datasets and the availability of specialized visualizations. This paper highlights the use of five BI tools (Elastic Kibana, Siren Investigate, Microsoft Power BI, Salesforce Tableau and Apache Superset) onto which the massive data management repository engine called Elasticsearch is compatible. Four case studies will be discussed in which these BI tools were applied on biological datasets with different characteristics. We conclude that the performance of the tools depends on the complexity of the biological questions and the size of the datasets.
Marie-Pier Scott-Boyer, Pascal Dufour, François Belleau, Régis Ongaro-Carcy, Clément Plessis, Olivier Périn, Arnaud Droit
Briefings Bioinform.1
2022 timeOmics: an R package for longitudinal multi-omics data integration
abstract
MOTIVATION: Multi-omics data integration enables the global analysis of biological systems and discovery of new biological insights. Multi-omics experimental designs have been further extended with a longitudinal dimension to study dynamic relationships between molecules. However, methods that integrate longitudinal multi-omics data are still in their infancy. RESULTS: We introduce the R package timeOmics, a generic analytical framework for the integration of longitudinal multi-omics data. The framework includes pre-processing, modeling and clustering to identify molecular features strongly associated with time. We illustrate this framework in a case study to detect seasonal patterns of mRNA, metabolites, gut taxa and clinical variables in patients with diabetes mellitus from the integrative Human Microbiome Project. AVAILABILITYAND IMPLEMENTATION: timeOmics is available on Bioconductor and github.com/abodein/timeOmics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Antoine Bodein, Marie-Pier Scott-Boyer, Olivier Périn, Kim-Anh Lê Cao, Arnaud Droit
Bioinform.2
2021 KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology
abstract
MOTIVATION: The growing production of massive heterogeneous biological data offers opportunities for new discoveries. However, performing multi-omics data analysis is challenging, and researchers are forced to handle the ever-increasing complexity of both data management and evolution of our biological understanding. Substantial efforts have been made to unify biological datasets into integrated systems. Unfortunately, they are not easily scalable, deployable and searchable, locally or globally. RESULTS: This publication presents two tools with a simple structure that can help any data provider, organization or researcher, requiring a reliable data search and analysis base. The first tool is Kibio, a scalable and adaptable data storage based on Elasticsearch search engine. The second tool is KibioR, a R package to pull, push and search Kibio datasets or any accessible Elasticsearch-based databases. These tools apply a uniform data exchange model and minimize the burden of data management by organizing data into a decentralized, versatile, searchable and shareable structure. Several case studies are presented using multiple databases, from drug characterization to miRNAs and pathways identification, emphasizing the ease of use and versatility of the Kibio/KibioR framework. AVAILABILITYAND IMPLEMENTATION: Both KibioR and Elasticsearch are open source. KibioR package source is available at https://github.com/regisoc/kibior and the library on CRAN at https://cran.r-project.org/package=kibior. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Régis Ongaro-Carcy, Marie-Pier Scott-Boyer, Adrien Dessemond, François Belleau, Mickaël Leclercq, Olivier Périn, Arnaud Droit
Bioinform.2
2021 GWENA: gene co-expression networks analysis and extended modules characterization in a single Bioconductor package
abstract
BACKGROUND: Network-based analysis of gene expression through co-expression networks can be used to investigate modular relationships occurring between genes performing different biological functions. An extended description of each of the network modules is therefore a critical step to understand the underlying processes contributing to a disease or a phenotype. Biological integration, topology study and conditions comparison (e.g. wild vs mutant) are the main methods to do so, but to date no tool combines them all into a single pipeline. RESULTS: Here we present GWENA, a new R package that integrates gene co-expression network construction and whole characterization of the detected modules through gene set enrichment, phenotypic association, hub genes detection, topological metric computation, and differential co-expression. To demonstrate its performance, we applied GWENA on two skeletal muscle datasets from young and old patients of GTEx study. Remarkably, we prioritized a gene whose involvement was unknown in the muscle development and growth. Moreover, new insights on the variations in patterns of co-expression were identified. The known phenomena of connectivity loss associated with aging was found coupled to a global reorganization of the relationships leading to expression of known aging related functions. CONCLUSION: GWENA is an R package available through Bioconductor ( https://bioconductor.org/packages/release/bioc/html/GWENA.html ) that has been developed to perform extended analysis of gene co-expression networks. Thanks to biological and topological information as well as differential co-expression, the package helps to dissect the role of genes relationships in diseases conditions or targeted phenotypes. GWENA goes beyond existing packages that perform co-expression analysis by including new tools to fully characterize modules, such as differential co-expression, additional enrichment databases, and network visualization.
Gwenaëlle G. Lemoine, Marie-Pier Scott-Boyer, Bathilde Ambroise, Olivier Périn, Arnaud Droit
BMC Bioinform.2
2019 Multi-omics integration - a comparison of unsupervised clustering methodologies
abstract
With the recent developments in the field of multi-omics integration, the interest in factors such as data preprocessing, choice of the integration method and the number of different omics considered had increased. In this work, the impact of these factors is explored when solving the problem of sample classification, by comparing the performances of five unsupervised algorithms: Multiple Canonical Correlation Analysis, Multiple Co-Inertia Analysis, Multiple Factor Analysis, Joint and Individual Variation Explained and Similarity Network Fusion. These methods were applied to three real data sets taken from literature and several ad hoc simulated scenarios to discuss classification performance in different conditions of noise and signal strength across the data types. The impact of experimental design, feature selection and parameter training has been also evaluated to unravel important conditions that can affect the accuracy of the result.
Giulia Tini, Luca Marchetti, Corrado Priami, Marie-Pier Scott-Boyer
Briefings Bioinform.4
2013 iBMQ: a R/Bioconductor package for integrated Bayesian modeling of eQTL data
abstract
MOTIVATION: Recently, mapping studies of expression quantitative loci (eQTL) (where gene expression levels are viewed as quantitative traits) have provided insight into the biology of gene regulation. Bayesian methods provide natural modeling frameworks for analyzing eQTL studies, where information shared across markers and/or genes can increase the power to detect eQTLs. Bayesian approaches tend to be computationally demanding and require specialized software. As a result, most eQTL studies use univariate methods treating each gene independently, leading to suboptimal results. RESULTS: We present a powerful, computationally optimized and free open-source R package, iBMQ. Our package implements a joint hierarchical Bayesian model where all genes and SNPs are modeled concurrently. Model parameters are estimated using a Markov chain Monte Carlo algorithm. The free and widely used openMP parallel library speeds up computation. Using a mouse cardiac dataset, we show that iBMQ improves the detection of large trans-eQTL hotspots compared with other state-of-the-art packages for eQTL analysis. AVAILABILITY: The R-package iBMQ is available from the Bioconductor Web site at http://bioconductor.org and runs on Linux, Windows and MAC OS X. It is distributed under the Artistic Licence-2.0 terms. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Greg C. Imholte, Marie-Pier Scott-Boyer, Aurélie Labbe, Christian F. Deschepper, Raphael Gottardo
Bioinform.2