Arnaud Droit

dblp:35/533 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-7922-790XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
multi-omics data integration
1.122022
timeOmics: an R package for longitudinal multi-omics data integration · Bioinform. 2022
KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
deep clustering
0.612022
Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022
Machine learning › Representation and self-supervised learning
mutual information
0.612022
Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022
Information theory › information measures › mutual information
generalized mutual information
0.612022
Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022
Information theory › information measures
mutual information
0.612022
Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022
Bioinformatics and computational biology › bioinformatics infrastructure
biological data management
0.512021
KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021
Bioinformatics and computational biology › statistical genetics
fine-mapping
0.312017
VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017
Bioinformatics and computational biology › genomics
genome-wide association study
0.312017
VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017
Bioinformatics and computational biology › proteomics
protein identification
0.212014
rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014
Bioinformatics and computational biology
proteomics
0.212014
rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014
Bioinformatics and computational biology › proteomics › mass spectrometry data analysis
tandem mass spectrometry database searching
0.212014
rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014
Bioinformatics and computational biology › epigenomics
ChIP-chip analysis
0.112010
rMAT - an R/Bioconductor package for analyzing ChIP-chip experiments · Bioinform. 2010
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
0.112010
rMAT - an R/Bioconductor package for analyzing ChIP-chip experiments · Bioinform. 2010
Bioinformatics and computational biology › genomics
variant annotation
0.112017
VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017
Bioinformatics and computational biology › genomics
next-generation sequencing data analysis
0.012013
NGS++: a library for rapid prototyping of epigenomics software tools · Bioinform. 2013

Methods — techniques the papers use, named apart from their topics

kullback-leibler divergence · 1.1kernel methods · 1.1preprocessing · 0.6modeling · 0.6clustering · 0.6uniform data exchange model · 0.5elasticsearch-based storage · 0.5x!tandem algorithm · 0.2r/bioconductor integration · 0.2c++11 · 0.2statistical enrichment analysis · 0.1r package · 0.1
YearPublicationVenuePosition
2023 Use of Elasticsearch-based business intelligence tools for integration and visualization of biological data
abstract
The emergence of massive datasets exploring the multiple levels of molecular biology has made their analysis and knowledge transfer more complex. Flexible tools to manage big biological datasets could be of great help for standardizing the usage of developed data visualizations and integration methods. Business intelligence (BI) tools have been used in many fields as exploratory tools. They have numerous connectors to link numerous data repositories with a unified graphic interface, offering an overview of data and facilitating interpretation for decision makers. BI tools could be a flexible and user-friendly way of handling molecular biological data with interactive visualizations. However, it is rather uncommon to see such tools used for the exploration of massive and complex datasets in biological fields. We believe that two main obstacles could be the reason. Firstly, we posit that the way to import data into BI tools are not compatible with biological databases. Secondly, BI tools may not be adapted to certain particularities of complex biological data, namely, the size, the variability of datasets and the availability of specialized visualizations. This paper highlights the use of five BI tools (Elastic Kibana, Siren Investigate, Microsoft Power BI, Salesforce Tableau and Apache Superset) onto which the massive data management repository engine called Elasticsearch is compatible. Four case studies will be discussed in which these BI tools were applied on biological datasets with different characteristics. We conclude that the performance of the tools depends on the complexity of the biological questions and the size of the datasets.
Marie-Pier Scott-Boyer, Pascal Dufour, François Belleau, Régis Ongaro-Carcy, Clément Plessis, Olivier Périn, Arnaud Droit
Briefings Bioinform.7
2022 Generalised Mutual Information for Discriminative Clustering
abstract
In the last decade, recent successes in deep clustering majorly involved the mutual information (MI) as an unsupervised objective for training neural networks with increasing regularisations. While the quality of the regularisations have been largely discussed for improvements, little attention has been dedicated to the relevance of MI as a clustering objective. In this paper, we first highlight how the maximisation of MI does not lead to satisfying clusters. We identified the Kullback-Leibler divergence as the main reason of this behaviour. Hence, we generalise the mutual information by changing its core distance, introducing the generalised mutual information (GEMINI): a set of metrics for unsupervised neural network training. Unlike MI, some GEMINIs do not require regularisations when training. Some of these metrics are geometry-aware thanks to distances or kernels in the data space. Finally, we highlight that GEMINIs can automatically select a relevant number of clusters, a property that has been little studied in deep clustering context where the number of clusters is a priori unknown.
Louis Ohl, Pierre-Alexandre Mattei, Charles Bouveyron, Warith Harchaoui, Mickaël Leclercq, Arnaud Droit, Frédéric Precioso
NeurIPS6
2022 timeOmics: an R package for longitudinal multi-omics data integration
abstract
MOTIVATION: Multi-omics data integration enables the global analysis of biological systems and discovery of new biological insights. Multi-omics experimental designs have been further extended with a longitudinal dimension to study dynamic relationships between molecules. However, methods that integrate longitudinal multi-omics data are still in their infancy. RESULTS: We introduce the R package timeOmics, a generic analytical framework for the integration of longitudinal multi-omics data. The framework includes pre-processing, modeling and clustering to identify molecular features strongly associated with time. We illustrate this framework in a case study to detect seasonal patterns of mRNA, metabolites, gut taxa and clinical variables in patients with diabetes mellitus from the integrative Human Microbiome Project. AVAILABILITYAND IMPLEMENTATION: timeOmics is available on Bioconductor and github.com/abodein/timeOmics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Antoine Bodein, Marie-Pier Scott-Boyer, Olivier Périn, Kim-Anh Lê Cao, Arnaud Droit
Bioinform.5
2021 KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology
abstract
MOTIVATION: The growing production of massive heterogeneous biological data offers opportunities for new discoveries. However, performing multi-omics data analysis is challenging, and researchers are forced to handle the ever-increasing complexity of both data management and evolution of our biological understanding. Substantial efforts have been made to unify biological datasets into integrated systems. Unfortunately, they are not easily scalable, deployable and searchable, locally or globally. RESULTS: This publication presents two tools with a simple structure that can help any data provider, organization or researcher, requiring a reliable data search and analysis base. The first tool is Kibio, a scalable and adaptable data storage based on Elasticsearch search engine. The second tool is KibioR, a R package to pull, push and search Kibio datasets or any accessible Elasticsearch-based databases. These tools apply a uniform data exchange model and minimize the burden of data management by organizing data into a decentralized, versatile, searchable and shareable structure. Several case studies are presented using multiple databases, from drug characterization to miRNAs and pathways identification, emphasizing the ease of use and versatility of the Kibio/KibioR framework. AVAILABILITYAND IMPLEMENTATION: Both KibioR and Elasticsearch are open source. KibioR package source is available at https://github.com/regisoc/kibior and the library on CRAN at https://cran.r-project.org/package=kibior. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Régis Ongaro-Carcy, Marie-Pier Scott-Boyer, Adrien Dessemond, François Belleau, Mickaël Leclercq, Olivier Périn, Arnaud Droit
Bioinform.7
2021 GWENA: gene co-expression networks analysis and extended modules characterization in a single Bioconductor package
abstract
BACKGROUND: Network-based analysis of gene expression through co-expression networks can be used to investigate modular relationships occurring between genes performing different biological functions. An extended description of each of the network modules is therefore a critical step to understand the underlying processes contributing to a disease or a phenotype. Biological integration, topology study and conditions comparison (e.g. wild vs mutant) are the main methods to do so, but to date no tool combines them all into a single pipeline. RESULTS: Here we present GWENA, a new R package that integrates gene co-expression network construction and whole characterization of the detected modules through gene set enrichment, phenotypic association, hub genes detection, topological metric computation, and differential co-expression. To demonstrate its performance, we applied GWENA on two skeletal muscle datasets from young and old patients of GTEx study. Remarkably, we prioritized a gene whose involvement was unknown in the muscle development and growth. Moreover, new insights on the variations in patterns of co-expression were identified. The known phenomena of connectivity loss associated with aging was found coupled to a global reorganization of the relationships leading to expression of known aging related functions. CONCLUSION: GWENA is an R package available through Bioconductor ( https://bioconductor.org/packages/release/bioc/html/GWENA.html ) that has been developed to perform extended analysis of gene co-expression networks. Thanks to biological and topological information as well as differential co-expression, the package helps to dissect the role of genes relationships in diseases conditions or targeted phenotypes. GWENA goes beyond existing packages that perform co-expression analysis by including new tools to fully characterize modules, such as differential co-expression, additional enrichment databases, and network visualization.
Gwenaëlle G. Lemoine, Marie-Pier Scott-Boyer, Bathilde Ambroise, Olivier Périn, Arnaud Droit
BMC Bioinform.5
2017 VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis
abstract
Motivation: The identification of the functional variants responsible for observed genome-wide association studies (GWAS) signals is one of the most challenging tasks of the post-GWAS research era. Several tools have been developed to annotate genetic variants by their genomic location and potential functional implications. Each of these tools has its own requirements and internal logic, which forces the user to become acquainted with each interface. Results: From an awareness of the amount of work needed to analyze a single locus, we have built a flexible, versatile and easy-to-use web interface designed to help in prioritizing variants and predicting their potential functional implications. This interface acts as a single-point of entry linking association results with reference tools and relevant experiments. Availability and Implementation: VEXOR is an integrative web application implemented through the Shiny framework and available at: http://romix.genome.ulaval.ca/vexor. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Audrey Lemaçon, Charles Joly Beauparlant, Penny Soucy, Jamie Allen, Douglas F. Easton, Peter Kraft, Jacques Simard, Arnaud Droit
Bioinform.8
2016 metagene Profiles Analyses Reveal Regulatory Element's Factor-Specific Recruitment Patterns
abstract
ChIP-Sequencing (ChIP-Seq) provides a vast amount of information regarding the localization of proteins across the genome. The aggregation of ChIP-Seq enrichment signal in a metagene plot is an approach commonly used to summarize data complexity and to obtain a high level visual representation of the general occupancy pattern of a protein. Here we present the R package metagene, the graphical interface Imetagene and the companion package similaRpeak. Together, they provide a framework to integrate, summarize and compare the ChIP-Seq enrichment signal from complex experimental designs. Those packages identify and quantify similarities or dissimilarities in patterns between large numbers of ChIP-Seq profiles. We used metagene to investigate the differential occupancy of regulatory factors at noncoding regulatory regions (promoters and enhancers) in relation to transcriptional activity in GM12878 B-lymphocytes. The relationships between occupancy patterns and transcriptional activity suggest two different mechanisms of action for transcriptional control: i) a "gradient effect" where the regulatory factor occupancy levels follow transcription and ii) a "threshold effect" where the regulatory factor occupancy levels max out prior to reaching maximal transcription. metagene, Imetagene and similaRpeak are implemented in R under the Artistic license 2.0 and are available on Bioconductor.
Charles Joly Beauparlant, Fabien C. Lamaze, Astrid Deschênes, Rawane Samb, Audrey Lemaçon, Pascal Belleau, Steve Bilodeau, Arnaud Droit
PLoS Comput. Biol.8
2014 rTANDEM, an R/Bioconductor package for MS/MS protein identification
abstract
SUMMARY: rTANDEM is an R/Bioconductor package that interfaces the X!Tandem protein identification algorithm. The package can run the multi-threaded algorithm on proteomic data files directly from R. It also provides functions to convert search parameters and results to/from R as well as functions to manipulate parameters and automate searches. An associated R package, shinyTANDEM, provides a web-based graphical interface to visualize and interpret the results. Together, those two packages form an entry point for a general MS/MS-based proteomic pipeline in R/Bioconductor. AVAILABILITY AND IMPLEMENTATION: rTANDEM and shinyTANDEM are distributed in R/Bioconductor, http://bioconductor.org/packages/release/bioc/. The packages are under open licenses (GPL-3 and Artistice-1.0). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Frédéric Fournier, Charles Joly Beauparlant, René Paradis, Arnaud Droit
Bioinform.4
2013 NGS++: a library for rapid prototyping of epigenomics software tools
abstract
MOTIVATION: The development of computational tools to enable testing and analysis of high-throughput-sequencing data is essential to modern genomics research. However, although multiple frameworks have been developed to facilitate access to these tools, comparatively little effort has been made at implementing low-level programming libraries to increase the speed and ease of their development. RESULTS: We propose NGS++, a programming library in C++11 specialized in manipulating both next-generation sequencing (NGS) datasets and genomic information files. This library allows easy integration of new formats and rapid prototyping of new functionalities with a focus on the analysis of genomic regions and features. It offers a powerful, yet versatile and easily extensible interface to read, write and manipulate multiple genomic file formats. By standardizing the internal data structures and presenting a common interface to the data parser, NGS++ offers an effective framework for epigenomics tool development. AVAILABILITY: NGS++ was written in C++ using the C++11 standard. It requires minimal efforts to build and is well-documented via a complete docXygen guide, online documentation and tutorials. Source code, tests, code examples and documentation are available via the website at http://www.ngsplusplus.ca and the github repository at https://github.com/NGS-lib/NGSplusplus. CONTACT: [email protected] or [email protected].
Alexei Nordell-Markovits, Charles Joly Beauparlant, Dominique Toupin, Shengrui Wang, Arnaud Droit, Nicolas Gevry
Bioinform.5
2010 rMAT - an R/Bioconductor package for analyzing ChIP-chip experiments
abstract
SUMMARY: Chromatin immunoprecipitation combined with DNA microarrays (ChIP-chip) has evolved as a popular technique to study DNA-protein binding or post-translational chromatin/histone modifications at the genomic level. However, the raw microarray intensities generate a massive amount of data, creating a need for efficient analysis algorithms and statistical methods to identify enriched regions. RESULTS: We present a fast, free and powerful, open source R package, rMAT, that allows the identification of regions enriched for transcription factor binding sites in ChIP-chip experiments on Affymetrix tiling arrays. AVAILABILITY: The R-package rMAT is available from the Bioconductor web site at http://bioconductor.org and runs on Linux, MAC OS and MS-Windows. rMAT is distributed under the terms of the Artistic Licence 2.0.
Arnaud Droit, Charles Cheung, Raphael Gottardo
Bioinform.1
2008 Quality assessment of peptide tandem mass spectra
abstract
BACKGROUND: Tandem mass spectrometry has emerged as a cornerstone of high throughput proteomic studies owing in part to various high throughput search engines which are used to interpret these tandem mass spectra. However, majority of experimental tandem mass spectra cannot be interpreted by any existing methods. There are many reasons why this happens. However, one of the most important reasons is that majority of experimental spectra are of too poor quality to be interpretable. It wastes time to interpret these "uninterpretable" spectra by any methods. On the other hand, some spectra of high quality are not able to get a score high enough to be interpreted by existing search engines because there are many similar peptides in the searched database. However, such spectra may be good enough to be interpreted by de novo methods or manually verifying methods. Therefore, it is worth in developing a method for assessing spectral quality, which can used for filtering the spectra of poor quality before any interpretation attempts or for finding the most potential candidates for de novo methods or manually verifying methods. RESULTS: This paper develops a novel method to assess the quality of tandem mass spectra, which can eliminate majority of poor quality spectra while losing very minority of high quality spectra. First, a number of features are proposed to describe the quality of tandem mass spectra. The proposed method maps each tandem spectrum into a feature vector. Then Fisher linear discriminant analysis (FLDA) is employed to construct the classifier (the filter) which discriminates the high quality spectra from the poor quality ones. The proposed method has been tested on two tandem mass spectra datasets acquired by ion trap mass spectrometers. CONCLUSION: Computational experiments illustrate that the proposed method outperforms the existing ones. The proposed method is generic, and is expected to be applicable to assessing the quality of spectra acquired by instruments other than ion trap mass spectrometers.
Fang-Xiang Wu, Pierre Gagné, Arnaud Droit, Guy G. Poirier
BMC Bioinform.3
2007 PARPs database: A LIMS systems for protein-protein interaction data mining or laboratory information management system
abstract
BACKGROUND: In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteins and the rapid advancement of this technique, in combination with other proteomics methods, results in an increasing amount of proteome data. This data must be archived and analysed using specialized bioinformatics tools. DESCRIPTION: We herein describe "PARPs database," a data analysis and management pipeline for liquid chromatography tandem mass spectrometry (LC-MS/MS) proteomics. PARPs database is a web-based tool whose features include experiment annotation, protein database searching, protein sequence management, as well as data-mining of the peptides and proteins identified. CONCLUSION: Using this pipeline, we have successfully identified several interactions of biological significance between PARP-1 and other proteins, namely RFC-1, 2, 3, 4 and 5.
Arnaud Droit, Joanna M. Hunter, Michèle Rouleau, Chantal Ethier, Aude Picard-Cloutier, David Bourgais, Guy G. Poirier
BMC Bioinform.1