VLDB 2026 Research / reviewers in the wild / expert
Arnaud Droit
dblp:35/533
· DBLP profile ↗
12ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-7922-790XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
multi-omics data integration |
1.1 | 2 | 2022 | timeOmics: an R package for longitudinal multi-omics data integration · Bioinform. 2022 KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
deep clustering |
0.6 | 1 | 2022 | Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning
mutual information |
0.6 | 1 | 2022 | Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022 |
Information theory › information measures › mutual information
generalized mutual information |
0.6 | 1 | 2022 | Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022 |
Information theory › information measures
mutual information |
0.6 | 1 | 2022 | Generalised Mutual Information for Discriminative Clustering · NeurIPS 2022 |
Bioinformatics and computational biology › bioinformatics infrastructure
biological data management |
0.5 | 1 | 2021 | KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biology · Bioinform. 2021 |
Bioinformatics and computational biology › statistical genetics
fine-mapping |
0.3 | 1 | 2017 | VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
genome-wide association study |
0.3 | 1 | 2017 | VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017 |
Bioinformatics and computational biology › proteomics
protein identification |
0.2 | 1 | 2014 | rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014 |
Bioinformatics and computational biology
proteomics |
0.2 | 1 | 2014 | rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014 |
Bioinformatics and computational biology › proteomics › mass spectrometry data analysis
tandem mass spectrometry database searching |
0.2 | 1 | 2014 | rTANDEM, an R/Bioconductor package for MS/MS protein identification · Bioinform. 2014 |
Bioinformatics and computational biology › epigenomics
ChIP-chip analysis |
0.1 | 1 | 2010 | rMAT - an R/Bioconductor package for analyzing ChIP-chip experiments · Bioinform. 2010 |
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction |
0.1 | 1 | 2010 | rMAT - an R/Bioconductor package for analyzing ChIP-chip experiments · Bioinform. 2010 |
Bioinformatics and computational biology › genomics
variant annotation |
0.1 | 1 | 2017 | VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysis · Bioinform. 2017 |
Bioinformatics and computational biology › genomics
next-generation sequencing data analysis |
0.0 | 1 | 2013 | NGS++: a library for rapid prototyping of epigenomics software tools · Bioinform. 2013 |
Methods — techniques the papers use, named apart from their topics
kullback-leibler divergence · 1.1kernel methods · 1.1preprocessing · 0.6modeling · 0.6clustering · 0.6uniform data exchange model · 0.5elasticsearch-based storage · 0.5x!tandem algorithm · 0.2r/bioconductor integration · 0.2c++11 · 0.2statistical enrichment analysis · 0.1r package · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Use of Elasticsearch-based business intelligence tools for integration and visualization of biological dataabstractThe emergence of massive datasets exploring the multiple levels of molecular biology has made their analysis and knowledge transfer more complex. Flexible tools to manage big biological datasets could be of great help for standardizing the usage of developed data visualizations and integration methods. Business intelligence (BI) tools have been used in many fields as exploratory tools. They have numerous connectors to link numerous data repositories with a unified graphic interface, offering an overview of data and facilitating interpretation for decision makers. BI tools could be a flexible and user-friendly way of handling molecular biological data with interactive visualizations. However, it is rather uncommon to see such tools used for the exploration of massive and complex datasets in biological fields. We believe that two main obstacles could be the reason. Firstly, we posit that the way to import data into BI tools are not compatible with biological databases. Secondly, BI tools may not be adapted to certain particularities of complex biological data, namely, the size, the variability of datasets and the availability of specialized visualizations. This paper highlights the use of five BI tools (Elastic Kibana, Siren Investigate, Microsoft Power BI, Salesforce Tableau and Apache Superset) onto which the massive data management repository engine called Elasticsearch is compatible. Four case studies will be discussed in which these BI tools were applied on biological datasets with different characteristics. We conclude that the performance of the tools depends on the complexity of the biological questions and the size of the datasets. Marie-Pier Scott-Boyer, Pascal Dufour, François Belleau, Régis Ongaro-Carcy, Clément Plessis, Olivier Périn, Arnaud Droit |
Briefings Bioinform. | 7 |
| 2022 | Generalised Mutual Information for Discriminative ClusteringabstractIn the last decade, recent successes in deep clustering majorly involved the mutual information (MI) as an unsupervised objective for training neural networks with increasing regularisations. While the quality of the regularisations have been largely discussed for improvements, little attention has been dedicated to the relevance of MI as a clustering objective. In this paper, we first highlight how the maximisation of MI does not lead to satisfying clusters. We identified the Kullback-Leibler divergence as the main reason of this behaviour. Hence, we generalise the mutual information by changing its core distance, introducing the generalised mutual information (GEMINI): a set of metrics for unsupervised neural network training. Unlike MI, some GEMINIs do not require regularisations when training. Some of these metrics are geometry-aware thanks to distances or kernels in the data space. Finally, we highlight that GEMINIs can automatically select a relevant number of clusters, a property that has been little studied in deep clustering context where the number of clusters is a priori unknown. Louis Ohl, Pierre-Alexandre Mattei, Charles Bouveyron, Warith Harchaoui, Mickaël Leclercq, Arnaud Droit, Frédéric Precioso |
NeurIPS | 6 |
| 2022 | timeOmics: an R package for longitudinal multi-omics data integrationabstractMOTIVATION: Multi-omics data integration enables the global analysis of biological systems and discovery of new biological insights. Multi-omics experimental designs have been further extended with a longitudinal dimension to study dynamic relationships between molecules. However, methods that integrate longitudinal multi-omics data are still in their infancy. RESULTS: We introduce the R package timeOmics, a generic analytical framework for the integration of longitudinal multi-omics data. The framework includes pre-processing, modeling and clustering to identify molecular features strongly associated with time. We illustrate this framework in a case study to detect seasonal patterns of mRNA, metabolites, gut taxa and clinical variables in patients with diabetes mellitus from the integrative Human Microbiome Project. AVAILABILITYAND IMPLEMENTATION: timeOmics is available on Bioconductor and github.com/abodein/timeOmics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antoine Bodein, Marie-Pier Scott-Boyer, Olivier Périn, Kim-Anh Lê Cao, Arnaud Droit |
Bioinform. | 5 |
| 2021 | KibioR & Kibio: a new architecture for next-generation data querying and sharing in big biologyabstractMOTIVATION: The growing production of massive heterogeneous biological data offers opportunities for new discoveries. However, performing multi-omics data analysis is challenging, and researchers are forced to handle the ever-increasing complexity of both data management and evolution of our biological understanding. Substantial efforts have been made to unify biological datasets into integrated systems. Unfortunately, they are not easily scalable, deployable and searchable, locally or globally. RESULTS: This publication presents two tools with a simple structure that can help any data provider, organization or researcher, requiring a reliable data search and analysis base. The first tool is Kibio, a scalable and adaptable data storage based on Elasticsearch search engine. The second tool is KibioR, a R package to pull, push and search Kibio datasets or any accessible Elasticsearch-based databases. These tools apply a uniform data exchange model and minimize the burden of data management by organizing data into a decentralized, versatile, searchable and shareable structure. Several case studies are presented using multiple databases, from drug characterization to miRNAs and pathways identification, emphasizing the ease of use and versatility of the Kibio/KibioR framework. AVAILABILITYAND IMPLEMENTATION: Both KibioR and Elasticsearch are open source. KibioR package source is available at https://github.com/regisoc/kibior and the library on CRAN at https://cran.r-project.org/package=kibior. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Régis Ongaro-Carcy, Marie-Pier Scott-Boyer, Adrien Dessemond, François Belleau, Mickaël Leclercq, Olivier Périn, Arnaud Droit |
Bioinform. | 7 |
| 2021 | GWENA: gene co-expression networks analysis and extended modules characterization in a single Bioconductor packageabstractBACKGROUND: Network-based analysis of gene expression through co-expression networks can be used to investigate modular relationships occurring between genes performing different biological functions. An extended description of each of the network modules is therefore a critical step to understand the underlying processes contributing to a disease or a phenotype. Biological integration, topology study and conditions comparison (e.g. wild vs mutant) are the main methods to do so, but to date no tool combines them all into a single pipeline. RESULTS: Here we present GWENA, a new R package that integrates gene co-expression network construction and whole characterization of the detected modules through gene set enrichment, phenotypic association, hub genes detection, topological metric computation, and differential co-expression. To demonstrate its performance, we applied GWENA on two skeletal muscle datasets from young and old patients of GTEx study. Remarkably, we prioritized a gene whose involvement was unknown in the muscle development and growth. Moreover, new insights on the variations in patterns of co-expression were identified. The known phenomena of connectivity loss associated with aging was found coupled to a global reorganization of the relationships leading to expression of known aging related functions. CONCLUSION: GWENA is an R package available through Bioconductor ( https://bioconductor.org/packages/release/bioc/html/GWENA.html ) that has been developed to perform extended analysis of gene co-expression networks. Thanks to biological and topological information as well as differential co-expression, the package helps to dissect the role of genes relationships in diseases conditions or targeted phenotypes. GWENA goes beyond existing packages that perform co-expression analysis by including new tools to fully characterize modules, such as differential co-expression, additional enrichment databases, and network visualization. Gwenaëlle G. Lemoine, Marie-Pier Scott-Boyer, Bathilde Ambroise, Olivier Périn, Arnaud Droit |
BMC Bioinform. | 5 |
| 2017 | VEXOR: an integrative environment for prioritization of functional variants in fine-mapping analysisabstractMotivation: The identification of the functional variants responsible for observed genome-wide association studies (GWAS) signals is one of the most challenging tasks of the post-GWAS research era. Several tools have been developed to annotate genetic variants by their genomic location and potential functional implications. Each of these tools has its own requirements and internal logic, which forces the user to become acquainted with each interface. Results: From an awareness of the amount of work needed to analyze a single locus, we have built a flexible, versatile and easy-to-use web interface designed to help in prioritizing variants and predicting their potential functional implications. This interface acts as a single-point of entry linking association results with reference tools and relevant experiments. Availability and Implementation: VEXOR is an integrative web application implemented through the Shiny framework and available at: http://romix.genome.ulaval.ca/vexor. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Audrey Lemaçon, Charles Joly Beauparlant, Penny Soucy, Jamie Allen, Douglas F. Easton, Peter Kraft, Jacques Simard, Arnaud Droit |
Bioinform. | 8 |
| 2016 | metagene Profiles Analyses Reveal Regulatory Element's Factor-Specific Recruitment PatternsabstractChIP-Sequencing (ChIP-Seq) provides a vast amount of information regarding the localization of proteins across the genome. The aggregation of ChIP-Seq enrichment signal in a metagene plot is an approach commonly used to summarize data complexity and to obtain a high level visual representation of the general occupancy pattern of a protein. Here we present the R package metagene, the graphical interface Imetagene and the companion package similaRpeak. Together, they provide a framework to integrate, summarize and compare the ChIP-Seq enrichment signal from complex experimental designs. Those packages identify and quantify similarities or dissimilarities in patterns between large numbers of ChIP-Seq profiles. We used metagene to investigate the differential occupancy of regulatory factors at noncoding regulatory regions (promoters and enhancers) in relation to transcriptional activity in GM12878 B-lymphocytes. The relationships between occupancy patterns and transcriptional activity suggest two different mechanisms of action for transcriptional control: i) a "gradient effect" where the regulatory factor occupancy levels follow transcription and ii) a "threshold effect" where the regulatory factor occupancy levels max out prior to reaching maximal transcription. metagene, Imetagene and similaRpeak are implemented in R under the Artistic license 2.0 and are available on Bioconductor. Charles Joly Beauparlant, Fabien C. Lamaze, Astrid Deschênes, Rawane Samb, Audrey Lemaçon, Pascal Belleau, Steve Bilodeau, Arnaud Droit |
PLoS Comput. Biol. | 8 |
| 2014 | rTANDEM, an R/Bioconductor package for MS/MS protein identificationabstractSUMMARY: rTANDEM is an R/Bioconductor package that interfaces the X!Tandem protein identification algorithm. The package can run the multi-threaded algorithm on proteomic data files directly from R. It also provides functions to convert search parameters and results to/from R as well as functions to manipulate parameters and automate searches. An associated R package, shinyTANDEM, provides a web-based graphical interface to visualize and interpret the results. Together, those two packages form an entry point for a general MS/MS-based proteomic pipeline in R/Bioconductor. AVAILABILITY AND IMPLEMENTATION: rTANDEM and shinyTANDEM are distributed in R/Bioconductor, http://bioconductor.org/packages/release/bioc/. The packages are under open licenses (GPL-3 and Artistice-1.0). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Frédéric Fournier, Charles Joly Beauparlant, René Paradis, Arnaud Droit |
Bioinform. | 4 |
| 2013 | NGS++: a library for rapid prototyping of epigenomics software toolsabstractMOTIVATION: The development of computational tools to enable testing and analysis of high-throughput-sequencing data is essential to modern genomics research. However, although multiple frameworks have been developed to facilitate access to these tools, comparatively little effort has been made at implementing low-level programming libraries to increase the speed and ease of their development. RESULTS: We propose NGS++, a programming library in C++11 specialized in manipulating both next-generation sequencing (NGS) datasets and genomic information files. This library allows easy integration of new formats and rapid prototyping of new functionalities with a focus on the analysis of genomic regions and features. It offers a powerful, yet versatile and easily extensible interface to read, write and manipulate multiple genomic file formats. By standardizing the internal data structures and presenting a common interface to the data parser, NGS++ offers an effective framework for epigenomics tool development. AVAILABILITY: NGS++ was written in C++ using the C++11 standard. It requires minimal efforts to build and is well-documented via a complete docXygen guide, online documentation and tutorials. Source code, tests, code examples and documentation are available via the website at http://www.ngsplusplus.ca and the github repository at https://github.com/NGS-lib/NGSplusplus. CONTACT: [email protected] or [email protected]. Alexei Nordell-Markovits, Charles Joly Beauparlant, Dominique Toupin, Shengrui Wang, Arnaud Droit, Nicolas Gevry |
Bioinform. | 5 |
| 2010 | rMAT - an R/Bioconductor package for analyzing ChIP-chip experimentsabstractSUMMARY: Chromatin immunoprecipitation combined with DNA microarrays (ChIP-chip) has evolved as a popular technique to study DNA-protein binding or post-translational chromatin/histone modifications at the genomic level. However, the raw microarray intensities generate a massive amount of data, creating a need for efficient analysis algorithms and statistical methods to identify enriched regions. RESULTS: We present a fast, free and powerful, open source R package, rMAT, that allows the identification of regions enriched for transcription factor binding sites in ChIP-chip experiments on Affymetrix tiling arrays. AVAILABILITY: The R-package rMAT is available from the Bioconductor web site at http://bioconductor.org and runs on Linux, MAC OS and MS-Windows. rMAT is distributed under the terms of the Artistic Licence 2.0. Arnaud Droit, Charles Cheung, Raphael Gottardo |
Bioinform. | 1 |
| 2008 | Quality assessment of peptide tandem mass spectraabstractBACKGROUND: Tandem mass spectrometry has emerged as a cornerstone of high throughput proteomic studies owing in part to various high throughput search engines which are used to interpret these tandem mass spectra. However, majority of experimental tandem mass spectra cannot be interpreted by any existing methods. There are many reasons why this happens. However, one of the most important reasons is that majority of experimental spectra are of too poor quality to be interpretable. It wastes time to interpret these "uninterpretable" spectra by any methods. On the other hand, some spectra of high quality are not able to get a score high enough to be interpreted by existing search engines because there are many similar peptides in the searched database. However, such spectra may be good enough to be interpreted by de novo methods or manually verifying methods. Therefore, it is worth in developing a method for assessing spectral quality, which can used for filtering the spectra of poor quality before any interpretation attempts or for finding the most potential candidates for de novo methods or manually verifying methods. RESULTS: This paper develops a novel method to assess the quality of tandem mass spectra, which can eliminate majority of poor quality spectra while losing very minority of high quality spectra. First, a number of features are proposed to describe the quality of tandem mass spectra. The proposed method maps each tandem spectrum into a feature vector. Then Fisher linear discriminant analysis (FLDA) is employed to construct the classifier (the filter) which discriminates the high quality spectra from the poor quality ones. The proposed method has been tested on two tandem mass spectra datasets acquired by ion trap mass spectrometers. CONCLUSION: Computational experiments illustrate that the proposed method outperforms the existing ones. The proposed method is generic, and is expected to be applicable to assessing the quality of spectra acquired by instruments other than ion trap mass spectrometers. Fang-Xiang Wu, Pierre Gagné, Arnaud Droit, Guy G. Poirier |
BMC Bioinform. | 3 |
| 2007 | PARPs database: A LIMS systems for protein-protein interaction data mining or laboratory information management systemabstractBACKGROUND: In the "post-genome" era, mass spectrometry (MS) has become an important method for the analysis of proteins and the rapid advancement of this technique, in combination with other proteomics methods, results in an increasing amount of proteome data. This data must be archived and analysed using specialized bioinformatics tools. DESCRIPTION: We herein describe "PARPs database," a data analysis and management pipeline for liquid chromatography tandem mass spectrometry (LC-MS/MS) proteomics. PARPs database is a web-based tool whose features include experiment annotation, protein database searching, protein sequence management, as well as data-mining of the peptides and proteins identified. CONCLUSION: Using this pipeline, we have successfully identified several interactions of biological significance between PARP-1 and other proteins, namely RFC-1, 2, 3, 4 and 5. Arnaud Droit, Joanna M. Hunter, Michèle Rouleau, Chantal Ethier, Aude Picard-Cloutier, David Bourgais, Guy G. Poirier |
BMC Bioinform. | 1 |