Guillaume Bourque

dblp:78/833 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-3933-9656ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Information theory · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Network and information security
1 paper
Privacy and data protection · 100%
Artificial intelligence
1 paper
Optimization for machine learning · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
epigenomics
1.532024
EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024
The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasets · Bioinform. 2019
eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data · Bioinform. 2019
Bioinformatics and computational biology
genomics
0.922024
RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023
EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024
Bioinformatics and computational biology › genomics › genomic data management
genomic data sharing
0.712023
RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023
Bioinformatics and computational biology › epigenomics › ChIP-seq analysis
peak detection
0.522017
Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017
PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count data · ICML 2015
Bioinformatics and computational biology
comparative genomics
0.522019
The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasets · Bioinform. 2019
webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010
Information theory › hypothesis testing
change-point detection
0.412020
Constrained Dynamic Programming and Supervised Penalty Learning Algorithms for Peak Detection in Genomic Data · J. Mach. Learn. Res. 2020
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis
0.312017
Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017
Data mining › predictive modeling
supervised learning
0.312017
Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017
Bioinformatics and computational biology › statistical genetics
genotype-phenotype analysis
0.212024
EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024
Privacy and data protection › health data privacy
genomic privacy
0.212024
EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024
Bioinformatics and computational biology › genomics
genomic data analysis
0.212015
PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count data · ICML 2015
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation analysis
0.112019
eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data · Bioinform. 2019
Bioinformatics and computational biology › comparative genomics
genome rearrangement
0.112010
webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010
Bioinformatics and computational biology
phylogenetics
0.112010
webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010

Methods — techniques the papers use, named apart from their topics

data aggregation · 1.5controlled-access data sharing · 1.5supervised learning · 0.9penalty learning · 0.9dynamic programming · 0.9matrix slicing · 0.7visual labeling · 0.6supervised machine learning · 0.6galaxy workflow · 0.4correlation analysis · 0.4maximum likelihood segmentation · 0.2constrained dynamic programming · 0.2
YearPublicationVenuePosition
2024 EpiVar Browser: advanced exploration of epigenomics data under controlled access
abstract
MOTIVATION: Human epigenomic data has been generated by large consortia for thousands of cell types to be used as a reference map of normal and disease chromatin states. Since epigenetic data contains potentially identifiable information, similarly to genetic data, most raw files generated by these consortia are stored in controlled-access databases. It is important to protect identifiable information, but this should not hinder secure sharing of these valuable datasets. RESULTS: Guided by the Framework for responsible sharing of genomic and health-related data from the Global Alliance for Genomics and Health (GA4GH), we have developed an approach and a tool to facilitate the exploration of epigenomics datasets' aggregate results, while filtering out identifiable information. Specifically, the EpiVar Browser allows a user to navigate an epigenetic dataset from a cohort of individuals and enables direct exploration of genotype-chromatin phenotype relationships. Because individual genotypes and epigenetic signal tracks are not directly accessible, and rather aggregated in the portal output, no identifiable data is released, yet the interface allows for dynamic genotype-epigenome interrogation. This approach has the potential to accelerate analyses that would otherwise require a lengthy multi-step approval process and provides a generalizable strategy to facilitate responsible access to sensitive epigenomics data. AVAILABILITY AND IMPLEMENTATION: Online portal: https://computationalgenomics.ca/tools/epivar; EpiVar Browser source code: https://github.com/c3g/epivar-browser; bw-merge-window tool source code: https://github.com/c3g/bw-merge-window.
David R. Lougheed, Hanshi Liu, Katherine A. Aracena, Romain Grégoire, Alain Pacis, Tomi Pastinen, Luis B. Barreiro, Yann Joly, David Bujold, Guillaume Bourque
Bioinform.10
2023 RNAget: an API to securely retrieve RNA quantifications
abstract
SUMMARY: Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. AVAILABILITY AND IMPLEMENTATION: https://ga4gh-rnaseq.github.io/schema/docs/index.html.
Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared Nedzel, Keenan Graham, Meenakshi S. Kagda, Pedro Assis, Benjamin C. Hitz, Emilio Righi, Roderic Guigó, Barbara J. Wold, Alvis Brazma, Julia Burchard, Joe Capka, Michael Cherry, Laura Clarke, Brian Craft, Manolis Dermitzakis, Mark Diekhans, John Dursi, Michael Sean Fitzsimons, Zac Flaming, Romina Garrido, Alfred Gil, Paul Godden, Matt Green, Mitch Guttman, Brian Haas, Max Haeussler, Sten Linnarsson, Adam Lipski, Simonne Longerich, David R. Lougheed, Jonathan Manning, John C. Marioni, Christopher Meyer, Stephen B. Montgomery, Alyssa Morrow, Alfonso Muñoz-Pomer Fuentes, Jared L. Nedzel, Kevin Osborn, Francis Ouellette, Irene Papatheodorou, Dmitri D. Pervouchine, Arun K. Ramani, Jordi Rambla De Argila, Bashir Sadjad, David Steinberg, Jeremiah Talkar, Timothy Tickle, Kathy Tzeng, Saman Vaisipour, Sean Watford, Barbara Wold
Bioinform.5
2020 Constrained Dynamic Programming and Supervised Penalty Learning Algorithms for Peak Detection in Genomic Data
abstract
Peak detection in genomic data involves segmenting counts of DNA sequence reads aligned to different locations of a chromosome. The goal is to detect peaks with higher counts, and filter out background noise with lower counts. Most existing algorithms for this problem are unsupervised heuristics tailored to patterns in specific data types. We propose a supervised framework for this problem, using optimal changepoint detection models with learned penalty functions. We propose the first dynamic programming algorithm that is guaranteed to compute the optimal solution to changepoint detection problems with constraints between adjacent segment mean parameters. Implementing this algorithm requires the choice of penalty parameter that determines the number of segments that are estimated. We show how the supervised learning ideas of Rigaill et al. (2013) can be used to choose this penalty. We compare the resulting implementation of our algorithm to several baselines in a benchmark of labeled ChIP-seq data sets with two different patterns (broad H3K36me3 data and sharp H3K4me3 data). Whereas baseline unsupervised methods only provide accurate peak detection for a single pattern, our supervised method achieves state-of-the-art accuracy in all data sets. The log-linear timings of our proposed dynamic programming algorithm make it scalable to the large genomic data sets that are now common. Our implementation is available in the PeakSegOptimal R package on CRAN.
Toby Hocking, Guillem Rigaill, Paul Fearnhead, Guillaume Bourque
J. Mach. Learn. Res.4
2019 The use of hyperspectral remote sensing to detect PCB contaminated soils in the 0.35 to 12 micron spectral range
abstract
Laboratory spectral measurements were conducted to evaluate the potential of hyperspectral technologies, both reflective and emissive, to detect Polychlorinated Biphenyl (PCB) in contaminated soils. Soil sample standards of silt, clay, sand, and mixed textures were contaminated with Polychlorinated Biphenyl (PCB) oil with concentrations varying between 0% and 10% (100 000 ppm) of PCB. An ASD (0.35 to 2.5 microns) and a FSR (1.6 to 12 microns) spectrometers were used to measure the reflectance of pure and contaminated soil. Two spectral regions, one in the SWIR and one in the MWIR were significantly correlated to the soil PCB concentrations. It was shown that 5W30 engine oil and PCB absorb similarly in the two spectral regions. Sand and clay soils did not respond the same way as the other soil type. The MWIR CH absorption band between 3330 and 3630 was found the most promising to detect PCB contaminated soils using hyperspectral remote sensing,
Josée Lévesque, Eldon Puckrin, Luc Levert, Guillaume Bourque
IGARSS4
2019 eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data
abstract
SUMMARY: The Illumina Infinium EPIC BeadChip is a new high-throughput array for DNA methylation analysis, extending the earlier 450k array by over 400 000 new sites. Previously, a method named eFORGE was developed to provide insights into cell type-specific and cell-composition effects for 450k data. Here, we present a significantly updated and improved version of eFORGE that can analyze both EPIC and 450k array data. New features include analysis of chromatin states, transcription factor motifs and DNase I footprints, providing tools for epigenome-wide association study interpretation and epigenome editing. AVAILABILITY AND IMPLEMENTATION: eFORGE v2.0 is implemented as a web tool available from https://eforge.altiusinstitute.org and https://eforge-tf.altiusinstitute.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Charles E. Breeze, Alex P. Reynolds, Jenny van Dongen, Ian Dunham, John Lazar, Shane J. Neph, Jeff Vierstra, Guillaume Bourque, Andrew E. Teschendorff, John A. Stamatoyannopoulos, Stephan Beck 0002
Bioinform.8
2019 The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasets
abstract
SUMMARY: In recent years, major initiatives such as the International Human Epigenome Consortium have generated thousands of high-quality genome-wide datasets for a large variety of assays and cell types. This data can be used as a reference to assess whether the signal from a user-provided dataset corresponds to its expected experiment, as well as to help reveal unexpected biological associations. We have developed the epiGenomic Efficient Correlator (epiGeEC) tool to enable genome-wide comparisons of very large numbers of datasets. A public Galaxy implementation of epiGeEC allows comparison of user datasets with thousands of public datasets in a few minutes. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://bitbucket.org/labjacquespe/epigeec and the Galaxy implementation at http://epigeec.genap.ca. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jonathan Laperle, Simon Hébert-Deschamps, Joanny Raby, David A. de Lima Morais, Michel Barrette, David Bujold, Charlotte Bastin, Marc-Antoine Robert, Jean-François Nadeau, Marie Harel, Alexei Nordell-Markovits, Alain Veilleux, Guillaume Bourque, Pierre-Étienne Jacques
Bioinform.13
2017 Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning
abstract
Motivation: Many peak detection algorithms have been proposed for ChIP-seq data analysis, but it is not obvious which algorithm and what parameters are optimal for any given dataset. In contrast, regions with and without obvious peaks can be easily labeled by visual inspection of aligned read counts in a genome browser. We propose a supervised machine learning approach for ChIP-seq data analysis, using labels that encode qualitative judgments about which genomic regions contain or do not contain peaks. The main idea is to manually label a small subset of the genome, and then learn a model that makes consistent peak predictions on the rest of the genome. Results: We created 7 new histone mark datasets with 12 826 visually determined labels, and analyzed 3 existing transcription factor datasets. We observed that default peak detection parameters yield high false positive rates, which can be reduced by learning parameters using a relatively small training set of labeled data from the same experiment type. We also observed that labels from different people are highly consistent. Overall, these data indicate that our supervised labeling method is useful for quantitatively training and testing peak detection algorithms. Availability and Implementation: Labeled histone mark data http://cbio.ensmp.fr/~thocking/chip-seq-chunk-db/ , R package to compute the label error of predicted peaks https://github.com/tdhock/PeakError. Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Toby Hocking, Patricia Goerner-Potvin, Andreanne Morin, Xiaojian Shao, Tomi Pastinen, Guillaume Bourque
Bioinform.6
2015 PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count data
abstract
Peak detection is a central problem in genomic data analysis, and current algorithms for this task are unsupervised and mostly effective for a single data type and pattern (e.g. H3K4me3 data with a sharp peak pattern). We propose PeakSeg, a new constrained maximum likelihood segmentation model for peak detection with an efficient inference algorithm: constrained dynamic programming. We investigate unsupervised and supervised learning of penalties for the critical model selection problem. We show that the supervised method has state-of-the-art peak detection across all data sets in a benchmark that includes both sharp H3K4me3 and broad H3K36me3 patterns.
Toby Hocking, Guillem Rigaill, Guillaume Bourque
ICML3
2010 Prediction of low coverage prone regions for Illumina sequencing projects using a support vector machine
abstract
Applications of next-generation sequencing technologies have the potential to bring revolutionary changes to medicine and biology. However, coverage bias can pose a challenge to short read data analysis tools, which rely on high coverage. To address this issue we have developed a support vector machine (SVM) based method for predicting low coverage prone (LCP) regions on a given genome. The developed SVM-based prediction of LCP regions on a given genome can assist data processing procedures based on Illumina sequencing technology, such as de novo sequencing and transcriptome analysis.
Zejun Zheng, Bertil Schmidt, Guillaume Bourque
BIBM3
2010 webMGR: an online tool for the multiple genome rearrangement problem
abstract
SUMMARY: The algorithm MGR enables the reconstruction of rearrangement phylogenies based on gene or synteny block order in multiple genomes. Although MGR has been successfully applied to study the evolution of different sets of species, its utilization has been hampered by the prohibitive running time for some applications. In the current work, we have designed new heuristics that significantly speed up the tool without compromising its accuracy. Moreover, we have developed a web server (webMGR) that includes elaborate web output to facilitate navigation through the results. AVAILABILITY: webMGR can be accessed via http://www.gis.a-star.edu.sg/~bourque. The source code of the improved standalone version of MGR is also freely available from the web site. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Chi Ho Lin, Sean Harry Lowcay, Atif Shahab, Guillaume Bourque
Bioinform.5