EDBT 2026 Demo / reviewers in the wild / expert
Guillaume Bourque
dblp:78/833
· DBLP profile ↗
10ranked-venue papers
0as first author
2since 2021 · last 2024
0000-0002-3933-9656ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Artificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Information theory · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Network and information security
1 paper |
Privacy and data protection · 100% | |
| Artificial intelligence
1 paper |
Optimization for machine learning · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
epigenomics |
1.5 | 3 | 2024 | EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024 The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasets · Bioinform. 2019 eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data · Bioinform. 2019 |
Bioinformatics and computational biology
genomics |
0.9 | 2 | 2024 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024 |
Bioinformatics and computational biology › genomics › genomic data management
genomic data sharing |
0.7 | 1 | 2023 | RNAget: an API to securely retrieve RNA quantifications · Bioinform. 2023 |
Bioinformatics and computational biology › epigenomics › ChIP-seq analysis
peak detection |
0.5 | 2 | 2017 | Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017 PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count data · ICML 2015 |
Bioinformatics and computational biology
comparative genomics |
0.5 | 2 | 2019 | The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasets · Bioinform. 2019 webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010 |
Information theory › hypothesis testing
change-point detection |
0.4 | 1 | 2020 | Constrained Dynamic Programming and Supervised Penalty Learning Algorithms for Peak Detection in Genomic Data · J. Mach. Learn. Res. 2020 |
Bioinformatics and computational biology › epigenomics
ChIP-seq analysis |
0.3 | 1 | 2017 | Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017 |
Data mining › predictive modeling
supervised learning |
0.3 | 1 | 2017 | Optimizing ChIP-seq peak detectors using visual labels and supervised machine learning · Bioinform. 2017 |
Bioinformatics and computational biology › statistical genetics
genotype-phenotype analysis |
0.2 | 1 | 2024 | EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024 |
Privacy and data protection › health data privacy
genomic privacy |
0.2 | 1 | 2024 | EpiVar Browser: advanced exploration of epigenomics data under controlled access · Bioinform. 2024 |
Bioinformatics and computational biology › genomics
genomic data analysis |
0.2 | 1 | 2015 | PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count data · ICML 2015 |
Bioinformatics and computational biology › epigenomics › DNA methylation
DNA methylation analysis |
0.1 | 1 | 2019 | eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data · Bioinform. 2019 |
Bioinformatics and computational biology › comparative genomics
genome rearrangement |
0.1 | 1 | 2010 | webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010 |
Bioinformatics and computational biology
phylogenetics |
0.1 | 1 | 2010 | webMGR: an online tool for the multiple genome rearrangement problem · Bioinform. 2010 |
Methods — techniques the papers use, named apart from their topics
data aggregation · 1.5controlled-access data sharing · 1.5supervised learning · 0.9penalty learning · 0.9dynamic programming · 0.9matrix slicing · 0.7visual labeling · 0.6supervised machine learning · 0.6galaxy workflow · 0.4correlation analysis · 0.4maximum likelihood segmentation · 0.2constrained dynamic programming · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EpiVar Browser: advanced exploration of epigenomics data under controlled accessabstractMOTIVATION: Human epigenomic data has been generated by large consortia for thousands of cell types to be used as a reference map of normal and disease chromatin states. Since epigenetic data contains potentially identifiable information, similarly to genetic data, most raw files generated by these consortia are stored in controlled-access databases. It is important to protect identifiable information, but this should not hinder secure sharing of these valuable datasets. RESULTS: Guided by the Framework for responsible sharing of genomic and health-related data from the Global Alliance for Genomics and Health (GA4GH), we have developed an approach and a tool to facilitate the exploration of epigenomics datasets' aggregate results, while filtering out identifiable information. Specifically, the EpiVar Browser allows a user to navigate an epigenetic dataset from a cohort of individuals and enables direct exploration of genotype-chromatin phenotype relationships. Because individual genotypes and epigenetic signal tracks are not directly accessible, and rather aggregated in the portal output, no identifiable data is released, yet the interface allows for dynamic genotype-epigenome interrogation. This approach has the potential to accelerate analyses that would otherwise require a lengthy multi-step approval process and provides a generalizable strategy to facilitate responsible access to sensitive epigenomics data. AVAILABILITY AND IMPLEMENTATION: Online portal: https://computationalgenomics.ca/tools/epivar; EpiVar Browser source code: https://github.com/c3g/epivar-browser; bw-merge-window tool source code: https://github.com/c3g/bw-merge-window. David R. Lougheed, Hanshi Liu, Katherine A. Aracena, Romain Grégoire, Alain Pacis, Tomi Pastinen, Luis B. Barreiro, Yann Joly, David Bujold, Guillaume Bourque |
Bioinform. | 10 |
| 2023 | RNAget: an API to securely retrieve RNA quantificationsabstractSUMMARY: Large-scale sharing of genomic quantification data requires standardized access interfaces. In this Global Alliance for Genomics and Health project, we developed RNAget, an API for secure access to genomic quantification data in matrix form. RNAget provides for slicing matrices to extract desired subsets of data and is applicable to all expression matrix-format data, including RNA sequencing and microarrays. Further, it generalizes to quantification matrices of other sequence-based genomics such as ATAC-seq and ChIP-seq. AVAILABILITY AND IMPLEMENTATION: https://ga4gh-rnaseq.github.io/schema/docs/index.html. Sean Upchurch, Emilio Palumbo, Jeremy Adams, David Bujold, Guillaume Bourque, Jared Nedzel, Keenan Graham, Meenakshi S. Kagda, Pedro Assis, Benjamin C. Hitz, Emilio Righi, Roderic Guigó, Barbara J. Wold, Alvis Brazma, Julia Burchard, Joe Capka, Michael Cherry, Laura Clarke, Brian Craft, Manolis Dermitzakis, Mark Diekhans, John Dursi, Michael Sean Fitzsimons, Zac Flaming, Romina Garrido, Alfred Gil, Paul Godden, Matt Green, Mitch Guttman, Brian Haas, Max Haeussler, Sten Linnarsson, Adam Lipski, Simonne Longerich, David R. Lougheed, Jonathan Manning, John C. Marioni, Christopher Meyer, Stephen B. Montgomery, Alyssa Morrow, Alfonso Muñoz-Pomer Fuentes, Jared L. Nedzel, Kevin Osborn, Francis Ouellette, Irene Papatheodorou, Dmitri D. Pervouchine, Arun K. Ramani, Jordi Rambla De Argila, Bashir Sadjad, David Steinberg, Jeremiah Talkar, Timothy Tickle, Kathy Tzeng, Saman Vaisipour, Sean Watford, Barbara Wold |
Bioinform. | 5 |
| 2020 | Constrained Dynamic Programming and Supervised Penalty Learning Algorithms for Peak Detection in Genomic DataabstractPeak detection in genomic data involves segmenting counts of DNA sequence reads aligned to different locations of a chromosome. The goal is to detect peaks with higher counts, and filter out background noise with lower counts. Most existing algorithms for this problem are unsupervised heuristics tailored to patterns in specific data types. We propose a supervised framework for this problem, using optimal changepoint detection models with learned penalty functions. We propose the first dynamic programming algorithm that is guaranteed to compute the optimal solution to changepoint detection problems with constraints between adjacent segment mean parameters. Implementing this algorithm requires the choice of penalty parameter that determines the number of segments that are estimated. We show how the supervised learning ideas of Rigaill et al. (2013) can be used to choose this penalty. We compare the resulting implementation of our algorithm to several baselines in a benchmark of labeled ChIP-seq data sets with two different patterns (broad H3K36me3 data and sharp H3K4me3 data). Whereas baseline unsupervised methods only provide accurate peak detection for a single pattern, our supervised method achieves state-of-the-art accuracy in all data sets. The log-linear timings of our proposed dynamic programming algorithm make it scalable to the large genomic data sets that are now common. Our implementation is available in the PeakSegOptimal R package on CRAN. Toby Hocking, Guillem Rigaill, Paul Fearnhead, Guillaume Bourque |
J. Mach. Learn. Res. | 4 |
| 2019 | The use of hyperspectral remote sensing to detect PCB contaminated soils in the 0.35 to 12 micron spectral rangeabstractLaboratory spectral measurements were conducted to evaluate the potential of hyperspectral technologies, both reflective and emissive, to detect Polychlorinated Biphenyl (PCB) in contaminated soils. Soil sample standards of silt, clay, sand, and mixed textures were contaminated with Polychlorinated Biphenyl (PCB) oil with concentrations varying between 0% and 10% (100 000 ppm) of PCB. An ASD (0.35 to 2.5 microns) and a FSR (1.6 to 12 microns) spectrometers were used to measure the reflectance of pure and contaminated soil. Two spectral regions, one in the SWIR and one in the MWIR were significantly correlated to the soil PCB concentrations. It was shown that 5W30 engine oil and PCB absorb similarly in the two spectral regions. Sand and clay soils did not respond the same way as the other soil type. The MWIR CH absorption band between 3330 and 3630 was found the most promising to detect PCB contaminated soils using hyperspectral remote sensing, Josée Lévesque, Eldon Puckrin, Luc Levert, Guillaume Bourque |
IGARSS | 4 |
| 2019 | eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic dataabstractSUMMARY: The Illumina Infinium EPIC BeadChip is a new high-throughput array for DNA methylation analysis, extending the earlier 450k array by over 400 000 new sites. Previously, a method named eFORGE was developed to provide insights into cell type-specific and cell-composition effects for 450k data. Here, we present a significantly updated and improved version of eFORGE that can analyze both EPIC and 450k array data. New features include analysis of chromatin states, transcription factor motifs and DNase I footprints, providing tools for epigenome-wide association study interpretation and epigenome editing. AVAILABILITY AND IMPLEMENTATION: eFORGE v2.0 is implemented as a web tool available from https://eforge.altiusinstitute.org and https://eforge-tf.altiusinstitute.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Charles E. Breeze, Alex P. Reynolds, Jenny van Dongen, Ian Dunham, John Lazar, Shane J. Neph, Jeff Vierstra, Guillaume Bourque, Andrew E. Teschendorff, John A. Stamatoyannopoulos, Stephan Beck 0002 |
Bioinform. | 8 |
| 2019 | The epiGenomic Efficient Correlator (epiGeEC) tool allows fast comparison of user datasets with thousands of public epigenomic datasetsabstractSUMMARY: In recent years, major initiatives such as the International Human Epigenome Consortium have generated thousands of high-quality genome-wide datasets for a large variety of assays and cell types. This data can be used as a reference to assess whether the signal from a user-provided dataset corresponds to its expected experiment, as well as to help reveal unexpected biological associations. We have developed the epiGenomic Efficient Correlator (epiGeEC) tool to enable genome-wide comparisons of very large numbers of datasets. A public Galaxy implementation of epiGeEC allows comparison of user datasets with thousands of public datasets in a few minutes. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://bitbucket.org/labjacquespe/epigeec and the Galaxy implementation at http://epigeec.genap.ca. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jonathan Laperle, Simon Hébert-Deschamps, Joanny Raby, David A. de Lima Morais, Michel Barrette, David Bujold, Charlotte Bastin, Marc-Antoine Robert, Jean-François Nadeau, Marie Harel, Alexei Nordell-Markovits, Alain Veilleux, Guillaume Bourque, Pierre-Étienne Jacques |
Bioinform. | 13 |
| 2017 | Optimizing ChIP-seq peak detectors using visual labels and supervised machine learningabstractMotivation: Many peak detection algorithms have been proposed for ChIP-seq data analysis, but it is not obvious which algorithm and what parameters are optimal for any given dataset. In contrast, regions with and without obvious peaks can be easily labeled by visual inspection of aligned read counts in a genome browser. We propose a supervised machine learning approach for ChIP-seq data analysis, using labels that encode qualitative judgments about which genomic regions contain or do not contain peaks. The main idea is to manually label a small subset of the genome, and then learn a model that makes consistent peak predictions on the rest of the genome. Results: We created 7 new histone mark datasets with 12 826 visually determined labels, and analyzed 3 existing transcription factor datasets. We observed that default peak detection parameters yield high false positive rates, which can be reduced by learning parameters using a relatively small training set of labeled data from the same experiment type. We also observed that labels from different people are highly consistent. Overall, these data indicate that our supervised labeling method is useful for quantitatively training and testing peak detection algorithms. Availability and Implementation: Labeled histone mark data http://cbio.ensmp.fr/~thocking/chip-seq-chunk-db/ , R package to compute the label error of predicted peaks https://github.com/tdhock/PeakError. Contacts: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Toby Hocking, Patricia Goerner-Potvin, Andreanne Morin, Xiaojian Shao, Tomi Pastinen, Guillaume Bourque |
Bioinform. | 6 |
| 2015 | PeakSeg: constrained optimal segmentation and supervised penalty learning for peak detection in count dataabstractPeak detection is a central problem in genomic data analysis, and current algorithms for this task are unsupervised and mostly effective for a single data type and pattern (e.g. H3K4me3 data with a sharp peak pattern). We propose PeakSeg, a new constrained maximum likelihood segmentation model for peak detection with an efficient inference algorithm: constrained dynamic programming. We investigate unsupervised and supervised learning of penalties for the critical model selection problem. We show that the supervised method has state-of-the-art peak detection across all data sets in a benchmark that includes both sharp H3K4me3 and broad H3K36me3 patterns. Toby Hocking, Guillem Rigaill, Guillaume Bourque |
ICML | 3 |
| 2010 | Prediction of low coverage prone regions for Illumina sequencing projects using a support vector machineabstractApplications of next-generation sequencing technologies have the potential to bring revolutionary changes to medicine and biology. However, coverage bias can pose a challenge to short read data analysis tools, which rely on high coverage. To address this issue we have developed a support vector machine (SVM) based method for predicting low coverage prone (LCP) regions on a given genome. The developed SVM-based prediction of LCP regions on a given genome can assist data processing procedures based on Illumina sequencing technology, such as de novo sequencing and transcriptome analysis. Zejun Zheng, Bertil Schmidt, Guillaume Bourque |
BIBM | 3 |
| 2010 | webMGR: an online tool for the multiple genome rearrangement problemabstractSUMMARY: The algorithm MGR enables the reconstruction of rearrangement phylogenies based on gene or synteny block order in multiple genomes. Although MGR has been successfully applied to study the evolution of different sets of species, its utilization has been hampered by the prohibitive running time for some applications. In the current work, we have designed new heuristics that significantly speed up the tool without compromising its accuracy. Moreover, we have developed a web server (webMGR) that includes elaborate web output to facilitate navigation through the results. AVAILABILITY: webMGR can be accessed via http://www.gis.a-star.edu.sg/~bourque. The source code of the improved standalone version of MGR is also freely available from the web site. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chi Ho Lin, Sean Harry Lowcay, Atif Shahab, Guillaume Bourque |
Bioinform. | 5 |