Gaetan De Waele

dblp:311/2909 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
0000-0003-0367-9699ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
2 papers
Bioinformatics and computational biology · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
gene regulation
1.012026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
1.012026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology › epigenomics
DNA methylation
0.612022
CpG Transformer for imputation of single-cell methylomes · Bioinform. 2022
Bioinformatics and computational biology
epigenomics
0.612022
CpG Transformer for imputation of single-cell methylomes · Bioinform. 2022
Bioinformatics and computational biology › genomics
machine learning for genomics
0.312026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology
negative sampling
0.312026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026

Methods — techniques the papers use, named apart from their topics

deep learning · 1.0ChIP-seq · 1.0ATAC-seq · 1.0transformer · 0.6transfer learning · 0.6sliding window self-attention · 0.6axial attention · 0.6
YearPublicationVenuePosition
2026 How negative sampling shapes the performance of transcription factor binding site prediction models
abstract
MOTIVATION: Transcription factors (TFs) are key players in gene regulation and development, where they activate and repress gene expression through DNA binding. Predicting transcription factor binding sites (TFBSs) has long been an active area of research, with many deep learning methods developed to tackle this problem. These models are often trained on TF ChIP-seq data, which is generally seen as only providing positive samples. The choice of datasets and negative sampling techniques is a critical yet often overlooked aspect of this work. RESULTS: In this study, we investigate the impact of different negative sampling techniques on TFBS prediction performance. We create high-quality test datasets based on ChIP-seq and ATAC-seq data, where true negatives can be identified as positions that are accessible but not bound by the TF in question. We then train models using various negative sampling techniques, including genomic sampling, shuffling, dinucleotide shuffling, neighborhood sampling, and cell line specific sampling, simulating cases where matching ATAC-seq data is not available. Our results show that, generally, metrics calculated on training datasets give inflated performance scores. Of the tested techniques, genomic sampling of negatives based on similarity to the positives performed by far the best, although still not reaching the performance of baseline models trained on high-quality datasets. Models trained on dinucleotide shuffled negatives performed poorly, despite being a common practice in the field. Our findings highlight the importance of carefully selecting negative sampling techniques for TFBS prediction, as they can significantly impact model performance and the interpretation of results. AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/NatanTourne/TFBS-negatives (DOI: 10.5281/zenodo.18007567).
Natan Tourne, Gaetan De Waele, Vanessa Vermeirssen, Willem Waegeman
Bioinform.2
2022 CpG Transformer for imputation of single-cell methylomes
abstract
MOTIVATION: The adoption of current single-cell DNA methylation sequencing protocols is hindered by incomplete coverage, outlining the need for effective imputation techniques. The task of imputing single-cell (methylation) data requires models to build an understanding of underlying biological processes. RESULTS: We adapt the transformer neural network architecture to operate on methylation matrices through combining axial attention with sliding window self-attention. The obtained CpG Transformer displays state-of-the-art performances on a wide range of scBS-seq and scRRBS-seq datasets. Furthermore, we demonstrate the interpretability of CpG Transformer and illustrate its rapid transfer learning properties, allowing practitioners to train models on new datasets with a limited computational and time budget. AVAILABILITY AND IMPLEMENTATION: CpG Transformer is freely available at https://github.com/gdewael/cpg-transformer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gaetan De Waele, Jim Clauwaert, Gerben Menschaert, Willem Waegeman
Bioinform.1