Natan Tourne

dblp:428/3612 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
0009-0000-5046-010XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
gene regulation
1.012026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology › gene regulation
transcription factor binding site prediction
1.012026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology › genomics
machine learning for genomics
0.312026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026
Bioinformatics and computational biology
negative sampling
0.312026
How negative sampling shapes the performance of transcription factor binding site prediction models · Bioinform. 2026

Methods — techniques the papers use, named apart from their topics

deep learning · 1.0ChIP-seq · 1.0ATAC-seq · 1.0
YearPublicationVenuePosition
2026 How negative sampling shapes the performance of transcription factor binding site prediction models
abstract
MOTIVATION: Transcription factors (TFs) are key players in gene regulation and development, where they activate and repress gene expression through DNA binding. Predicting transcription factor binding sites (TFBSs) has long been an active area of research, with many deep learning methods developed to tackle this problem. These models are often trained on TF ChIP-seq data, which is generally seen as only providing positive samples. The choice of datasets and negative sampling techniques is a critical yet often overlooked aspect of this work. RESULTS: In this study, we investigate the impact of different negative sampling techniques on TFBS prediction performance. We create high-quality test datasets based on ChIP-seq and ATAC-seq data, where true negatives can be identified as positions that are accessible but not bound by the TF in question. We then train models using various negative sampling techniques, including genomic sampling, shuffling, dinucleotide shuffling, neighborhood sampling, and cell line specific sampling, simulating cases where matching ATAC-seq data is not available. Our results show that, generally, metrics calculated on training datasets give inflated performance scores. Of the tested techniques, genomic sampling of negatives based on similarity to the positives performed by far the best, although still not reaching the performance of baseline models trained on high-quality datasets. Models trained on dinucleotide shuffled negatives performed poorly, despite being a common practice in the field. Our findings highlight the importance of carefully selecting negative sampling techniques for TFBS prediction, as they can significantly impact model performance and the interpretation of results. AVAILABILITY AND IMPLEMENTATION: The code used in this study is available at https://github.com/NatanTourne/TFBS-negatives (DOI: 10.5281/zenodo.18007567).
Natan Tourne, Gaetan De Waele, Vanessa Vermeirssen, Willem Waegeman
Bioinform.1