VLDB 2026 Research / reviewers in the wild / expert
Jens Van Nooten
dblp:315/9377
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-0165-5709ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Artificial intelligence
2 papers |
Language models and text generation · 61% Representation and self-supervised learning · 21% Trustworthy machine learning · 18% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Data mining › predictive modeling › classification
multi-label classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Data mining › text mining
text classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.3 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Methods — techniques the papers use, named apart from their topics
threshold optimization · 2.0sentence encoder · 2.0benchmark analysis · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Push and Pull: Training Sentence Encoders with Contrastive Losses for Distance-Based Multi-Label Text Classification
Jens Van Nooten, Andriy Kosar |
LREC | 1 |
| 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text ClassificationabstractDistance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast inference and adaptability to expanding label sets, as opposed to zero-shot, few-shot, and fine-tuned neural networks that require re-training in such cases. In multi-label distance-based classification and information retrieval algorithms, thresholds are required to determine whether a text instance is “similar” to a label or query. Similarity between a text and label is determined in a dense embedding space, usually generated by state-of-the-art sentence encoders. Multi-label classification complicates matters, as a text instance can have multiple true labels, unlike in multi-class or binary classification, where each instance is assigned only one label. We expand upon previous literature on this underexplored topic by thoroughly examining and evaluating the ability of sentence encoders to perform distance-based classification. First, we perform an exploratory study to verify whether the semantic relationships between texts and labels vary across models, datasets, and label sets by conducting experiments on a diverse collection of realistic multi-label text classification (MLTC) datasets. We find that similarity distributions show statistically significant differences across models, datasets and even label sets. We propose a novel method for optimizing label-specific thresholds using a validation set. Our label-specific thresholding method achieves an average improvement of 46% over normalized 0.5 thresholding and outperforms uniform thresholding approaches from previous work by an average of 14%. Additionally, the method demonstrates strong performance even with limited labeled examples. Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Jump To Hyperspace: Comparing Euclidean and Hyperbolic Loss Functions for Hierarchical Multi-Label Text ClassificationabstractHierarchical Multi-Label Text Classification (HMTC) is a challenging machine learning task where multiple labels from a hierarchically organized label set are assigned to a single text. In this study, we examine the effectiveness of Euclidean and hyperbolic loss functions to improve the performance of BERT models on HMTC, which very few previous studies have adopted. We critically evaluate label-aware losses as well as contrastive losses in the Euclidean and hyperbolic space, demonstrating that hyperbolic loss functions perform comparably with non-hyperbolic loss functions on four commonly used HMTC datasets in most scenarios. While hyperbolic label-aware losses perform the best on low-level labels, the overall consistency and micro-averaged performance is compromised. Additionally, we find that our contrastive losses are less effective for HMTC when deployed in the hyperbolic space than non-hyperbolic counterparts. Our research highlights that with the right metrics and training objectives, hyperbolic space does not provide any additional benefits compared to Euclidean space for HMTC, thereby prompting a reevaluation of how different geometric spaces are used in other AI applications. Jens Van Nooten, Walter Daelemans |
COLING | 1 |
| 2025 | In Benchmarks We Trust ... Or Not?abstractIne Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi 0002, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans |
EMNLP | 3 |
| 2022 | CoNTACT: A Dutch COVID-19 Adapted BERT for Vaccine Hesitancy and Argumentation DetectionabstractWe present CoNTACT: a Dutch language model adapted to the domain of COVID-19 tweets. The model was developed by continuing the pre-training phase of RobBERT (Delobelle et al., 2020) by using 2.8M Dutch COVID-19 related tweets posted in 2021. In order to test the performance of the model and compare it to RobBERT, the two models were tested on two tasks: (1) binary vaccine hesitancy detection and (2) detection of arguments for vaccine hesitancy. For both tasks, not only Twitter but also Facebook data was used to show cross-genre performance. In our experiments, CoNTACT showed statistically significant gains over RobBERT in all experiments for task 1. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions. Jens Lemmens, Jens Van Nooten, Tim Kreutz, Walter Daelemans |
COLING | 2 |