Jens Van Nooten

dblp:315/9377 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0002-0165-5709ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Artificial intelligence
2 papers
Language models and text generation · 61% Representation and self-supervised learning · 21% Trustworthy machine learning · 18%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Data mining › predictive modeling › classification
multi-label classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Data mining › text mining
text classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.312026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026

Methods — techniques the papers use, named apart from their topics

threshold optimization · 2.0sentence encoder · 2.0benchmark analysis · 0.9
YearPublicationVenuePosition
2026 Push and Pull: Training Sentence Encoders with Contrastive Losses for Distance-Based Multi-Label Text Classification
Jens Van Nooten, Andriy Kosar
LREC1
2026 One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification
abstract
Distance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast inference and adaptability to expanding label sets, as opposed to zero-shot, few-shot, and fine-tuned neural networks that require re-training in such cases. In multi-label distance-based classification and information retrieval algorithms, thresholds are required to determine whether a text instance is “similar” to a label or query. Similarity between a text and label is determined in a dense embedding space, usually generated by state-of-the-art sentence encoders. Multi-label classification complicates matters, as a text instance can have multiple true labels, unlike in multi-class or binary classification, where each instance is assigned only one label. We expand upon previous literature on this underexplored topic by thoroughly examining and evaluating the ability of sentence encoders to perform distance-based classification. First, we perform an exploratory study to verify whether the semantic relationships between texts and labels vary across models, datasets, and label sets by conducting experiments on a diverse collection of realistic multi-label text classification (MLTC) datasets. We find that similarity distributions show statistically significant differences across models, datasets and even label sets. We propose a novel method for optimizing label-specific thresholds using a validation set. Our label-specific thresholding method achieves an average improvement of 46% over normalized 0.5 thresholding and outperforms uniform thresholding approaches from previous work by an average of 14%. Additionally, the method demonstrates strong performance even with limited labeled examples.
Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans
IEEE Trans. Knowl. Data Eng.1
2025 Jump To Hyperspace: Comparing Euclidean and Hyperbolic Loss Functions for Hierarchical Multi-Label Text Classification
abstract
Hierarchical Multi-Label Text Classification (HMTC) is a challenging machine learning task where multiple labels from a hierarchically organized label set are assigned to a single text. In this study, we examine the effectiveness of Euclidean and hyperbolic loss functions to improve the performance of BERT models on HMTC, which very few previous studies have adopted. We critically evaluate label-aware losses as well as contrastive losses in the Euclidean and hyperbolic space, demonstrating that hyperbolic loss functions perform comparably with non-hyperbolic loss functions on four commonly used HMTC datasets in most scenarios. While hyperbolic label-aware losses perform the best on low-level labels, the overall consistency and micro-averaged performance is compromised. Additionally, we find that our contrastive losses are less effective for HMTC when deployed in the hyperbolic space than non-hyperbolic counterparts. Our research highlights that with the right metrics and training objectives, hyperbolic space does not provide any additional benefits compared to Euclidean space for HMTC, thereby prompting a reevaluation of how different geometric spaces are used in other AI applications.
Jens Van Nooten, Walter Daelemans
COLING1
2025 In Benchmarks We Trust ... Or Not?
abstract
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi 0002, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
EMNLP3
2022 CoNTACT: A Dutch COVID-19 Adapted BERT for Vaccine Hesitancy and Argumentation Detection
abstract
We present CoNTACT: a Dutch language model adapted to the domain of COVID-19 tweets. The model was developed by continuing the pre-training phase of RobBERT (Delobelle et al., 2020) by using 2.8M Dutch COVID-19 related tweets posted in 2021. In order to test the performance of the model and compare it to RobBERT, the two models were tested on two tasks: (1) binary vaccine hesitancy detection and (2) detection of arguments for vaccine hesitancy. For both tasks, not only Twitter but also Facebook data was used to show cross-genre performance. In our experiments, CoNTACT showed statistically significant gains over RobBERT in all experiments for task 1. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions.
Jens Lemmens, Jens Van Nooten, Tim Kreutz, Walter Daelemans
COLING2