VLDB 2026 Research / reviewers in the wild / expert
Guy De Pauw
dblp:p/GuyDePauw
· DBLP profile ↗
13ranked-venue papers
4as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% | |
| Artificial intelligence
2 papers |
Representation and self-supervised learning · 88% Multi-agent systems · 12% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Data mining › predictive modeling › classification
multi-label classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Data mining › text mining
text classification |
1.0 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.3 | 1 | 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026 |
Methods — techniques the papers use, named apart from their topics
threshold optimization · 2.0sentence encoder · 2.0evolutionary computing · 0.1agent-based modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text ClassificationabstractDistance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast inference and adaptability to expanding label sets, as opposed to zero-shot, few-shot, and fine-tuned neural networks that require re-training in such cases. In multi-label distance-based classification and information retrieval algorithms, thresholds are required to determine whether a text instance is “similar” to a label or query. Similarity between a text and label is determined in a dense embedding space, usually generated by state-of-the-art sentence encoders. Multi-label classification complicates matters, as a text instance can have multiple true labels, unlike in multi-class or binary classification, where each instance is assigned only one label. We expand upon previous literature on this underexplored topic by thoroughly examining and evaluating the ability of sentence encoders to perform distance-based classification. First, we perform an exploratory study to verify whether the semantic relationships between texts and labels vary across models, datasets, and label sets by conducting experiments on a diverse collection of realistic multi-label text classification (MLTC) datasets. We find that similarity distributions show statistically significant differences across models, datasets and even label sets. We propose a novel method for optimizing label-specific thresholds using a validation set. Our label-specific thresholding method achieves an average improvement of 46% over normalized 0.5 thresholding and outperforms uniform thresholding approaches from previous work by an average of 14%. Additionally, the method demonstrates strong performance even with limited labeled examples. Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Multimodular Text Normalization of Dutch User-Generated ContentabstractAs social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization. Sarah Schulz, Guy De Pauw, Orphée De Clercq, Bart Desmet, Véronique Hoste, Walter Daelemans, Lieve Macken |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Self-taught assistive vocal interfaces: an overview of the ALADIN projectabstractThis paper gives an overview of research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing ap-proaches, the vocal interface is trained by the end-user himself, which means it can be used with any vocabulary and grammar, and that it is maximally adapted to the — possibly dysarthric — speech of the user. This paper describes the overall learn-ing framework, the user-centred design and evaluation aspects, database collection and approaches taken to combat problems such as noise and erroneous input. Index Terms: vocal user interface, user-centred design, self-taught learning, speech database, dysarthric speech Jort F. Gemmeke, Bart Ons, Netsanet M. Tessema, Hugo Van hamme, Janneke van de Loo, Guy De Pauw, Walter Daelemans, Jonathan Huyghe, Jan Derboven, Lode Vuegen, Bert Van Den Broeck, Peter Karsmakers, Bart Vanrumste |
INTERSPEECH | 6 |
| 2012 | A Self-Learning Assistive Vocal Interface Based on Vocabulary Learning and Grammar InductionabstractThis paper introduces research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing approaches, the vocal interface is self-learning which means it can be used with any language, dialect, vocabulary and grammar. The paper describes the overall learning framework, and the two components that will provide vocabulary learning and grammar induction. In addition, the paper describes encouraging results of early implementations of these vocabulary and grammar learning components, applied to recorded sessions of a vocally guided card game, patience. Index Terms: language acquisition, word finding, grammar induction, non-negative matrix factorization, concept tagging Jort F. Gemmeke, Janneke van de Loo, Guy De Pauw, Joris Driesen, Hugo Van hamme, Walter Daelemans |
INTERSPEECH | 3 |
| 2012 | The Netlog Corpus. A Resource for the Study of Flemish Dutch Internet Language
Mike Kestemont, Claudia Peersman, Benny De Decker, Guy De Pauw, Kim Luyckx, Roser Morante, Frederik Vaassen, Janneke van de Loo, Walter Daelemans |
LREC | 4 |
| 2007 | Bootstrapping morphological analysis of gĩkũyũ using unsupervised maximum entropy learningabstractThis paper describes a proof-of-the-principle experiment in which maximum entropy learning is used for the automatic induction of shallow morphological features for the resourcescarce Bantu language of Gĩkũyũ. This novel approach circumvents the limitations of typical unsupervised morphological induction methods that employ minimum-edit distance metrics to establish morphological similarity between words. The experimental results show that the unsupervised maximum entropy learning approach compares favorably to those of the established AutoMorphology method. Index Terms: unsupervised learning, morphology, Bantu languages Guy De Pauw, Peter Waiganjo Wagacha |
INTERSPEECH | 1 |
| 2006 | KNACK-2002: a Richly Annotated Corpus of Dutch Written Text
Véronique Hoste, Guy De Pauw |
LREC | 2 |
| 2006 | A mixed word / morphological approach for extending CELEX for high coverage on contemporary large corpora
Joris Vaneyghen, Guy De Pauw, Dirk Van Compernolle, Walter Daelemans |
LREC | 2 |
| 2006 | A Grapheme-Based Approach for Accent Restoration in Gikuyu
Peter Waiganjo Wagacha, Guy De Pauw, Pauline Githinji |
LREC | 2 |
| 2004 | Evaluation and Adaptation of the Celex Dutch Morphological Database
Tom Laureys, Guy De Pauw, Hugo Van hamme, Walter Daelemans, Dirk Van Compernolle |
LREC | 2 |
| 2003 | Evolutionary Computing as a Tool for Grammar Development
Guy De Pauw |
GECCO | 1 |
| 2003 | GRAEL: an agent-based evolutionary computing approach for natural language grammar development
Guy De Pauw |
IJCAI | 1 |
| 2000 | Aspects of Pattern-matching in Data-Oriented Parsing
Guy De Pauw |
COLING | 1 |