Guy De Pauw

dblp:p/GuyDePauw · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Artificial intelligence
2 papers
Representation and self-supervised learning · 88% Multi-agent systems · 12%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › predictive modeling › classification › nearest neighbor classification
distance-based classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Data mining › predictive modeling › classification
multi-label classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Data mining › text mining
text classification
1.012026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding
0.312026
One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification · IEEE Trans. Knowl. Data Eng. 2026

Methods — techniques the papers use, named apart from their topics

threshold optimization · 2.0sentence encoder · 2.0evolutionary computing · 0.1agent-based modeling · 0.1
YearPublicationVenuePosition
2026 One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text Classification
abstract
Distance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast inference and adaptability to expanding label sets, as opposed to zero-shot, few-shot, and fine-tuned neural networks that require re-training in such cases. In multi-label distance-based classification and information retrieval algorithms, thresholds are required to determine whether a text instance is “similar” to a label or query. Similarity between a text and label is determined in a dense embedding space, usually generated by state-of-the-art sentence encoders. Multi-label classification complicates matters, as a text instance can have multiple true labels, unlike in multi-class or binary classification, where each instance is assigned only one label. We expand upon previous literature on this underexplored topic by thoroughly examining and evaluating the ability of sentence encoders to perform distance-based classification. First, we perform an exploratory study to verify whether the semantic relationships between texts and labels vary across models, datasets, and label sets by conducting experiments on a diverse collection of realistic multi-label text classification (MLTC) datasets. We find that similarity distributions show statistically significant differences across models, datasets and even label sets. We propose a novel method for optimizing label-specific thresholds using a validation set. Our label-specific thresholding method achieves an average improvement of 46% over normalized 0.5 thresholding and outperforms uniform thresholding approaches from previous work by an average of 14%. Additionally, the method demonstrates strong performance even with limited labeled examples.
Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans
IEEE Trans. Knowl. Data Eng.3
2016 Multimodular Text Normalization of Dutch User-Generated Content
abstract
As social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization.
Sarah Schulz, Guy De Pauw, Orphée De Clercq, Bart Desmet, Véronique Hoste, Walter Daelemans, Lieve Macken
ACM Trans. Intell. Syst. Technol.2
2013 Self-taught assistive vocal interfaces: an overview of the ALADIN project
abstract
This paper gives an overview of research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing ap-proaches, the vocal interface is trained by the end-user himself, which means it can be used with any vocabulary and grammar, and that it is maximally adapted to the — possibly dysarthric — speech of the user. This paper describes the overall learn-ing framework, the user-centred design and evaluation aspects, database collection and approaches taken to combat problems such as noise and erroneous input. Index Terms: vocal user interface, user-centred design, self-taught learning, speech database, dysarthric speech
Jort F. Gemmeke, Bart Ons, Netsanet M. Tessema, Hugo Van hamme, Janneke van de Loo, Guy De Pauw, Walter Daelemans, Jonathan Huyghe, Jan Derboven, Lode Vuegen, Bert Van Den Broeck, Peter Karsmakers, Bart Vanrumste
INTERSPEECH6
2012 A Self-Learning Assistive Vocal Interface Based on Vocabulary Learning and Grammar Induction
abstract
This paper introduces research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing approaches, the vocal interface is self-learning which means it can be used with any language, dialect, vocabulary and grammar. The paper describes the overall learning framework, and the two components that will provide vocabulary learning and grammar induction. In addition, the paper describes encouraging results of early implementations of these vocabulary and grammar learning components, applied to recorded sessions of a vocally guided card game, patience. Index Terms: language acquisition, word finding, grammar induction, non-negative matrix factorization, concept tagging
Jort F. Gemmeke, Janneke van de Loo, Guy De Pauw, Joris Driesen, Hugo Van hamme, Walter Daelemans
INTERSPEECH3
2012 The Netlog Corpus. A Resource for the Study of Flemish Dutch Internet Language
Mike Kestemont, Claudia Peersman, Benny De Decker, Guy De Pauw, Kim Luyckx, Roser Morante, Frederik Vaassen, Janneke van de Loo, Walter Daelemans
LREC4
2007 Bootstrapping morphological analysis of gĩkũyũ using unsupervised maximum entropy learning
abstract
This paper describes a proof-of-the-principle experiment in which maximum entropy learning is used for the automatic induction of shallow morphological features for the resourcescarce Bantu language of Gĩkũyũ. This novel approach circumvents the limitations of typical unsupervised morphological induction methods that employ minimum-edit distance metrics to establish morphological similarity between words. The experimental results show that the unsupervised maximum entropy learning approach compares favorably to those of the established AutoMorphology method. Index Terms: unsupervised learning, morphology, Bantu languages
Guy De Pauw, Peter Waiganjo Wagacha
INTERSPEECH1
2006 KNACK-2002: a Richly Annotated Corpus of Dutch Written Text
Véronique Hoste, Guy De Pauw
LREC2
2006 A mixed word / morphological approach for extending CELEX for high coverage on contemporary large corpora
Joris Vaneyghen, Guy De Pauw, Dirk Van Compernolle, Walter Daelemans
LREC2
2006 A Grapheme-Based Approach for Accent Restoration in Gikuyu
Peter Waiganjo Wagacha, Guy De Pauw, Pauline Githinji
LREC2
2004 Evaluation and Adaptation of the Celex Dutch Morphological Database
Tom Laureys, Guy De Pauw, Hugo Van hamme, Walter Daelemans, Dirk Van Compernolle
LREC2
2003 Evolutionary Computing as a Tool for Grammar Development
Guy De Pauw
GECCO1
2003 GRAEL: an agent-based evolutionary computing approach for natural language grammar development
Guy De Pauw
IJCAI1
2000 Aspects of Pattern-matching in Data-Oriented Parsing
Guy De Pauw
COLING1