EDBT 2026 Demo / reviewers in the wild / expert
Walter Daelemans
dblp:d/WalterDaelemans
· DBLP profile ↗
116ranked-venue papers
21as first author
14since 2021 · last 2026
0000-0002-9832-7890ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 100 · 17 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12Databases, data management, data science and information retrieval · 9 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One Size Does Not Fit All: Exploring Variable Thresholds for Distance-Based Multi-Label Text ClassificationabstractDistance-based unsupervised text classification is a method within text classification that leverages the semantic similarity between a label and a text to determine label relevance. This method provides numerous benefits, including fast inference and adaptability to expanding label sets, as opposed to zero-shot, few-shot, and fine-tuned neural networks that require re-training in such cases. In multi-label distance-based classification and information retrieval algorithms, thresholds are required to determine whether a text instance is “similar” to a label or query. Similarity between a text and label is determined in a dense embedding space, usually generated by state-of-the-art sentence encoders. Multi-label classification complicates matters, as a text instance can have multiple true labels, unlike in multi-class or binary classification, where each instance is assigned only one label. We expand upon previous literature on this underexplored topic by thoroughly examining and evaluating the ability of sentence encoders to perform distance-based classification. First, we perform an exploratory study to verify whether the semantic relationships between texts and labels vary across models, datasets, and label sets by conducting experiments on a diverse collection of realistic multi-label text classification (MLTC) datasets. We find that similarity distributions show statistically significant differences across models, datasets and even label sets. We propose a novel method for optimizing label-specific thresholds using a validation set. Our label-specific thresholding method achieves an average improvement of 46% over normalized 0.5 thresholding and outperforms uniform thresholding approaches from previous work by an average of 14%. Additionally, the method demonstrates strong performance even with limited labeled examples. Jens Van Nooten, Andriy Kosar, Guy De Pauw, Walter Daelemans |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | Jump To Hyperspace: Comparing Euclidean and Hyperbolic Loss Functions for Hierarchical Multi-Label Text ClassificationabstractHierarchical Multi-Label Text Classification (HMTC) is a challenging machine learning task where multiple labels from a hierarchically organized label set are assigned to a single text. In this study, we examine the effectiveness of Euclidean and hyperbolic loss functions to improve the performance of BERT models on HMTC, which very few previous studies have adopted. We critically evaluate label-aware losses as well as contrastive losses in the Euclidean and hyperbolic space, demonstrating that hyperbolic loss functions perform comparably with non-hyperbolic loss functions on four commonly used HMTC datasets in most scenarios. While hyperbolic label-aware losses perform the best on low-level labels, the overall consistency and micro-averaged performance is compromised. Additionally, we find that our contrastive losses are less effective for HMTC when deployed in the hyperbolic space than non-hyperbolic counterparts. Our research highlights that with the right metrics and training objectives, hyperbolic space does not provide any additional benefits compared to Euclidean space for HMTC, thereby prompting a reevaluation of how different geometric spaces are used in other AI applications. Jens Van Nooten, Walter Daelemans |
COLING | 2 |
| 2025 | In Benchmarks We Trust ... Or Not?abstractIne Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi 0002, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans |
EMNLP | 10 |
| 2024 | Model Priming with Triplet Loss for Few-Shot Emotion Classification in TextabstractAutomatically detecting emotions in text data is a challenging task, especially when only little supervised training data is available. Therefore, we attempt to boost model performance in few-shot experiments by learning emotion label information. We explore triplet loss for emotion detection from text, and utilize this technique to cluster text representations that express the same emotion in the embedding space before learning the classification task. We show that this method, which we call emotion priming, outperforms baseline results, multi-task fine-tuning with cross-entropy loss, and an existing label infusion method that adds label words to the input sequence to alter the model’s attention weights. An analysis of the emotion class representations after priming shows that the observed performance gain can be attributed to the redistribution of the text representations. The results also indicate that this method is robust in datasets that contain many classes and few examples per class. In contrast to earlier work, we found that the label infusion method leads to a substantial performance decrease compared to the baseline model, especially for the datasets with more complex label schemes. Finally, we report results for zero-shot experiments with ChatGPT as a larger alternative to smaller fine-tuned language models, and show that it fails to produce accurate results, indicating the complexity of the studied task. Jens Lemmens, Walter Daelemans |
ECAI | 2 |
| 2022 | Open-Domain Dialog Evaluation Using Follow-Ups LikelihoodabstractAutomatic evaluation of open-domain dialogs remains an unsolved problem. Existing methods do not correlate strongly with human annotations. In this paper, we present a new automated evaluation method based on the use of follow-ups. We measure the probability that a language model will continue the conversation with a fixed set of follow-ups (e.g. not really relevant here, what are you trying to say?). When compared against twelve existing methods, our new evaluation achieves the highest correlation with human evaluations. Maxime De Bruyn, Ehsan Lotfi 0002, Jeska Buhmann, Walter Daelemans |
COLING | 4 |
| 2022 | Domain- and Task-Adaptation for VaccinChatNL, a Dutch COVID-19 FAQ Answering Corpus and Classification ModelabstractFAQs are important resources to find information. However, especially if a FAQ concerns many question-answer pairs, it can be a difficult and time-consuming job to find the answer you are looking for. A FAQ chatbot can ease this process by automatically retrieving the relevant answer to a user’s question. We present VaccinChatNL, a Dutch FAQ corpus on the topic of COVID-19 vaccination. Starting with 50 question-answer pairs we built VaccinChat, a FAQ chatbot, which we used to gather more user questions that were also annotated with the appropriate or new answer classes. This iterative process of gathering user questions, annotating them, and retraining the model with the increased data set led to a corpus that now contains 12,883 user questions divided over 181 answers. We provide the first publicly available Dutch FAQ answering data set of this size with large groups of semantically equivalent human-paraphrased questions. Furthermore, our study shows that before fine-tuning a classifier, continued pre-training of Dutch language models with task- and/or domain-specific data improves classification results. In addition, we show that large groups of semantically similar questions are important for obtaining well-performing intent classification models. Jeska Buhmann, Maxime De Bruyn, Ehsan Lotfi 0002, Walter Daelemans |
COLING | 4 |
| 2022 | CoNTACT: A Dutch COVID-19 Adapted BERT for Vaccine Hesitancy and Argumentation DetectionabstractWe present CoNTACT: a Dutch language model adapted to the domain of COVID-19 tweets. The model was developed by continuing the pre-training phase of RobBERT (Delobelle et al., 2020) by using 2.8M Dutch COVID-19 related tweets posted in 2021. In order to test the performance of the model and compare it to RobBERT, the two models were tested on two tasks: (1) binary vaccine hesitancy detection and (2) detection of arguments for vaccine hesitancy. For both tasks, not only Twitter but also Facebook data was used to show cross-genre performance. In our experiments, CoNTACT showed statistically significant gains over RobBERT in all experiments for task 1. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions. For task 2, we observed substantial improvements in virtually all classes in all experiments. An error analysis indicated that the domain adaptation yielded better representations of domain-specific terminology, causing CoNTACT to make more accurate classification decisions. Jens Lemmens, Jens Van Nooten, Tim Kreutz, Walter Daelemans |
COLING | 4 |
| 2022 | Cyberbullying Classifiers are Sensitive to Model-Agnostic PerturbationsabstractA limited amount of studies investigates the role of model-agnostic adversarial behavior in toxic content classification. As toxicity classifiers predominantly rely on lexical cues, (deliberately) creative and evolving language-use can be detrimental to the utility of current corpora and state-of-the-art models when they are deployed for content moderation. The less training data is available, the more vulnerable models might become. This study is, to our knowledge, the first to investigate the effect of adversarial behavior and augmentation for cyberbullying detection. We demonstrate that model-agnostic lexical substitutions significantly hurt classifier performance. Moreover, when these perturbed samples are used for augmentation, we show models become robust against word-level perturbations at a slight trade-off in overall task performance. Augmentations proposed in prior work on toxicity prove to be less effective. Our results underline the need for such evaluations in online harm areas with small corpora. Chris Emmery, Ákos Kádár, Grzegorz Chrupala, Walter Daelemans |
LREC | 4 |
| 2022 | Detecting Vaccine Skepticism on Twitter Using Heterogeneous Information Networks
Tim Kreutz, Walter Daelemans |
NLDB | 2 |
| 2022 | An Ensemble Approach for Dutch Cross-Domain Hate Speech Detection
Ilia Markov, Ine Gevers, Walter Daelemans |
NLDB | 3 |
| 2022 | EmoLabel: Semi-Automatic Methodology for Emotion Annotation of Social Media TextabstractThe exponential growth of the amount of subjective information on the Web 2.0. has caused an increasing interest from researchers willing to develop methods to extract emotion data from these new sources. One of the most important challenges in textual emotion detection is the gathering of data with emotion labels because of the subjectivity of assigning these labels. Basing on this rationale, the main objective of our research is to contribute to the resolution of this important challenge. This is tackled by proposing EmoLabel: a semi-automatic methodology based on pre-annotation, which consists of two main phases: (1) an automatic process to pre-annotate the unlabelled English sentences; and (2) a manual process of refinement where human annotators determine which is the dominant emotion. Our objective is to assess the influence of this automatic pre-annotation method on manual emotion annotation from two points of view: agreement and time needed for annotation. The evaluation performed demonstrates the benefits of pre-annotation processes since the results on annotation time show a gain of near 20 percent when the pre-annotation process is applied (Pre-ML) without reducing annotator performance. Moreover, the benefits of pre-annotation are higher in those contributors whose performance is low (inaccurate annotators). Lea Canales, Walter Daelemans, Ester Boldrini, Patricio Martínez-Barco |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Conceptual Grounding Constraints for Truly Robust Biomedical Name RepresentationsabstractEffective representation of biomedical names for downstream NLP tasks requires the encoding of both lexical as well as domain-specific semantic information.Ideally, the synonymy and semantic relatedness of names should be consistently reflected by their closeness in an embedding space.To achieve such robustness, prior research has considered multi-task objectives when training neural encoders.In this paper, we take a next step towards truly robust representations, which capture more domainspecific semantics while remaining universally applicable across different biomedical corpora and domains.To this end, we use conceptual grounding constraints which more effectively align encoded names to pretrained embeddings of their concept identifiers.These constraints are effective even when using a Deep Averaging Network, a simple feedforward encoding architecture that allows for scaling to large corpora while remaining sufficiently expressive.We empirically validate our approach using multiple tasks and benchmarks, which assess both literal synonymy as well as more general semantic relatedness.Our code is open-source and available at www.github. com/clips/conceptualgrounding. Pieter Fivez, Simon Suster, Walter Daelemans |
EACL | 3 |
| 2021 | Mapping probability word problems to executable representationsabstractWhile solving math word problems automatically has received considerable attention in the NLP community, few works have addressed probability word problems specifically.In this paper, we employ and analyse various neural models for answering such word problems.In a two-step approach, the problem text is first mapped to a formal representation in a declarative language using a sequence-to-sequence model, and then the resulting representation is executed using a probabilistic programming system to provide the answer.Our best performing model incorporates general-domain contextualised word representations that were finetuned using transfer learning on another in-domain dataset.We also apply end-to-end models to this task, which bring out the importance of the two-step approach in obtaining correct solutions to probability problems. Simon Suster, Pieter Fivez, Pietro Totis, Angelika Kimmig, Jesse Davis, Luc De Raedt, Walter Daelemans |
EMNLP (1) | 7 |
| 2021 | Multi-modal Label Retrieval for the Visual Arts: The Case of IconclassabstractIconclass is an iconographic classification system from the domain of cultural heritage which is used to annotate subjects represented in the visual arts.In this work, we investigate the feasibility of automatically assigning Iconclass codes to visual artworks using a cross-modal retrieval set-up.We explore the text and image branches of the cross-modal network.In addition, we describe a multi-modal architecture that can jointly capitalize on multiple feature sources: textual features, coming from the titles for these artworks (in multiple languages) and visual features, extracted from photographic reproductions of the artworks.We utilize Iconclass definitions in English as matching labels.We evaluate our approach on a publicly available dataset of artworks (containing English and Dutch titles).Our results demonstrate that, in isolation, textual features strongly outperform visual features, although visual features can still offer a useful complement to purely linguistic features.Moreover, we show the cross-lingual (Dutch-English) strategy to be on par with the monolingual approach (English-English), which opens important perspectives for applications of this approach beyond resource-rich languages. Nikolay Banar, Walter Daelemans, Mike Kestemont |
ICAART (1) | 2 |
| 2020 | A Deep Generative Approach to Native Language IdentificationabstractNative language identification (NLI) -identifying the native language (L1) of a person based on his/her writing in the second language (L2) -is useful for a variety of purposes, including marketing, security, and educational applications.From a traditional machine learning perspective, NLI is usually framed as a multi-class classification task, where numerous designed features are combined in order to achieve state-of-the-art results.We introduce a deep generative language modelling (LM) approach to NLI, which consists in fine-tuning a GPT-2 model separately on texts written by the authors with the same L1, and assigning a label to an unseen text based on the minimum LM loss with respect to one of these fine-tuned GPT-2 models.Our method outperforms traditional machine learning approaches and currently achieves the best results on the benchmark NLI datasets. Ehsan Lotfi 0002, Ilia Markov, Walter Daelemans |
COLING | 3 |
| 2020 | Transfer Learning for Digital Heritage Collections: Comparing Neural Machine Translation at the Subword-level and Character-levelabstractTransfer learning via pre-training has become an important strategy for the efficient application of NLP methods in domains where only limited training data is available.This paper reports on a focused case study in which we apply transfer learning in the context of neural machine translation (French-Dutch) for cultural heritage metadata (i.e.titles of artistic works).Nowadays, neural machine translation (NMT) is commonly applied at the subword level using byte-pair encoding (BPE), because word-level models struggle with rare and out-of-vocabulary words.Because unseen vocabulary is a significant issue in domain adaptation, BPE seems a better fit for transfer learning across text varieties.We discuss an experiment in which we compare a subword-level to a character-level NMT approach.We pre-trained models on a large, generic corpus and fine-tuned them in a two-stage process: first, on a domain-specific dataset extracted from Wikipedia, and then on our metadata.While our experiments show comparable performance for character-level and BPEbased models on the general dataset, we demonstrate that the character-level approach nevertheless yields major downstream performance gains during the subsequent stages of fine-tuning.We therefore conclude that character-level translation can be beneficial compared to the popular subword-level approach in the cultural heritage domain. Nikolay Banar, Karine Lasaracina, Walter Daelemans, Mike Kestemont |
ICAART (1) | 3 |
| 2020 | The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual EuropeabstractMultilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions. Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon |
LREC | 26 |
| 2020 | Orthographic Codes and the Neighborhood Effect: Lessons from Information TheoryabstractWe consider the orthographic neighborhood effect: the effect that words with more orthographic similarity to other words are read faster. The neighborhood effect serves as an important control variable in psycholinguistic studies of word reading, and explains variance in addition to word length and word frequency. Following previous work, we model the neighborhood effect as the average distance to neighbors in feature space for three feature sets: slots, character ngrams and skipgrams. We optimize each of these feature sets and find evidence for language-independent optima, across five megastudy corpora from five alphabetic languages. Additionally, we show that weighting features using the inverse of mutual information (MI) improves the neighborhood effect significantly for all languages. We analyze the inverse feature weighting, and show that, across languages, grammatical morphemes get the lowest weights. Finally, we perform the same experiments on Korean Hangul, a non-alphabetic writing system, where we find the opposite results: slower responses as a function of denser neighborhoods, and a negative effect of inverse feature weighting. This raises the question of whether this is a cognitive effect, or an effect of the way we represent Hangul orthography, and indicates more research is needed. Stéphan Tulkens, Dominiek Sandra, Walter Daelemans |
LREC | 3 |
| 2019 | Unsupervised concept extraction from clinical text through semantic composition
Stéphan Tulkens, Simon Suster, Walter Daelemans |
J. Biomed. Informatics | 3 |
| 2018 | Enhancing General Sentiment Lexicons for Domain-Specific UseabstractLexicon based methods for sentiment analysis rely on high quality polarity lexicons. In recent years, automatic methods for inducing lexicons have increased the viability of lexicon based methods for polarity classification. SentProp is a framework for inducing domain-specific polarities from word embeddings. We elaborate on SentProp by evaluating its use for enhancing DuOMan, a general-purpose lexicon, for use in the political domain. By adding only top sentiment bearing words from the vocabulary and applying small polarity shifts in the general-purpose lexicon, we increase accuracy in an in-domain classification task. The enhanced lexicon performs worse than the original lexicon in an out-domain task, showing that the words we added and the polarity shifts we applied are domain-specific and do not translate well to an out-domain setting. Tim Kreutz, Walter Daelemans |
COLING | 2 |
| 2018 | From Strings to Other Things: Linking the Neighborhood and Transposition Effects in Word ReadingabstractWe investigate the relation between the transposition and deletion effects in word reading, i.e., the finding that readers can successfully read "SLAT" as "SALT", or "WRK" as "WORK", and the neighborhood effect.In particular, we investigate whether lexical orthographic neighborhoods take into account transposition and deletion in determining neighbors.If this is the case, it is more likely that the neighborhood effect takes place early during processing, and does not solely rely on similarity of internal representations.We introduce a new neighborhood measure, rd20, which can be used to quantify neighborhood effects over arbitrary feature spaces.We calculate the rd20 over large sets of words in three languages using various feature sets and show that feature sets that do not allow for transposition or deletion explain more variance in Reaction Time (RT) measurements.We also show that the rd20 can be calculated using the hidden state representations of an Multi-Layer Perceptron, and show that these explain less variance than the raw features.We conclude that the neighborhood effect is unlikely to have a perceptual basis, but is more likely to be the result of items co-activating after recognition.All code is available at: www.github.com/clips/conll2018 Stéphan Tulkens, Dominiek Sandra, Walter Daelemans |
CoNLL | 3 |
| 2018 | WordKit: a Python Package for Orthographic and Phonological Featurization
Stéphan Tulkens, Dominiek Sandra, Walter Daelemans |
LREC | 3 |
| 2018 | CliCR: a Dataset of Clinical Case Reports for Machine Reading ComprehensionabstractSimon Šuster, Walter Daelemans. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Simon Suster, Walter Daelemans |
NAACL-HLT | 2 |
| 2018 | Patient representation learning and interpretable evaluation using clinical notes
Madhumita Sushil, Simon Suster, Kim Luyckx, Walter Daelemans |
J. Biomed. Informatics | 4 |
| 2017 | Clinical Machine Comprehension Using Case Reports
Simon Suster, Walter Daelemans |
AMIA | 2 |
| 2017 | Distributional learning and lexical category acquisition: What makes words easy to categorize?
Giovanni Cassani, Robert Grimm 0003, Steven Gillis, Walter Daelemans |
CogSci | 4 |
| 2017 | Evidence for a facilitatory effect of multi-word units on child word learning
Robert Grimm 0003, Giovanni Cassani, Steven Gillis, Walter Daelemans |
CogSci | 4 |
| 2017 | Selecting relevant features from the electronic health record for clinical code prediction
Elyne Scheurwegs, Boris Cule, Kim Luyckx, Léon Luyten, Walter Daelemans |
J. Biomed. Informatics | 5 |
| 2017 | Assigning clinical codes with data-driven concept representation on Dutch clinical free text
Elyne Scheurwegs, Kim Luyckx, Léon Luyten, Bart Goethals, Walter Daelemans |
J. Biomed. Informatics | 5 |
| 2016 | Constraining the Search Space in Cross-Situational Word Learning: Different Models Make Different Predictions
Giovanni Cassani, Robert Grimm 0003, Steven Gillis, Walter Daelemans |
CogSci | 4 |
| 2016 | Evaluating Unsupervised Dutch Word Embeddings as a Linguistic Resource
Stéphan Tulkens, Chris Emmery, Walter Daelemans |
LREC | 3 |
| 2016 | TwiSty: A Multilingual Twitter Stylometry Corpus for Gender and Personality Profiling
Ben Verhoeven, Walter Daelemans, Barbara Plank |
LREC | 2 |
| 2016 | Authenticating the writings of Julius Caesar
Mike Kestemont, Justin Anthony Stover, Moshe Koppel, Folgert Karsdorp, Walter Daelemans |
Expert Syst. Appl. | 5 |
| 2016 | Data integration of structured and unstructured sources for assigning clinical codes to patient staysabstractOBJECTIVE: Enormous amounts of healthcare data are becoming increasingly accessible through the large-scale adoption of electronic health records. In this work, structured and unstructured (textual) data are combined to assign clinical diagnostic and procedural codes (specifically ICD-9-CM) to patient stays. We investigate whether integrating these heterogeneous data types improves prediction strength compared to using the data types in isolation. METHODS: Two separate data integration approaches were evaluated. Early data integration combines features of several sources within a single model, and late data integration learns a separate model per data source and combines these predictions with a meta-learner. This is evaluated on data sources and clinical codes from a broad set of medical specialties. RESULTS: When compared with the best individual prediction source, late data integration leads to improvements in predictive power (eg, overall F-measure increased from 30.6% to 38.3% for International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) diagnostic codes), while early data integration is less consistent. The predictive strength strongly differs between medical specialties, both for ICD-9-CM diagnostic and procedural codes. DISCUSSION: Structured data provides complementary information to unstructured data (and vice versa) for predicting ICD-9-CM codes. This can be captured most effectively by the proposed late data integration approach. CONCLUSIONS: We demonstrated that models using multiple electronic health record data sources systematically outperform models using data sources in isolation in the task of predicting ICD-9-CM codes over a broad range of medical specialties. Elyne Scheurwegs, Kim Luyckx, Léon Luyten, Walter Daelemans, Tim Van den Bulcke |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | Multimodular Text Normalization of Dutch User-Generated ContentabstractAs social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization. Sarah Schulz, Guy De Pauw, Orphée De Clercq, Bart Desmet, Véronique Hoste, Walter Daelemans, Lieve Macken |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2014 | Creative Web Services with Pattern
Tom De Smedt, Lucas Nijs, Walter Daelemans |
ICCC | 3 |
| 2014 | The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite |
LREC | 10 |
| 2014 | CLiPS Stylometry Investigation (CSI) corpus: A Dutch corpus for the detection of age, gender, personality, sentiment and deception in text
Ben Verhoeven, Walter Daelemans |
LREC | 2 |
| 2014 | Using Wiktionary to Build an Italian Part-of-Speech Tagger
Tom De Smedt, Fabio Marfia, Matteo Matteucci, Walter Daelemans |
NLDB | 4 |
| 2014 | Evaluating and understanding text-based stock price prediction models
Enric Junqué de Fortuny, Tom De Smedt, David Martens, Walter Daelemans |
Inf. Process. Manag. | 4 |
| 2013 | Explanation in Computational Stylometry
Walter Daelemans |
CICLing (2) | 1 |
| 2013 | Self-taught assistive vocal interfaces: an overview of the ALADIN projectabstractThis paper gives an overview of research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing ap-proaches, the vocal interface is trained by the end-user himself, which means it can be used with any vocabulary and grammar, and that it is maximally adapted to the — possibly dysarthric — speech of the user. This paper describes the overall learn-ing framework, the user-centred design and evaluation aspects, database collection and approaches taken to combat problems such as noise and erroneous input. Index Terms: vocal user interface, user-centred design, self-taught learning, speech database, dysarthric speech Jort F. Gemmeke, Bart Ons, Netsanet M. Tessema, Hugo Van hamme, Janneke van de Loo, Guy De Pauw, Walter Daelemans, Jonathan Huyghe, Jan Derboven, Lode Vuegen, Bert Van Den Broeck, Peter Karsmakers, Bart Vanrumste |
INTERSPEECH | 7 |
| 2012 | Improving Topic Classification for Highly Inflective Languages
Jurgita Kapociute-Dzikiene, Frederik Vaassen, Walter Daelemans, Algis Krupavicius |
COLING | 3 |
| 2012 | A Statistical Relational Learning Approach to Identifying Evidence Based Medicine Categories
Mathias Verbeke, Vincent Van Asch, Roser Morante, Paolo Frasconi, Walter Daelemans, Luc De Raedt |
EMNLP-CoNLL | 5 |
| 2012 | A Self-Learning Assistive Vocal Interface Based on Vocabulary Learning and Grammar InductionabstractThis paper introduces research within the ALADIN project, which aims to develop an assistive vocal interface for people with a physical impairment. In contrast to existing approaches, the vocal interface is self-learning which means it can be used with any language, dialect, vocabulary and grammar. The paper describes the overall learning framework, and the two components that will provide vocabulary learning and grammar induction. In addition, the paper describes encouraging results of early implementations of these vocabulary and grammar learning components, applied to recorded sessions of a vocally guided card game, patience. Index Terms: language acquisition, word finding, grammar induction, non-negative matrix factorization, concept tagging Jort F. Gemmeke, Janneke van de Loo, Guy De Pauw, Joris Driesen, Hugo Van hamme, Walter Daelemans |
INTERSPEECH | 6 |
| 2012 | The Netlog Corpus. A Resource for the Study of Flemish Dutch Internet Language
Mike Kestemont, Claudia Peersman, Benny De Decker, Guy De Pauw, Kim Luyckx, Roser Morante, Frederik Vaassen, Janneke van de Loo, Walter Daelemans |
LREC | 9 |
| 2012 | ConanDoyle-neg: Annotation of negation cues and their scope in Conan Doyle stories
Roser Morante, Walter Daelemans |
LREC | 2 |
| 2012 | "Vreselijk mooi!" (terribly beautiful): A Subjectivity Lexicon for Dutch Adjectives
Tom De Smedt, Walter Daelemans |
LREC | 2 |
| 2012 | Media coverage in times of political crisis: A text mining approach
Enric Junqué de Fortuny, Tom De Smedt, David Martens, Walter Daelemans |
Expert Syst. Appl. | 4 |
| 2012 | Pattern for Python
Tom De Smedt, Walter Daelemans |
J. Mach. Learn. Res. | 2 |
| 2011 | Generative Art Inspired by Nature, Using NodeBox
Tom De Smedt, Ludivine Lechat, Walter Daelemans |
EvoApplications (2) | 3 |
| 2011 | Kernel-Based Logical and Relational Learning with kLog for Hedge Cue Detection
Mathias Verbeke, Paolo Frasconi, Vincent Van Asch, Roser Morante, Walter Daelemans, Luc De Raedt |
ILP | 5 |
| 2010 | A Chunk-Driven Bootstrapping Approach to Extracting Translation Patterns
Lieve Macken, Walter Daelemans |
CICLing | 2 |
| 2010 | Highlights of the BioTM 2010 workshop on advances in bio text miningabstractRecently, the application of text mining (TM) and natural language processing (NLP) techniques to the biological and medical sciences has received increasing interest. In addition to many new workshops and conferences arising in this domain, recently also a number of community-wide tasks were conducted to benchmark text mining techniques on specific challenges (e.g. BioCreative, BioNLP Shared Task, ...) Thomas Abeel, Sofie Van Landeghem, Roser Morante, Vincent Van Asch, Yves Van de Peer, Walter Daelemans, Yvan Saeys |
BMC Bioinform. | 6 |
| 2010 | Colin de la Higuera: Grammatical inference: learning automata and grammars - Cambridge University Press, 2010, iv + 417 pages
Walter Daelemans |
Mach. Transl. | 1 |
| 2009 | A Metalearning Approach to Processing the Scope of Negation
Roser Morante, Walter Daelemans |
CoNLL | 2 |
| 2009 | A Robust and Extensible Exemplar-Based Model of Thematic Fit
Bram Vandekerckhove, Dominiek Sandra, Walter Daelemans |
EACL | 3 |
| 2008 | Semantic and Syntactic Features for Dutch Coreference Resolution
Iris Hendrickx, Véronique Hoste, Walter Daelemans |
CICLing | 3 |
| 2008 | Authorship Attribution and Verification with Many Authors and Limited Data
Kim Luyckx, Walter Daelemans |
COLING | 2 |
| 2008 | A Combined Memory-Based Semantic Role Labeler of English
Roser Morante, Walter Daelemans, Vincent Van Asch |
CoNLL | 2 |
| 2008 | Learning the Scope of Negation in Biomedical Texts
Roser Morante, Anthony M. L. Liekens, Walter Daelemans |
EMNLP | 3 |
| 2008 | CNTS: Memory-Based Learning of Generating Repeated References
Iris Hendrickx, Walter Daelemans, Kim Luyckx, Roser Morante, Vincent Van Asch |
INLG | 2 |
| 2008 | A Coreference Corpus and Resolution System for Dutch
Iris Hendrickx, Gosse Bouma, Frederik Coppens, Walter Daelemans, Véronique Hoste, Geert Kloosterman, Anne-Marie Mineur, Joeri Van Der Vloet, Jean-Luc Verschelde |
LREC | 4 |
| 2008 | Personae: a Corpus for Author and Personality Prediction from Text
Kim Luyckx, Walter Daelemans |
LREC | 2 |
| 2008 | Guest Editors' Introduction: Special issue of Selected Papers from ECML PKDD 2008
Walter Daelemans, Bart Goethals, Katharina Morik |
Data Min. Knowl. Discov. | 1 |
| 2008 | Guest Editors' introduction: special issue of selected papers from ECML PKDD 2008
Walter Daelemans, Bart Goethals, Katharina Morik |
Mach. Learn. | 1 |
| 2007 | Letter to the Editor
Walter Daelemans, Antal van den Bosch |
Comput. Linguistics | 1 |
| 2006 | A Mission for Computational Natural Language Learning
Walter Daelemans |
CoNLL | 1 |
| 2006 | Investigating Lexical Substitution Scoring for Subtitle Generation
Oren Glickman, Ido Dagan, Walter Daelemans, Mikaela Keller, Samy Bengio |
CoNLL | 3 |
| 2006 | A mixed word / morphological approach for extending CELEX for high coverage on contemporary large corpora
Joris Vaneyghen, Guy De Pauw, Dirk Van Compernolle, Walter Daelemans |
LREC | 4 |
| 2005 | Improving Sequence Segmentation Learning by Predicting Trigrams
Antal van den Bosch, Walter Daelemans |
CoNLL | 2 |
| 2004 | Memory-based semantic role labeling: Optimizing features, algorithm, and output
Antal van den Bosch, Sander Canisius, Walter Daelemans, Iris Hendrickx, Erik F. Tjong Kim Sang |
CoNLL | 3 |
| 2004 | Automatic Sentence Simplification for Subtitling in Dutch and English
Walter Daelemans, Anja Höthker, Erik F. Tjong Kim Sang |
LREC | 1 |
| 2004 | Evaluation and Adaptation of the Celex Dutch Morphological Database
Tom Laureys, Guy De Pauw, Hugo Van hamme, Walter Daelemans, Dirk Van Compernolle |
LREC | 4 |
| 2004 | Multimodal, Multilingual Resources in the Subtitling Process
Stelios Piperidis, Iason Demiros, Prokopis Prokopidis, Peter Vanroose, Anja Höthker, Walter Daelemans, Elsa Sklavounou, Manos Konstantinou, Yannis Karavidas |
LREC | 6 |
| 2004 | Unsupervised Text Mining for Ontology Extraction: An Evaluation of Statistical Measures
Marie-Laure Reinberger, Walter Daelemans |
LREC | 2 |
| 2004 | Recent Advances in Example-Based Machine Translation edited by Michael CarlAndy Way
Walter Daelemans |
Comput. Linguistics | 1 |
| 2004 | Using rule-induction techniques to model pronunciation variation in Dutch
Véronique Hoste, Walter Daelemans, Steven Gillis |
Comput. Speech Lang. | 2 |
| 2003 | Learning to Predict Pitch Accents and Prosodic Boundaries in DutchabstractWe train a decision tree inducer (CART) and a memory-based classifier (MBL) on predicting prosodic pitch accents and breaks in Dutch text, on the basis of shallow, easy-to-compute features. We train the algorithms on both tasks individually and on the two tasks simultaneously. The parameters of both algorithms and the selection of features are optimized per task with iterative deepening, an efficient wrapper procedure that uses progressive sampling of training data. Results show a consistent significant advantage of MBL over CART, and also indicate that task combination can be done at the cost of little generalization score loss. Tests on cross-validated data and on held-out data yield F-scores of MBL on accent placement of 84 and 87, respectively, and on breaks of 88 and 91, respectively. Accent placement is shown to outperform an informed baseline rule; reliably predicting breaks other than those already indicated by intra-sentential punctuation, however, appears to be more challenging. Erwin Marsi, Martin Reynaert, Antal van den Bosch, Walter Daelemans, Véronique Hoste |
ACL | 4 |
| 2003 | Is Shallow Parsing Useful for Unsupervised Learning of Semantic Clusters?
Marie-Laure Reinberger, Walter Daelemans |
CICLing | 2 |
| 2003 | Memory-Based Named Entity Recognition using Unannotated Data
Fien De Meulder, Walter Daelemans |
CoNLL | 2 |
| 2003 | Combined Optimization of Feature Selection and Algorithm Parameters in Machine Learning of Language
Walter Daelemans, Véronique Hoste, Fien De Meulder, Bart Naudts |
ECML | 1 |
| 2002 | Transcription of out-of-vocabulary words in large vocabulary speech recognition based on phoneme-to-grapheme conversionabstractIn this paper, we describe a method to enhance the readability of the textual output in a large vocabulary continuous speech recognizer when out-of-vocabulary words occur. Bart Decadt, Jacques Duchateau, Walter Daelemans, Patrick Wambacq |
ICASSP | 3 |
| 2002 | Combining information sources for memory-based pitch accent placementabstractWe describe results on pitch accent placement in Dutch text obtained with a memory-based learning approach. The training material consists of newspaper texts that have been prosodically annotated by humans, and subsequently enriched with linguistic features and informational metrics using generally available, lowcost, shallow, knowledge-poor tools. We report on the effects of context-modelling and the nearest neighbours parameter (k), and show the advantage of combining features of a different nature, where the best performance yields a cross-validated F-score of 82. Evaluation on an independent test corpus shows that our approach outperforms existing TTS systems for Dutch. 1. Erwin Marsi, Bertjan Busser, Walter Daelemans, Véronique Hoste, Martin Reynaert, Antal van den Bosch |
INTERSPEECH | 3 |
| 2002 | Dutch HLT resources: from BLARK to priority listsabstract\n Contains fulltext :\n 76208.pdf (author's version ) (Open Access)\n Helmer Strik, Walter Daelemans, Diana Binnenpoorte, Janienke Sturm, Folkert de Vriend, Catia Cucchiarini |
INTERSPEECH | 2 |
| 2002 | A Field Survey for Establishing Priorities in the Development of HLT Resources for Dutch
Diana Binnenpoorte, Folkert de Vriend, Janienke Sturm, Walter Daelemans, Helmer Strik, Catia Cucchiarini |
LREC | 4 |
| 2002 | Evaluation of Machine Learning Methods for Natural Language Processing Tasks
Walter Daelemans, Véronique Hoste |
LREC | 1 |
| 2002 | Logistic-based patient grouping for multi-disciplinary treatment
Laura Maruster, A. J. M. M. Weijters, Geerhard de Vries, Antal van den Bosch, Walter Daelemans |
Artif. Intell. Medicine | 5 |
| 2002 | Introduction to Special Issue on Machine Learning Approaches to Shallow Parsing
James Alistair Hammerton, Miles Osborne, Susan Armstrong, Walter Daelemans |
J. Mach. Learn. Res. | 4 |
| 2002 | Parameter optimization for machine-learning of word sense disambiguationabstractVarious Machine Learning (ML) approaches have been demonstrated to produce relatively successful Word Sense Disambiguation (WSD) systems. There are still unexplained differences among the performance measurements of different algorithms, hence it is warranted to deepen the investigation into which algorithm has the right ‘bias’ for this task. In this paper, we show that this is not easy to accomplish, due to intricate interactions between information sources, parameter settings, and properties of the training data. We investigate the impact of parameter optimization on generalization accuracy in a memory-based learning approach to English and Dutch WSD. A ‘word-expert’ architecture was adopted, yielding a set of classifiers, each specialized in one single wordform. The experts consist of multiple memory-based learning classifiers, each taking different information sources as input, combined in a voting scheme. We optimized the architectural and parametric settings for each individual word-expert by performing cross-validation experiments on the learning material. The results of these experiments show that the variation of both the algorithmic parameters and the information sources available to the classifiers leads to large fluctuations in accuracy. We demonstrate that optimization per word-expert leads to an overall significant improvement in the generalization accuracies of the produced WSD systems. Véronique Hoste, Iris Hendrickx, Walter Daelemans, Antal van den Bosch |
Nat. Lang. Eng. | 3 |
| 2001 | Improving Accuracy in NLP Through Combination of Machine Learning SystemsabstractWe examine how differences in language models, learned by different data-driven systems performing the same NLP task, can be exploited to yield a higher accuracy than the best individual system. We do this by means of experiments involving the task of morphosyntactic word class tagging, on the basis of three different tagged corpora. Four well-known tagger generators (hidden Markov model, memory-based, transformation rules, and maximum entropy) are trained on the same corpus data. After comparison, their outputs are combined using several voting strategies and second-stage classifiers. All combination taggers outperform their best component. The reduction in error rate varies with the material in question, but can be as high as 24.3% with the LOB corpus. Hans van Halteren, Jakub Zavrel, Walter Daelemans |
Comput. Linguistics | 3 |
| 2001 | Complex answers: a case study using a WWW question answering systemabstractWe investigate the problem of complex answers in question answering. Complex answers consist of several simple answers. We describe the online question answering system SHAPAQA, and using data from this system we show that the problem of complex answers is quite common. We define nine types of complex questions, and suggest two approaches, based on answer frequencies, that allow question answering systems to tackle the problem. Sabine Buchholz, Walter Daelemans |
Nat. Lang. Eng. | 2 |
| 2000 | A Rule Induction Approach to Modeling Regional Pronunciation Variation
Véronique Hoste, Steven Gillis, Walter Daelemans |
COLING | 3 |
| 2000 | Applying System Combination to Base Noun Phrase Identification
Erik F. Tjong Kim Sang, Walter Daelemans, Hervé Déjean, Rob Koeling, Yuval Krymolowski, Vasin Punyakanok, Dan Roth 0001 |
COLING | 2 |
| 2000 | Meta-Learning for Phonemic Annotation of Corpora
Véronique Hoste, Walter Daelemans, Erik F. Tjong Kim Sang, Steven Gillis |
ICML | 2 |
| 2000 | Part of Speech Tagging and Lemmatisation for the Spoken Dutch Corpus
Frank Van Eynde, Jakub Zavrel, Walter Daelemans |
LREC | 3 |
| 2000 | Bootstrapping a Tagged Corpus through Combination of Existing Heterogeneous Taggers
Jakub Zavrel, Walter Daelemans |
LREC | 2 |
| 1999 | Memory-Based Morphological AnalysisabstractWe present a general architecture for efficient and deterministic morphological analysis based on memory-based learning, and apply it to morphological analysis of Dutch.The system makes direct mappings from letters in context to rich categories that encode morphological boundaries, syntactic class labels, and spelling changes.Both precision and recall of labeled morphemes are over 84% on held-out dictionary test words and estimated to be over 93% in free text. Antal van den Bosch, Walter Daelemans |
ACL | 2 |
| 1999 | Cascaded Grammatical Relation Assignment
Sabine Buchholz, Jorn Veenstra, Walter Daelemans |
EMNLP | 3 |
| 1999 | Machine learning of word pronunciation: the case against abstractionabstractAn adequate approach to speech translation for small to medium sized tasks is the use of subsequential trans-ducers —a finite state model — as language model for a speech recognizer. These transducers can be automati-cally trained from sample corpora. The use of manually defined categories improves the training of the subsequential transducers when the avail-able data are scarce. These categories depend on the source and target languages we want to translate. We introduce an automatic approach to derive cate-gories that can be used in training subsequential transduc-ers. This approach extends monolingual word clustering methods to the bilingual case using alignments obtained from statistical models. Experimental results indicate that the models trained with these categories have lower trans-lation errors. 1 Bertjan Busser, Walter Daelemans, Antal van den Bosch |
EUROSPEECH | 2 |
| 1999 | Introduction to the special issue on memory-based language processingabstractMemory-based language processing (MBLP) views language processing as being based on the direct reuse of previous experience rather than on the use of rules or other structures extracted from that experience. In such a framework, language acquisition is modelled as the storage of examples in memory,and language processing as analogical or similarity-based reasoning. We briefly discuss the properties and origins of this family of techniques, and provide an overview of current approaches and issues. Walter Daelemans |
J. Exp. Theor. Artif. Intell. | 1 |
| 1999 | Forgetting Exceptions is Harmful in Language Learning
Walter Daelemans, Antal van den Bosch, Jakub Zavrel |
Mach. Learn. | 1 |
| 1998 | Do Not Forget: Full Memory in Memory-Based Learning of Word Pronunciation
Antal van den Bosch, Walter Daelemans |
CoNLL | 2 |
| 1998 | Modularity in Inductively-Learned Word Pronunciation Systems
Antal van den Bosch, A. J. M. M. Weijters, Walter Daelemans |
CoNLL | 3 |
| 1998 | Abstraction is Harmful in Language Learning
Walter Daelemans |
CoNLL | 1 |
| 1997 | Memory-Based Learning: Using Similarity for SmoothingabstractThis paper analyses the relation between the use of similarity in Memory-Based Learning and the notion of backed-off smoothing in statistical language modeling. We show that the two approaches are closely related, and we argue that feature weighting methods in the Memory-Based paradigm can offer the advantage of automatically specifying a suitable domain-specific hierarchy between most specific and most general conditioning information without the need for a large number of parameters. We report two applications of this approach: PP-attachment and POS-tagging. Our method achieves state-of-the-art performance in both domains, and allows the easy integration of diverse information sources, such as rich lexical representations. Jakub Zavrel, Walter Daelemans |
ACL | 2 |
| 1997 | Resolving PP attachment Ambiguities with Memory-Based Learning
Jakub Zavrel, Walter Daelemans, Jorn Veenstra |
CoNLL | 2 |
| 1997 | Empirical Learning of Natural Language Processing Task
Walter Daelemans, Antal van den Bosch, A. J. M. M. Weijters |
ECML | 1 |
| 1996 | Unsupervised Discovery of Phonological Categories through Supervised Learning of Morphological Rules
Walter Daelemans, Peter Berck, Steven Gillis |
COLING | 1 |
| 1994 | The Acquisition of Stress: A Data-Oriented Approach
Walter Daelemans, Gert Durieux, Steven Gillis |
Comput. Linguistics | 1 |
| 1994 | Default inheritance in an object-oriented representation of linguistic categories
Walter Daelemans, Koenraad De Smedt |
Int. J. Hum. Comput. Stud. | 1 |
| 1993 | Data-Oriented Methods for Grapheme-to-Phoneme Conversion
Antal van den Bosch, Walter Daelemans |
EACL | 2 |
| 1993 | Tabtalk: reusability in data-oriented grapheme-to-phoneme conversionabstractIn the traditional (knowledge-based) approach to the design of grapheme-to-phoneme modules in text-to-speech systems, it is claimed that various explicitly coded, language-specific, linguistic knowledge sources are necessary for a good performance. Due to knowledge acquisition bottlenecks, this implies long development cycles. As an alternative, we propose to use inductive methods from machine learning in a simple combined Trie Search and Similarity-Based Reasoning approach and show that, for Dutch, its performance is better than that of the knowledge-based approach and backpropagation learning. Furthermore, we show that our approach is reusable for any language for which a training corpus exists. Keywords: grapheme-to-phoneme conversion, text-tospeech, trie search, similarity-based reasoning, machine learning INTRODUCTION The larger part of research on grapheme-to-phoneme conversion focuses on developing systems that implement various levels of language-specific linguistic knowledge.... Walter Daelemans, Antal van den Bosch |
EUROSPEECH | 1 |
| 1992 | Inheritance in Natural Language Processing
Walter Daelemans, Koenraad De Smedt, Gerald Gazdar |
Comput. Linguistics | 1 |
| 1988 | GRAFON: a grapheme-to-phoneme conversion system for Duth
Walter Daelemans |
COLING | 1 |
| 1987 | A Tool For The Automatic Creation, Extension And Updating Of Lexical Knowledge Bases
Walter Daelemans |
EACL | 1 |