Véronique Hoste

dblp:h/VeroniqueHoste · DBLP profile ↗
← Back
64ranked-venue papers
6as first author
16since 2021 · last 2026
0000-0002-0539-4630ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Counter-Hypothesis Generation: Towards Evaluating How LLMs Reason about Alternatives
abstract
Reasoning about alternatives is a fundamental component of human cognition and argumentation, yet it remains unclear whether large language models (LLMs) can coherently generate and assess them. This paper introduces Counter-Hypothesis Generation (CHG), a novel task for evaluating how LLMs construct plausible hypotheses when contextual information changes. Inspired by open-domain commonsense reasoning, where models infer and compare multiple explanations, CHG bridges commonsense and counterfactual reasoning by requiring models to generate hypotheses that remain logically consistent with modified premises. We present a test set annotated by a human expert and complemented with counter-hypotheses generated by OpenAI-o3 and DeepSeek-r1. Experimental results reveal that even advanced reasoning models exhibit notable limitations in counter-hypothesis generation.
Marzieh Abdolmaleki, Aaron Maladry, Véronique Hoste, Els Lefever
LREC3
2026 Learning through News: Bridging the Gap between Algorithmic Recommendation and Human Curation
abstract
News recommendation systems play a central role in how readers access and process current events. Most recommenders’ underlying algorithmic strategies, however, prioritize user engagement over comprehension, amplifying risks of misinformation and filter bubbles. This study investigates whether fine-grained content-based recommendation strategies favor human knowledge retention and explores how such a content-based recommendation can be operationalized using event coreference–based document modeling. To this purpose, we first measure the effect of manually curated content-based news recommendation on knowledge retention across five news topics with 126 Dutch speaking participants. Next, we investigate document retrieval by comparing a state-of-the-art event coreference resolution system for Dutch which recommends news articles based on event chains with a document similarity retrieval baseline using state-of-the-art embedding models in three increasingly more complex test settings. The results demonstrate that human-curated content-based recommendation can positively and significantly impact readers’ knowledge retention. Moreover, we show that a fine-grained coreference system can approach said level of human curation better than state-of-the-art document retrieval methods. In general, this holds potential for scalable, comprehension-oriented news recommendation.
Florian Debaene, Loic De Langhe, Orphée De Clercq, Véronique Hoste
LREC4
2026 LoveHate: Stance Detection and Generation for Multiple Topics in User-generated Comments in Russian and English
abstract
This paper introduces LoveHate, a new multi-topic corpus of user-generated arguments in Russian, collected from the historical data of the debate platform lovehate.ru. The dataset contains nearly 19,000 posts spanning socially and politically relevant topics, each mapped to binary pro and con stances. We test multiple approaches to stance detection and stance generation across Russian and English data, including translated variants, using both classifier-based (Roberta, RuRoberta) and instruction-tuned generative (Llama, Qwen) models. Results demonstrate that language-specific pretraining yields the strongest performance for stance classification (F1 = 0.892 with RuRoberta), while multilingual generative models – when fine-tuned on sufficient data – can effectively generate stance in Russian without explicit Russian pretraining. Cross-domain experiments show that English datasets generalise better across corpora, whereas Russian data capture language- and culture-specific argumentation but are less effective for generalizable models. Generating topics remains a more challenging task for both Russian and English data. The dataset and accompanying results contribute to multilingual stance research and provide a valuable new resource for argument mining in Russian.
Natalia Evgrafova, Véronique Hoste, Els Lefever
LREC2
2026 Towards Reliable Evaluation of Emotional Text Generation in LLMs: Human vs. Automatic Metrics
Sadegh Jafari, Els Lefever, Véronique Hoste
LREC3
2026 Echoes of the Troubadours: A Corpus of Troubadour Poetry for Stylometric Analysis and Authorship Attribution
abstract
We present TrobaCor, a curated corpus of medieval troubadour poetry, which comprises 1668 unique Old Occitan texts by a large variety of authors. Clustering and stylometric experiments show that we can accurately model authorial style beyond topical content, even though formulaic or topically diverse genres remain challenging. Furthermore, we can model and detect traces of an author’s stylistic "DNA" even in short-form collaborative poetry, offering a uniquely fine-grained perspective in the field. In addition, we provide self-organizing map visualizations in order to provide an interpretable view of stylistic patterns across authors. TrobaCor is publicly released to support reproducible research in NLP and digital humanities on this low-resource historical corpus.
Loic De Langhe, Orphée De Clercq, Véronique Hoste
LREC3
2026 Exploring the Transfer of Irony Explanation Generation from English to Dutch
abstract
Explanation generation has gained increasing attention in the field of NLP because it makes the output of classification models more intuitively understandable for humans. This is particularly relevant for complex semantic tasks such as irony detection, where there may not be any explicit linguistic markers. Generative models have shown great potential for irony explanation in earlier work, but most studies have been limited to English. Since this is the highest-resourced language, these capabilities may not be available in languages other than English. To address this gap, this paper analyses the performance of generative models for explanation generation in Dutch, a lower-resourced but closely related language to English. Our work shows that larger proprietary models, like GPT-4, can generate meaningful explanations based on relevant world knowledge, whereas smaller open-source models still struggle to perform this task. Besides quality evaluation, we also analyse the limitations of these models, showing that GPT models struggle most with verbosity and that both open-source and proprietary models exhibit circular reasoning ("this text is ironic because the person expresses this in an ironic way”). Finally, open-source models struggle in particular for Dutch because they fail to produce the relevant world knowledge that is required to understand the irony. All models and data used for the experiments is available at iRONNIE on Hugging Face.
Aaron Maladry, Els Lefever, Cynthia Van Hee, Véronique Hoste
LREC4
2025 Evaluating Transformers for OCR Post-Correction in Early Modern Dutch Theatre
abstract
This paper explores the effectiveness of two types of transformer models — large generative models and sequence-to-sequence models — for automatically post-correcting Optical Character Recognition (OCR) output in early modern Dutch plays. To address the need for optimally aligned data, we create a parallel dataset based on the OCRed and ground truth versions from the EmDComF corpus using state-of-the-art alignment techniques. By combining character-based and semantic methods, we design and release a qualitative OCR-to-gold parallel dataset, selecting the alignment with the lowest Character Error Rate (CER) for all alignment pairs. We then fine-tune and evaluate five generative models and four sequence-to-sequence models on the OCR post-correction dataset. Results show that sequence-to-sequence models generally outperform generative models in this task, correcting more OCR errors and overgenerating and undergenerating less, with mBART as the best performing system.
Florian Debaene, Aaron Maladry, Els Lefever, Véronique Hoste
COLING4
2025 LDW: Label Divergence Weighting for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) traditionally assumes a unified emotional signal across modalities such as text, audio, and video. However, recent findings suggest that each modality may convey distinct affective perspectives. Motivated by perspectivist theories from cognitive science and natural language processing, this paper introduces Label Divergence Weighting (LDW), a modality-weighting strategy that dynamically adjusts trust in each modality based on its alignment with the overall sentiment label. The LDW framework leverages training-time supervision from the divergence between unimodal and multimodal sentiment annotations to learn modality reliability, and applies this learning to unseen data without requiring unimodal labels at inference time. Integrated into a multitask variant of the Tensor Fusion Network (MTFN), the proposed LDW-MTFN model achieves state-of-the-art results on both the acted Chinese dataset CH-SIMS and the authentic English dataset UniC. Extensive experiments and ablation studies demonstrate the robustness and generalizability of LDW across datasets with different cultural, linguistic, and environmental characteristics.
Quanqi Du, Loic De Langhe, Els Lefever, Véronique Hoste
ACM Multimedia4
2025 Why Robots Are Bad at Detecting Their Mistakes: Limitations of Miscommunication Detection in Human-Robot Dialogue
abstract
Detecting miscommunication in human-robot interaction is a critical function for maintaining user engagement and trust. While humans effortlessly detect communication errors in conversations through both verbal and non-verbal cues, robots face significant challenges in interpreting non-verbal feedback, despite advances in computer vision for recognizing affective expressions. This research evaluates the effectiveness of machine learning models in detecting miscommunications in robot dialogue. Using a multi-modal dataset of 240 human-robot conversations, where four distinct types of conversational failures were systematically introduced, we assess the performance of state-of-the-art computer vision models. After each conversational turn, users provided feedback on whether they perceived an error, enabling an analysis of the models’ ability to accurately detect robot mistakes. Despite using state-of-the-art models, the performance barely exceeds random chance in identifying miscommunication, while on a dataset with more expressive emotional content, they successfully identified confused states. To explore the underlying cause, we asked human raters to do the same. They could also only identify around half of the induced miscommunications, similarly to our model. These results uncover a fundamental limitation in identifying robot miscommunications in dialogue: even when users perceive the induced miscommunication as such, they often do not communicate this to their robotic conversation partner. This knowledge can shape expectations of the performance of computer vision models and can help researchers to design better human-robot conversations by deliberately eliciting feedback where needed.
Ruben Janssens, Jens De Bock, Sofie Labat, Eva Verhelst, Véronique Hoste, Tony Belpaeme
RO-MAN5
2024 Enhancing Unrestricted Cross-Document Event Coreference with Graph Reconstruction Networks
abstract
Event Coreference Resolution remains a challenging discourse-oriented task within the domain of Natural Language Processing. In this paper we propose a methodology where we combine traditional mention-pair coreference models with a lightweight and modular graph reconstruction algorithm. We show that building graph models on top of existing mention-pair models leads to improved performance for both a wide range of baseline mention-pair algorithms as well as a recently developed state-of-the-art model and this at virtually no added computational cost. Moreover, additional experiments seem to indicate that our method is highly robust in low-data settings and that its performance scales with increases in performance for the underlying mention-pair models.
Loic De Langhe, Orphée De Clercq, Véronique Hoste
LREC/COLING3
2024 Human and System Perspectives on the Expression of Irony: An Analysis of Likelihood Labels and Rationales
abstract
In this paper, we examine the recognition of irony by both humans and automatic systems. We achieve this by enhancing the annotations of an English benchmark data set for irony detection. This enhancement involves a layer of human-annotated irony likelihood using a 7-point Likert scale that combines binary annotation with a confidence measure. Additionally, the annotators indicated the trigger words that led them to perceive the text as ironic, which leveraged necessary theoretical insights into the definition of irony and its various forms. By comparing these trigger word spans across annotators, we determine the extent to which humans agree on the source of irony in a text. Finally, we compare the human-annotated spans with sub-token importance attributions for fine-tuned transformers using Layer Integrated Gradients, a state-of-the-art interpretability metric. Our results indicate that our model achieves better performance on tweets that were annotated with high confidence and high agreement. Although automatic systems can identify trigger words with relative success, they still attribute a significant amount of their importance to the wrong tokens.
Aaron Maladry, Alessandra Teresa Cignarella, Els Lefever, Cynthia Van Hee, Véronique Hoste
LREC/COLING5
2023 Fuzzy rough nearest neighbour methods for detecting emotions, hate speech and irony
Olha Kaminska, Chris Cornelis, Véronique Hoste
Inf. Sci.3
2022 Aspect-Based Emotion Analysis and Multimodal Coreference: A Case Study of Customer Comments on Adidas Instagram Posts
abstract
While aspect-based sentiment analysis of user-generated content has received a lot of attention in the past years, emotion detection at the aspect level has been relatively unexplored. Moreover, given the rise of more visual content on social media platforms, we want to meet the ever-growing share of multimodal content. In this paper, we present a multimodal dataset for Aspect-Based Emotion Analysis (ABEA). Additionally, we take the first steps in investigating the utility of multimodal coreference resolution in an ABEA framework. The presented dataset consists of 4,900 comments on 175 images and is annotated with aspect and emotion categories and the emotional dimensions of valence and arousal. Our preliminary experiments suggest that ABEA does not benefit from multimodal coreference resolution, and that aspect and emotion classification only requires textual information. However, when more specific information about the aspects is desired, image recognition could be essential.
Luna De Bruyne, Akbar Karimi 0001, Orphée De Clercq, Andrea Prati 0001, Véronique Hoste
LREC5
2022 Automatic classification of participant roles in cyberbullying: Can we detect victims, bullies, and bystanders in social media text?
abstract
Abstract Successful prevention of cyberbullying depends on the adequate detection of harmful messages. Given the impossibility of human moderation on the Social Web, intelligent systems are required to identify clues of cyberbullying automatically. Much work on cyberbullying detection focuses on detecting abusive language without analyzing the severity of the event nor the participants involved. Automatic analysis of participant roles in cyberbullying traces enables targeted bullying prevention strategies. In this paper, we aim to automatically detect different participant roles involved in textual cyberbullying traces, including bullies, victims, and bystanders. We describe the construction of two cyberbullying corpora (a Dutch and English corpus) that were both manually annotated with bullying types and participant roles and we perform a series of multiclass classification experiments to determine the feasibility of text-based cyberbullying participant role detection. The representative datasets present a data imbalance problem for which we investigate feature filtering and data resampling as skew mitigation techniques. We investigate the performance of feature-engineered single and ensemble classifier setups as well as transformer-based pretrained language models (PLMs). Cross-validation experiments revealed promising results for the detection of cyberbullying roles using PLM fine-tuning techniques, with the best classifier for English (RoBERTa) yielding a macro-averaged ${F_1}$ -score of 55.84%, and the best one for Dutch (RobBERT) yielding an ${F_1}$ -score of 56.73%. Experiment replication data and source code are available at https://osf.io/nb2r3 .
Gilles Jacobs, Cynthia Van Hee, Véronique Hoste
Nat. Lang. Eng.3
2021 A Million Tweets Are Worth a Few Points: Tuning Transformers for Customer Service Tasks
abstract
Amir Hadifar, Sofie Labat, Veronique Hoste, Chris Develder, Thomas Demeester. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Amir Hadifar, Sofie Labat, Véronique Hoste, Chris Develder, Thomas Demeester
NAACL-HLT3
2021 Is neural always better? SMT versus NMT for Dutch text normalization
Claudia Matos Veliz, Orphée De Clercq, Véronique Hoste
Expert Syst. Appl.3
2020 An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus
abstract
Seeing the myriad of existing emotion models, with the categorical versus dimensional opposition the most important dividing line, building an emotion-annotated corpus requires some well thought-out strategies concerning framework choice. In our work on automatic emotion detection in Dutch texts, we investigate this problem by means of two case studies. We find that the labels joy, love, anger, sadness and fear are well-suited to annotate texts coming from various domains and topics, but that the connotation of the labels strongly depends on the origin of the texts. Moreover, it seems that information is lost when an emotional state is forcedly classified in a limited set of categories, indicating that a bi-representational format is desirable when creating an emotion corpus.
Luna De Bruyne, Orphée De Clercq, Véronique Hoste
LREC3
2020 Estimating word-level quality of statistical machine translation output using monolingual information alone
abstract
Abstract Various studies show that statistical machine translation (SMT) systems suffer from fluency errors, especially in the form of grammatical errors and errors related to idiomatic word choices. In this study, we investigate the effectiveness of using monolingual information contained in the machine-translated text to estimate word-level quality of SMT output. We propose a recurrent neural network architecture which uses morpho-syntactic features and word embeddings as word representations within surface and syntactic n-grams. We test the proposed method on two language pairs and for two tasks, namely detecting fluency errors and predicting overall post-editing effort. Our results show that this method is effective for capturing all types of fluency errors at once. Moreover, on the task of predicting post-editing effort, while solely relying on monolingual information, it achieves on-par results with the state-of-the-art quality estimation systems which use both bilingual and monolingual information.
Arda Tezcan, Véronique Hoste, Lieve Macken
Nat. Lang. Eng.2
2019 Estimating post-editing time using a gold-standard set of machine translation errors
Arda Tezcan, Véronique Hoste, Lieve Macken
Comput. Speech Lang.2
2018 A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents
Ayla Rigouts Terryn, Véronique Hoste, Els Lefever
LREC2
2018 We Usually Don't Like Going to the Dentist: Using Common Sense to Detect Irony on Twitter
abstract
Although common sense and connotative knowledge come naturally to most people, computers still struggle to perform well on tasks for which such extratextual information is required. Automatic approaches to sentiment analysis and irony detection have revealed that the lack of such world knowledge undermines classification performance. In this article, we therefore address the challenge of modeling implicit or prototypical sentiment in the framework of automatic irony detection. Starting from manually annotated connoted situation phrases (e.g., “flight delays,” “sitting the whole day at the doctor’s office”), we defined the implicit sentiment held towards such situations automatically by using both a lexico-semantic knowledge base and a data-driven method. We further investigate how such implicit sentiment information affects irony detection by assessing a state-of-the-art irony classifier before and after it is informed with implicit sentiment information.
Cynthia Van Hee, Els Lefever, Véronique Hoste
Comput. Linguistics3
2018 Online suicide prevention through optimised text classification
Bart Desmet, Véronique Hoste
Inf. Sci.2
2016 Monday mornings are my fave : ) #not Exploring the Automatic Recognition of Irony in English tweets
abstract
Recognising and understanding irony is crucial for the improvement natural language processing tasks including sentiment analysis. In this study, we describe the construction of an English Twitter corpus and its annotation for irony based on a newly developed fine-grained annotation scheme. We also explore the feasibility of automatic irony recognition by exploiting a varied set of features including lexical, syntactic, sentiment and semantic (Word2Vec) information. Experiments on a held-out test set show that our irony classifier benefits from this combined information, yielding an F1-score of 67.66%. When explicit hashtag information like #irony is included in the data, the system even obtains an F1-score of 92.77%. A qualitative analysis of the output reveals that recognising irony that results from a polarity clash appears to be (much) more feasible than recognising other forms of ironic utterances (e.g., descriptions of situational irony).
Cynthia Van Hee, Els Lefever, Véronique Hoste
COLING3
2016 Detecting Grammatical Errors in Machine Translation Output Using Dependency Parsing and Treebank Querying
Arda Tezcan, Véronique Hoste, Lieve Macken
EAMT2
2016 Rude waiter but mouthwatering pastries! An exploratory study into Dutch Aspect-Based Sentiment Analysis
Orphée De Clercq, Véronique Hoste
LREC2
2016 Exploring the Realization of Irony in Twitter Data
Cynthia Van Hee, Els Lefever, Véronique Hoste
LREC3
2016 A Classification-based Approach to Economic Event Detection in Dutch News Text
Els Lefever, Véronique Hoste
LREC2
2016 All Mixed Up? Finding the Optimal Feature Set for General Readability Prediction and Its Application to English and Dutch
abstract
Readability research has a long and rich tradition, but there has been too little focus on general readability prediction without targeting a specific audience or text genre. Moreover, although NLP-inspired research has focused on adding more complex readability features, there is still no consensus on which features contribute most to the prediction. In this article, we investigate in close detail the feasibility of constructing a readability prediction system for English and Dutch generic text using supervised machine learning. Based on readability assessments by both experts and crowdsourcing, we implement different types of text characteristics ranging from easy-to-compute superficial text characteristics to features requiring deep linguistic processing, resulting in ten different feature groups. Both a regression and classification set-up are investigated reflecting the two possible readability prediction tasks: scoring individual texts or comparing two texts. We show that going beyond correlation calculations for readability optimization using a wrapper-based genetic algorithm optimization approach is a promising task that provides considerable insights in which feature combinations contribute to the overall readability prediction. Because we also have gold standard information available for those features requiring deep processing, we are able to investigate the true upper bound of our Dutch system. Interestingly, we will observe that the performance of our fully automatic readability prediction pipeline is on par with the pipeline using gold-standard deep syntactic and semantic information.
Orphée De Clercq, Véronique Hoste
Comput. Linguistics2
2016 Multimodular Text Normalization of Dutch User-Generated Content
abstract
As social media constitutes a valuable source for data analysis for a wide range of applications, the need for handling such data arises. However, the nonstandard language used on social media poses problems for natural language processing (NLP) tools, as these are typically trained on standard language material. We propose a text normalization approach to tackle this problem. More specifically, we investigate the usefulness of a multimodular approach to account for the diversity of normalization issues encountered in user-generated content (UGC). We consider three different types of UGC written in Dutch (SNS, SMS, and tweets) and provide a detailed analysis of the performance of the different modules and the overall system. We also apply an extrinsic evaluation by evaluating the performance of a part-of-speech tagger, lemmatizer, and named-entity recognizer before and after normalization.
Sarah Schulz, Guy De Pauw, Orphée De Clercq, Bart Desmet, Véronique Hoste, Walter Daelemans, Lieve Macken
ACM Trans. Intell. Syst. Technol.5
2015 Smart Computer Aided Translation Environment - SCATE
Vincent Vandeghinste, Tom Vanallemeersch, Frank Van Eynde, Geert Heyman, Marie-Francine Moens, Joris Pelemans, Patrick Wambacq, Iulianna Van der Lek-Ciudin, Arda Tezcan, Lieve Macken, Véronique Hoste, Eva Geurts, Mieke Haesen
EAMT11
2015 Fine-grained analysis of explicit and implicit sentiment in financial news articles
Marjan Van de Kauter, Diane Breesch, Véronique Hoste
Expert Syst. Appl.3
2014 Does size matter? A comparison of a large web corpus and a smaller focused corpus for medical term extraction
Klaar Vanopstal, Els Lefever, Véronique Hoste
AMIA3
2014 Towards Shared Datasets for Normalization Research
Orphée De Clercq, Sarah Schulz, Bart Desmet, Véronique Hoste
LREC4
2014 Recognising suicidal messages in Dutch social media
Bart Desmet, Véronique Hoste
LREC2
2014 Evaluation of Automatic Hypernym Extraction from Technical Corpora in English and Dutch
Els Lefever, Marjan Van de Kauter, Véronique Hoste
LREC3
2014 Using the crowd for readability prediction
abstract
Abstract While human annotation is crucial for many natural language processing tasks, it is often very expensive and time-consuming. Inspired by previous work on crowdsourcing, we investigate the viability of using non-expert labels instead of gold standard annotations from experts for a machine learning approach to automatic readability prediction. In order to do so, we evaluate two different methodologies to assess the readability of a wide variety of text material: A more traditional setup in which expert readers make readability judgments and a crowdsourcing setup for users who are not necessarily experts. To this purpose two assessment tools were implemented: a tool where expert readers can rank a batch of texts based on readability, and a lightweight crowdsourcing tool, which invites users to provide pairwise comparisons. To validate this approach, readability assessments for a corpus of written Dutch generic texts were gathered. By collecting multiple assessments per text, we explicitly wanted to level out readers' background knowledge and attitude. Our findings show that the assessments collected through both methodologies are highly consistent and that crowdsourcing is a viable alternative to expert labeling. This is a good news as crowdsourcing is more lightweight to use and can have access to a much wider audience of potential annotators. By performing a set of basic machine learning experiments using a feature set that mainly encodes basic lexical and morpho-syntactic information, we further illustrate how the collected data can be used to perform text comparisons or to assign an absolute readability score to an individual text. We do not focus on optimising the algorithms to achieve the best possible results for the learning tasks, but carry them out to illustrate the various possibilities of our data sets. The results on different data sets, however, show that our system outperforms the readability formulas and a baseline language modelling approach. We conclude that readability assessment by comparing texts is a polyvalent methodology, which can be adapted to specific domains and target audiences if required.
Orphée De Clercq, Véronique Hoste, Bart Desmet, Philip van Oosten, Martine De Cock, Lieve Macken
Nat. Lang. Eng.2
2013 Five Languages Are Better Than One: An Attempt to Bypass the Data Acquisition Bottleneck for WSD
Els Lefever, Véronique Hoste, Martine De Cock
CICLing (1)2
2013 Emotion detection in suicide notes
Bart Desmet, Véronique Hoste
Expert Syst. Appl.2
2012 Evaluating automatic cross-domain Dutch semantic role annotation
Orphée De Clercq, Véronique Hoste, Paola Monachesi
LREC2
2012 Discovering Missing Wikipedia Inter-language Links by means of Cross-lingual Word Sense Disambiguation
Els Lefever, Véronique Hoste, Martine De Cock
LREC2
2012 From keystrokes to annotated process data: Enriching the output of Inputlog with linguistic information
Lieve Macken, Véronique Hoste, Mariëlle Leijten, Luuk van Waes
LREC2
2012 Beyond SoNaR: towards the facilitation of large corpus building efforts
Martin Reynaert, Ineke Schuurman, Véronique Hoste, Nelleke Oostdijk, Maarten van Gompel
LREC3
2011 A Posteriori Agreement as a Quality Measure for Readability Prediction Systems
Philip van Oosten, Véronique Hoste, Dries Tanghe
CICLing (2)2
2010 Towards a Balanced Named Entity Corpus for Dutch
Bart Desmet, Véronique Hoste
LREC2
2010 Construction of a Benchmark Data Set for Cross-lingual Word Sense Disambiguation
Els Lefever, Véronique Hoste
LREC2
2010 Towards an Improved Methodology for Automated Readability Prediction
Philip van Oosten, Dries Tanghe, Véronique Hoste
LREC3
2010 Interacting Semantic Layers of Annotation in SoNaR, a Reference Corpus of Contemporary Written Dutch
Ineke Schuurman, Véronique Hoste, Paola Monachesi
LREC2
2010 Towards a Learning Approach for Abbreviation Detection and Resolution
Klaar Vanopstal, Bart Desmet, Véronique Hoste
LREC3
2010 Clustering web people search results using fuzzy ants
Els Lefever, Timur Fayruzov, Véronique Hoste, Martine De Cock
Inf. Sci.3
2009 Language-Independent Bilingual Terminology Extraction from a Multilingual Parallel Corpus
Els Lefever, Lieve Macken, Véronique Hoste
EACL3
2009 Linguistic feature analysis for protein interaction extraction
abstract
BACKGROUND: The rapid growth of the amount of publicly available reports on biomedical experimental results has recently caused a boost of text mining approaches for protein interaction extraction. Most approaches rely implicitly or explicitly on linguistic, i.e., lexical and syntactic, data extracted from text. However, only few attempts have been made to evaluate the contribution of the different feature types. In this work, we contribute to this evaluation by studying the relative importance of deep syntactic features, i.e., grammatical relations, shallow syntactic features (part-of-speech information) and lexical features. For this purpose, we use a recently proposed approach that uses support vector machines with structured kernels. RESULTS: Our results reveal that the contribution of the different feature types varies for the different data sets on which the experiments were conducted. The smaller the training corpus compared to the test data, the more important the role of grammatical relations becomes. Moreover, deep syntactic information based classifiers prove to be more robust on heterogeneous texts where no or only limited common vocabulary is shared. CONCLUSION: Our findings suggest that grammatical relations play an important role in the interaction extraction task. Moreover, the net advantage of adding lexical and shallow syntactic features is small related to the number of added features. This implies that efficient classifiers can be built by using only a small fraction of the features that are typically being used in recent approaches.
Timur Fayruzov, Martine De Cock, Chris Cornelis, Véronique Hoste
BMC Bioinform.4
2008 Semantic and Syntactic Features for Dutch Coreference Resolution
Iris Hendrickx, Véronique Hoste, Walter Daelemans
CICLing2
2008 Linguistically-Based Sub-Sentential Alignment for Terminology Extraction from a Bilingual Automotive Corpus
Lieve Macken, Els Lefever, Véronique Hoste
COLING3
2008 A Coreference Corpus and Resolution System for Dutch
Iris Hendrickx, Gosse Bouma, Frederik Coppens, Walter Daelemans, Véronique Hoste, Geert Kloosterman, Anne-Marie Mineur, Joeri Van Der Vloet, Jean-Luc Verschelde
LREC5
2008 Learning-based Detection of Scientific Terms in Patient Information
Véronique Hoste, Els Lefever, Klaar Vanopstal, Isabelle Delaere
LREC1
2006 KNACK-2002: a Richly Annotated Corpus of Dutch Written Text
Véronique Hoste, Guy De Pauw
LREC1
2004 Using rule-induction techniques to model pronunciation variation in Dutch
Véronique Hoste, Walter Daelemans, Steven Gillis
Comput. Speech Lang.1
2003 Learning to Predict Pitch Accents and Prosodic Boundaries in Dutch
abstract
We train a decision tree inducer (CART) and a memory-based classifier (MBL) on predicting prosodic pitch accents and breaks in Dutch text, on the basis of shallow, easy-to-compute features. We train the algorithms on both tasks individually and on the two tasks simultaneously. The parameters of both algorithms and the selection of features are optimized per task with iterative deepening, an efficient wrapper procedure that uses progressive sampling of training data. Results show a consistent significant advantage of MBL over CART, and also indicate that task combination can be done at the cost of little generalization score loss. Tests on cross-validated data and on held-out data yield F-scores of MBL on accent placement of 84 and 87, respectively, and on breaks of 88 and 91, respectively. Accent placement is shown to outperform an informed baseline rule; reliably predicting breaks other than those already indicated by intra-sentential punctuation, however, appears to be more challenging.
Erwin Marsi, Martin Reynaert, Antal van den Bosch, Walter Daelemans, Véronique Hoste
ACL5
2003 Combined Optimization of Feature Selection and Algorithm Parameters in Machine Learning of Language
Walter Daelemans, Véronique Hoste, Fien De Meulder, Bart Naudts
ECML2
2002 Combining information sources for memory-based pitch accent placement
abstract
We describe results on pitch accent placement in Dutch text obtained with a memory-based learning approach. The training material consists of newspaper texts that have been prosodically annotated by humans, and subsequently enriched with linguistic features and informational metrics using generally available, lowcost, shallow, knowledge-poor tools. We report on the effects of context-modelling and the nearest neighbours parameter (k), and show the advantage of combining features of a different nature, where the best performance yields a cross-validated F-score of 82. Evaluation on an independent test corpus shows that our approach outperforms existing TTS systems for Dutch. 1.
Erwin Marsi, Bertjan Busser, Walter Daelemans, Véronique Hoste, Martin Reynaert, Antal van den Bosch
INTERSPEECH4
2002 Evaluation of Machine Learning Methods for Natural Language Processing Tasks
Walter Daelemans, Véronique Hoste
LREC2
2002 Parameter optimization for machine-learning of word sense disambiguation
abstract
Various Machine Learning (ML) approaches have been demonstrated to produce relatively successful Word Sense Disambiguation (WSD) systems. There are still unexplained differences among the performance measurements of different algorithms, hence it is warranted to deepen the investigation into which algorithm has the right ‘bias’ for this task. In this paper, we show that this is not easy to accomplish, due to intricate interactions between information sources, parameter settings, and properties of the training data. We investigate the impact of parameter optimization on generalization accuracy in a memory-based learning approach to English and Dutch WSD. A ‘word-expert’ architecture was adopted, yielding a set of classifiers, each specialized in one single wordform. The experts consist of multiple memory-based learning classifiers, each taking different information sources as input, combined in a voting scheme. We optimized the architectural and parametric settings for each individual word-expert by performing cross-validation experiments on the learning material. The results of these experiments show that the variation of both the algorithmic parameters and the information sources available to the classifiers leads to large fluctuations in accuracy. We demonstrate that optimization per word-expert leads to an overall significant improvement in the generalization accuracies of the produced WSD systems.
Véronique Hoste, Iris Hendrickx, Walter Daelemans, Antal van den Bosch
Nat. Lang. Eng.1
2000 A Rule Induction Approach to Modeling Regional Pronunciation Variation
Véronique Hoste, Steven Gillis, Walter Daelemans
COLING1
2000 Meta-Learning for Phonemic Annotation of Corpora
Véronique Hoste, Walter Daelemans, Erik F. Tjong Kim Sang, Steven Gillis
ICML1