EDBT 2026 Demo / reviewers in the wild / expert
Thomas Demeester
dblp:39/10077
· DBLP profile ↗
53ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0002-9901-5768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 4 first-author · 21 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling Clinical Uncertainty in Radiology Reports: From Explicit Uncertainty Markers to Implicit Reasoning PathwaysabstractRadiology reports are invaluable for clinical decision-making and hold great potential for automated analysis when structured into machine-readable formats. These reports often contain uncertainty, which we categorize into two distinct types: (i) Explicit uncertainty reflects doubt about the presence or absence of findings, conveyed through hedging phrases. These vary in meaning depending on the context, making rule-based systems insufficient to quantify the level of uncertainty for specific findings; (ii) Implicit uncertainty arises when radiologists omit parts of their reasoning, recording only key findings or diagnoses. Here, it is often unclear whether omitted findings are truly absent or simply unmentioned for brevity. We address these challenges with a two-part framework. We quantify explicit uncertainty by creating an expert-validated, LLM-based reference ranking of common hedging phrases, and mapping each finding to a probability value based on this reference. In addition, we model implicit uncertainty through an expansion framework that systematically adds characteristic sub-findings derived from expert-defined diagnostic pathways for 14 common diagnoses. Using these methods, we release Lunguage++, an expanded, uncertainty-aware version of the Lunguage benchmark of fine-grained structured radiology reports. This enriched resource enables uncertainty-aware image classification, faithful diagnostic reasoning, and new investigations into the clinical impact of diagnostic uncertainty. Paloma Rabaey, Jong Hak Moon, Jung-Oh Lee, Min Gwan Kim, Hangyul Yoon, Thomas Demeester, Edward Choi 0003 |
LREC | 6 |
| 2026 | Patient-level information extraction by consistent integration of textual and tabular evidence with Bayesian networksabstractAbstract Electronic health records (EHRs) form an invaluable resource for training clinical decision support systems. To leverage the potential of such systems in high-risk applications, we need large, structured tabular datasets on which we can build transparent feature-based models. While part of the EHR already contains structured information (e.g. diagnosis codes, medications, and lab results), much of the information is contained within unstructured text (e.g. discharge summaries and nursing notes). In this work, we propose a method for multi-modal patient-level information extraction that leverages both the tabular features available in the patient’s EHR (using an expert-informed Bayesian network) as well as clinical notes describing the patient’s symptoms (using neural text classifiers). We propose the use of virtual evidence augmented with a consistency node to provide an interpretable, probabilistic fusion of the models’ predictions. The consistency node improves the calibration of the final predictions compared to virtual evidence alone, allowing the Bayesian network to better adjust the neural classifier’s output to handle missing information and resolve contradictions between the tabular and text data. We show the potential of our method on the SimSUM dataset, a simulated benchmark linking tabular EHRs with clinical notes through expert knowledge. Paloma Rabaey, Adrick Tench, Stefan Heytens, Thomas Demeester |
Appl. Intell. | 4 |
| 2025 | Improved Allergy Wheal Detection for the Skin Prick Automated Test Device
Rembert Daems, Sven Seys, Valérie Hox, Adam Chaker, Glynnis De Greve, Winde Lemmens, Anne-Lise Poirrier, Eline Beckers, Zuzana Diamant, Carmen Dierickx, Peter W. Hellings, Caroline Huart, Claudia Jerin, Mark Jorissen, Hanne Oscé, Karolien Roux, Sophie Tombu, Saartje Uyttebroek, Andrzej Zarowski, Senne Gorris, Laura Van Gerven, Dirk Loeckx, Thomas Demeester |
AIME (2) | 24 |
| 2025 | Leveraging Artificial Intelligence as a Decision Support System in Belgian Commercial CourtsabstractIn Belgium, each commercial court has at least one Chamber for Companies in Difficulties (CCD), tasked with the early detection and investigation of financially distressed companies. While the CCD’s primary objective is to promote recovery by encouraging financially distressed companies to take action, it also has a regulatory function by facilitating the orderly removal of companies that lack the willingness or capacity to recover. To identify companies, CCDs rely on a database of red flags. Based on these red flags, CCDs can decide to open files, investigate companies, and act accordingly. Given the economic stakes and the resource-intensive nature of the current manual process, we are working on a pilot project at the CCD of Antwerp to develop an AI-based judicial decision support system to assist CCD judges in selecting and prioritizing cases. Stijn Van Ruymbeke, Aruna Audenaert, Henri Arno, Tibe Habils, Joke Baeck, Klaas Mulier, Thomas Demeester |
ICAIL | 7 |
| 2025 | Dynamic Negative Guidance of Diffusion ModelsabstractNegative Prompting (NP) is widely utilized in diffusion models, particularly in text-to-image applications, to prevent the generation of undesired features. In this paper, we show that conventional NP is limited by the assumption of a constant guidance scale, which may lead to highly suboptimal results, or even complete failure, due to the non-stationarity and state-dependence of the reverse process. Based on this analysis, we derive a principled technique called ***D**ynamic **N**egative **G**uidance*, which relies on a near-optimal time and state dependent modulation of the guidance without requiring additional training. Unlike NP, negative guidance requires estimating the posterior class probability during the denoising process, which is achieved with limited additional computational overhead by tracking the discrete Markov Chain during the generative process. We evaluate the performance of DNG class-removal on MNIST and CIFAR10, where we show that DNG leads to higher safety, preservation of class balance and image quality when compared with baseline methods. Furthermore, we show that it is possible to use DNG with Stable Diffusion to obtain more accurate and less invasive guidance than NP. Felix Koulischer, Johannes Deleu, Gabriel Raya, Thomas Demeester, Luca Ambrogioni |
ICLR | 4 |
| 2025 | Feedback Guidance of Diffusion ModelsabstractWhile Classifier-Free Guidance (CFG) has become standard for improving sample fidelity in conditional diffusion models, it can harm diversity and induce memorization by applying constant guidance regardless of whether a particular sample needs correction. We propose **F**eed**B**ack **G**uidance (FBG), which uses a state-dependent coefficient to self-regulate guidance amounts based on need. Our approach is derived from first principles by assuming the learned conditional distribution is linearly corrupted by the unconditional distribution, contrasting with CFG's implicit multiplicative assumption. Our scheme relies on feedback of its own predictions about the conditional signal informativeness to adapt guidance dynamically during inference, challenging the view of guidance as a fixed hyperparameter. The approach is benchmarked on ImageNet512x512, where it significantly outperforms Classifier-Free Guidance and is competitive to Limited Interval Guidance (LIG) while benefitting from a strong mathematical framework. On Text-To-Image generation, we demonstrate that, as anticipated, our approach automatically applies higher guidance scales for complex prompts than for simpler ones and that it can be easily combined with existing guidance schemes such as CFG or LIG. Felix Koulischer, Florian Handke, Johannes Deleu, Thomas Demeester, Luca Ambrogioni |
NeurIPS | 4 |
| 2025 | Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in AlignmentabstractAbstract Large Language Models (LLMs) are often aligned using contrastive alignment objectives and preference pair datasets. The interaction between model, paired data, and objective makes alignment a complicated procedure, sometimes producing subpar results. We study this and find that (i) preference data gives a better learning signal when the underlying responses are contrastive, and (ii) alignment objectives lead to better performance when they specify more control over the model during training. Based on these insights, we introduce Contrastive Learning from AI Revisions (CLAIR), a data-creation method which leads to more contrastive preference pairs, and Anchored Preference Optimization (APO), a controllable and more stable alignment objective. We align Llama-3-8B-Instruct using various comparable datasets and alignment objectives and measure MixEval-Hard scores, which correlate highly with human judgments. The CLAIR preferences lead to the strongest performance out of all datasets, and APO consistently outperforms less controllable objectives. Our best model, trained on 32K CLAIR preferences with APO, improves Llama-3-8B-Instruct by 7.65%, closing the gap with GPT4-turbo by 45%. Our code and datasets are available. Karel D'Oosterlinck, Winnie Xu, Chris Develder, Thomas Demeester, Amanpreet Singh, Christopher Potts, Douwe Kiela, Shikib Mehri |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Integrating Visual Context Into Language Models for Situated Social Conversation StartersabstractEmbodied conversational agents that interact socially with people in the physical world require multi-modal capabilities, such as appropriately responding to visual features of users. While existing vision-and-language models can generate language based on visual input, this language is not situated in a social interaction in the physical world. We present a novel task called Visual Conversation Starters, where an agent generates a conversation-starting question referring to features visible in an image of the user. We collect a dataset of 4000 images of people with 12000 crowdsourced conversation starters, compare various model architectures: fine-tuning smaller seq2seq or image-to-text models versus zero-shot prompting of GPT-3.5, using image captions versus end-to-end image input, training on human data versus synthetic questions generated by GPT-3.5. Models were used to generate friendly conversation starters which were evaluated on criteria including language fluency, visual grounding, interestingness, politeness. Results show that GPT-3.5 generates more interesting, polite questions than smaller models that are fine-tuned on crowdsourced data, but vision-to-language models are better at referencing visual features, they can mimick GPT-3.5's performance. This demonstrates the feasibility of deep visiolinguistic models for situated social agents, forming an important first stage in creating situated multimodal social interaction. Ruben Janssens, Pieter Wolfert, Thomas Demeester, Tony Belpaeme |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Clinical Reasoning over Tabular Data and Text with Bayesian Networks
Paloma Rabaey, Johannes Deleu, Stefan Heytens, Thomas Demeester |
AIME (1) | 4 |
| 2024 | Debiasing Synthetic Data Generated by Deep Generative ModelsabstractWhile synthetic data hold great promise for privacy protection, their statistical analysis poses significant challenges that necessitate innovative solutions. The use of deep generative models (DGMs) for synthetic data generation is known to induce considerable bias and imprecision into synthetic data analyses, compromising their inferential utility as opposed to original data analyses. This bias and uncertainty can be substantial enough to impede statistical convergence rates, even in seemingly straightforward analyses like mean calculation. The standard errors of such estimators then exhibit slower shrinkage with sample size than the typical 1 over root-$n$ rate. This complicates fundamental calculations like p-values and confidence intervals, with no straightforward remedy currently available. In response to these challenges, we propose a new strategy that targets synthetic data created by DGMs for specific data analyses. Drawing insights from debiased and targeted machine learning, our approach accounts for biases, enhances convergence rates, and facilitates the calculation of estimators with easily approximated large sample variances. We exemplify our proposal through a simulation study on toy data and two case studies on real-world data, highlighting the importance of tailoring DGMs for targeted data analysis. This debiasing strategy contributes to advancing the reliability and applicability of synthetic data in statistical inference. Alexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Johan Decruyenaere, Christiaan Polet, Thomas Demeester, Stijn Vansteelandt |
NeurIPS | 6 |
| 2024 | The Real Deal Behind the Artificial Appeal: Inferential Utility of Tabular Synthetic DataabstractRecent advances in generative models facilitate the creation of synthetic data to be made available for research in privacy-sensitive contexts. However, the analysis of synthetic data raises a unique set of methodological challenges. In this work, we highlight the importance of inferential utility and provide empirical evidence against naive inference from synthetic data, whereby synthetic data are treated as if they were actually observed. Before publishing synthetic data, it is essential to develop statistical inference tools for such data. By means of a simulation study, we show that the rate of false-positive findings (type 1 error) will be unacceptably high, even when the estimates are unbiased. Despite the use of a previously proposed correction factor, this problem persists for deep generative models, in part due to slower convergence of estimators and resulting underestimation of the true standard error. We further demonstrate our findings through a case study. Alexander Decruyenaere, Heidelinde Dehaene, Paloma Rabaey, Christiaan Polet, Johan Decruyenaere, Stijn Vansteelandt, Thomas Demeester |
UAI | 7 |
| 2024 | Few-shot out-of-scope intent classification: analyzing the robustness of prompt-based learning
Maarten De Raedt, Johannes Deleu, Thomas Demeester, Chris Develder |
Appl. Intell. | 4 |
| 2024 | Revisiting clustering for efficient unsupervised dialogue structure inductionabstractAbstract In the development of a task-oriented dialogue system, defining the dialogue structure is a time-consuming task. Hence, several works have looked into automatically inferring it from data, e.g., actual conversations between a customer and a support agent. To recover such dialogue structure, recent methods based on discrete variational models learn to jointly encode and cluster utterances in dialogue states, but (i) represent utterances by only considering preceding dialogue context, and (ii) are slow to train since they are optimized with a compute-expensive decoding objective. We revisit and improve upon an existing efficient pipeline approach, commonly adopted as a baseline, that first encodes utterances and then clusters them with k-means to induce the dialogue structure. However, the existing approach represents utterances as bag-of-words or skip-thought vectors, which have been shown to perform poorly in semantic similarity tasks, and without considering dialogue context. We therefore first investigate the use of more powerful transformer-based encoders for encoding utterances. Next, we propose ellodar, a method for learning representations that capture both preceding and subsequent dialogue context, inspired by word-to-vec training strategies. ellodar is efficient since representations are learned directly in the encoding space by finetuning just a single linear layer on top of a frozen sentence encoder with a vector-to-vector regression training objective. Extensive experiments on representative datasets for dialogue structure induction (SimDial, Schema Guided Dialogues, DSTC2, and CamRest676) demonstrate that in terms of effectiveness to induce the correct dialogue structure, (i) clustering utterances represented by transformed-based encoders improves recent joint models by 13%–32% on standard cluster metrics, and (ii) clustering ellodar’s representations yields additional improvements ranging from +20% to +26%, with speedups of $$\times $$ × $$\textbf{10}$$ 10 – $$\textbf{10}^{\textbf{4}}$$ 10 4 compared to the recent joint models. Maarten De Raedt, Fréderic Godin, Chris Develder, Thomas Demeester |
Appl. Intell. | 4 |
| 2024 | BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insightsabstractOBJECTIVE: In this study, we investigate the potential of large language models (LLMs) to complement biomedical knowledge graphs in the training of semantic models for the biomedical and clinical domains. MATERIALS AND METHODS: Drawing on the wealth of the Unified Medical Language System knowledge graph and harnessing cutting-edge LLMs, we propose a new state-of-the-art approach for obtaining high-fidelity representations of biomedical concepts and sentences, consisting of 3 steps: an improved contrastive learning phase, a novel self-distillation phase, and a weight averaging phase. RESULTS: Through rigorous evaluations of diverse downstream tasks, we demonstrate consistent and substantial improvements over the previous state of the art for semantic textual similarity (STS), biomedical concept representation (BCR), and clinically named entity linking, across 15+ datasets. Besides our new state-of-the-art biomedical model for English, we also distill and release a multilingual model compatible with 50+ languages and finetuned on 7 European languages. DISCUSSION: Many clinical pipelines can benefit from our latest models. Our new multilingual model enables a range of languages to benefit from our advancements in biomedical semantic representation learning, opening a new avenue for bioinformatics researchers around the world. As a result, we hope to see BioLORD-2023 becoming a precious tool for future biomedical applications. CONCLUSION: In this article, we introduced BioLORD-2023, a state-of-the-art model for STS and BCR designed for the clinical domain. François Remy, Kris Demuynck, Thomas Demeester |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | CookDial: a dataset for task-oriented dialogs grounded in procedural documents
Klim Zaporojets, Johannes Deleu, Thomas Demeester, Chris Develder |
Appl. Intell. | 4 |
| 2022 | Robustifying Sentiment Classification by Maximally Exploiting Few CounterfactualsabstractFor text classification tasks, finetuned language models perform remarkably well.Yet, they tend to rely on spurious patterns in training data, thus limiting their performance on outof-distribution (OOD) test data.Among recent models aiming to avoid this spurious pattern problem, adding extra counterfactual samples to the training data has proven to be very effective.Yet, counterfactual data generation is costly since it relies on human annotation.Thus, we propose a novel solution that only requires annotation of a small fraction (e.g., 1%) of the original training data, and uses automatic generation of extra counterfactuals in an encoding vector space.We demonstrate the effectiveness of our approach in sentiment classification, using IMDb data for training and other sets for OOD tests (i.e., Amazon, SemEval and Yelp).We achieve noticeable accuracy improvements by adding only 1% manual counterfactuals: +3% compared to adding +100% in-distribution training samples, +1.3% compared to alternate counterfactual approaches. Maarten De Raedt, Fréderic Godin, Chris Develder, Thomas Demeester |
EMNLP | 4 |
| 2022 | "Cool glasses, where did you get them?": Generating Visually Grounded Conversation Starters for Human-Robot DialogueabstractVisually situated language interaction is an important challenge in multi-modal Human-Robot Interaction (URI). In this context we present a data-driven method to generate situated conversation starters based on visual context. We take visual data about the interactants and generate appropriate greetings for conversational agents in the context of HRI. For this, we constructed a novel open-source data set consisting of 4000 URI-oriented images of people facing the camera, each augmented by three conversation-starting questions. We compared a baseline retrieval-based model and a generative model. Human evaluation of the models using crowdsourcing shows that the generative model scores best, specifically at correctly referencing visual features. We also investigated how automated metrics can be used as a proxy for human evaluation and found that common automated metrics are a poor substitute for human judgement. Finally, we provide a proof-of-concept demonstrator through an interaction with a Furhat social robot. Ruben Janssens, Pieter Wolfert, Thomas Demeester, Tony Belpaeme |
HRI | 3 |
| 2022 | TempEL: Linking Dynamically Evolving and Newly Emerging EntitiesabstractIn our continuously evolving world, entities change over time and new, previously non-existing or unknown, entities appear. We study how this evolutionary scenario impacts the performance on a well established entity linking (EL) task. For that study, we introduce TempEL, an entity linking dataset that consists of time-stratified English Wikipedia snapshots from 2013 to 2022, from which we collect both anchor mentions of entities, and these target entities’ descriptions. By capturing such temporal aspects, our newly introduced TempEL resource contrasts with currently existing entity linking datasets, which are composed of fixed mentions linked to a single static version of a target Knowledge Base (e.g., Wikipedia 2010 for CoNLL-AIDA). Indeed, for each of our collected temporal snapshots, TempEL contains links to entities that are continual, i.e., occur in all of the years, as well as completely new entities that appear for the first time at some point. Thus, we enable to quantify the performance of current state-of-the-art EL models for: (i) entities that are subject to changes over time in their Knowledge Base descriptions as well as their mentions’ contexts, and (ii) newly created entities that were previously non-existing (e.g., at the time the EL model was trained). Our experimental results show that in terms of temporal performance degradation, (i) continual entities suffer a decrease of up to 3.1% EL accuracy, while (ii) for new entities this accuracy drop is up to 17.9%. This highlights the challenge of the introduced TempEL dataset and opens new research prospects in the area of time-evolving entity disambiguation. Klim Zaporojets, Lucie-Aimée Kaffee, Johannes Deleu, Thomas Demeester, Chris Develder, Isabelle Augenstein |
NeurIPS | 4 |
| 2021 | A Simple Geometric Method for Cross-Lingual Linguistic Transformations with Pre-trained AutoencodersabstractPowerful sentence encoders trained for multiple languages are on the rise.These systems are capable of embedding a wide range of linguistic properties into vector representations.While explicit probing tasks can be used to verify the presence of specific linguistic properties, it is unclear whether the vector representations can be manipulated to indirectly steer such properties.For efficient learning, we investigate the use of a geometric mapping in embedding space to transform linguistic properties, without any tuning of the pre-trained sentence encoder or decoder.We validate our approach on three linguistic properties using a pre-trained multilingual autoencoder and analyze the results in both monolingual and crosslingual settings. Maarten De Raedt, Fréderic Godin, Pieter Buteneers, Chris Develder, Thomas Demeester |
EMNLP (1) | 5 |
| 2021 | A Million Tweets Are Worth a Few Points: Tuning Transformers for Customer Service TasksabstractAmir Hadifar, Sofie Labat, Veronique Hoste, Chris Develder, Thomas Demeester. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Amir Hadifar, Sofie Labat, Véronique Hoste, Chris Develder, Thomas Demeester |
NAACL-HLT | 5 |
| 2021 | Neural probabilistic logic programming in DeepProbLog
Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt |
Artif. Intell. | 4 |
| 2021 | Overly optimistic prediction results on imbalanced data: a case study of flaws and benefits when applying over-sampling
Gilles Vandewiele, Isabelle Dehaene, György Kovács 0002, Lucas Sterckx, Olivier Janssens, Femke Ongenae, Femke De Backere, Filip De Turck, Kristien Roelens, Johan Decruyenaere, Sofie Van Hoecke, Thomas Demeester |
Artif. Intell. Medicine | 12 |
| 2021 | Solving arithmetic word problems by scoring equations with recursive neural networks
Klim Zaporojets, Giannis Bekoulis, Johannes Deleu, Thomas Demeester, Chris Develder |
Expert Syst. Appl. | 4 |
| 2021 | DWIE: An entity-centric dataset for multi-task document-level information extraction
Klim Zaporojets, Johannes Deleu, Chris Develder, Thomas Demeester |
Inf. Process. Manag. | 4 |
| 2021 | Exploration of block-wise dynamic sparseness
Amir Hadifar, Johannes Deleu, Chris Develder, Thomas Demeester |
Pattern Recognit. Lett. | 4 |
| 2020 | System Identification with Time-Aware Neural Sequence ModelsabstractEstablished recurrent neural networks are well-suited to solve a wide variety of prediction tasks involving discrete sequences. However, they do not perform as well in the task of dynamical system identification, when dealing with observations from continuous variables that are unevenly sampled in time, for example due to missing observations. We show how such neural sequence models can be adapted to deal with variable step sizes in a natural way. In particular, we introduce a ‘time-aware’ and stationary extension of existing models (including the Gated Recurrent Unit) that allows them to deal with unevenly sampled system observations by adapting to the observation times, while facilitating higher-order temporal behavior. We discuss the properties and demonstrate the validity of the proposed approach, based on samples from two industrial input/output processes. Thomas Demeester |
AAAI | 1 |
| 2020 | Clinical information extraction for preterm birth risk predictionabstractThis paper contributes to the pursuit of leveraging unstructured medical notes to structured clinical decision making. In particular, we present a pipeline for clinical information extraction from medical notes related to preterm birth, and discuss the main challenges as well as its potential for clinical practice. A large collection of medical notes, created by staff during hospitalizations of patients who were at risk of delivering preterm, was gathered and analyzed. Based on an annotated collection of notes, we trained and evaluated information extraction components to discover clinical entities such as symptoms, events, anatomical sites and procedures, as well as attributes linked to these clinical entities. In a retrospective study, we show that these are highly informative for clinical decision support models that are trained to predict whether delivery is likely to occur within specific time windows, in combination with structured information from electronic health records. Lucas Sterckx, Gilles Vandewiele, Isabelle Dehaene, Olivier Janssens, Femke Ongenae, Femke De Backere, Filip De Turck, Kristien Roelens, Johan Decruyenaere, Sofie Van Hoecke, Thomas Demeester |
J. Biomed. Informatics | 11 |
| 2019 | Time-to-Birth Prediction Models and the Influence of Expert Opinions
Gilles Vandewiele, Isabelle Dehaene, Olivier Janssens, Femke Ongenae, Femke De Backere, Filip De Turck, Kristien Roelens, Sofie Van Hoecke, Thomas Demeester |
AIME | 9 |
| 2019 | A Critical Look at Studies Applying Over-Sampling on the TPEHGDB Dataset
Gilles Vandewiele, Isabelle Dehaene, Olivier Janssens, Femke Ongenae, Femke De Backere, Filip De Turck, Kristien Roelens, Sofie Van Hoecke, Thomas Demeester |
AIME | 9 |
| 2019 | Character-level recurrent neural networks in practice: comparing training and sampling schemes
Cedric De Boom, Thomas Demeester, Bart Dhoedt |
Neural Comput. Appl. | 2 |
| 2018 | Predefined Sparseness in Recurrent Sequence ModelsabstractInducing sparseness while training neural networks has been shown to yield models with a lower memory footprint but similar effectiveness to dense models.However, sparseness is typically induced starting from a dense model, and thus this advantage does not hold during training.We propose techniques to enforce sparseness upfront in recurrent sequence models for NLP applications, to also benefit training.First, in language modeling, we show how to increase hidden state sizes in recurrent layers without increasing the number of parameters, leading to more expressive models.Second, for sequence labeling, we show that word embeddings with predefined sparseness lead to similar performance as dense embeddings, at a fraction of the number of trainable parameters. Thomas Demeester, Johannes Deleu, Fréderic Godin, Chris Develder |
CoNLL | 1 |
| 2018 | Adversarial training for multi-context joint entity and relation extractionabstractAdversarial training (AT) is a regularization method that can be used to improve the robustness of neural network methods by adding small perturbations in the training data.We show how to use AT for the tasks of entity recognition and relation extraction.In particular, we demonstrate that applying AT to a general purpose baseline model for jointly extracting entities and relations, allows improving the state-of-the-art effectiveness on several datasets in different contexts (i.e., news, biomedical, and real estate data) and for different languages (English and Dutch). Giannis Bekoulis, Johannes Deleu, Thomas Demeester, Chris Develder |
EMNLP | 3 |
| 2018 | Explaining Character-Aware Neural Networks for Word-Level Prediction: Do They Discover Linguistic Rules?abstractCharacter-level features are currently used in different neural network-based natural language processing algorithms.However, little is known about the character-level patterns those models learn.Moreover, models are often compared only quantitatively while a qualitative analysis is missing.In this paper, we investigate which character-level patterns neural networks learn and if those patterns coincide with manually-defined word segmentations and annotations.To that end, we extend the contextual decomposition (Murdoch et al., 2018) technique to convolutional neural networks which allows us to compare convolutional neural networks and bidirectional long short-term memory networks.We evaluate and compare these models for the task of morphological tagging on three morphologically different languages and show that these models implicitly discover understandable linguistic rules. Fréderic Godin, Kris Demuynck, Joni Dambre, Wesley De Neve, Thomas Demeester |
EMNLP | 5 |
| 2018 | DeepProbLog: Neural Probabilistic Logic ProgrammingabstractWe introduce DeepProbLog, a probabilistic logic programming language that incorporates deep learning by means of neural predicates. We show how existing inference and learning techniques can be adapted for the new language. Our experiments demonstrate that DeepProbLog supports (i) both symbolic and subsymbolic representations and inference, (ii) program induction, (iii) probabilistic (logic) programming, and (iv) (deep) learning from examples. To the best of our knowledge, this work is the first to propose a framework where general-purpose neural networks and expressive probabilistic-logical modeling and reasoning are integrated in a way that exploits the full expressiveness and strengths of both worlds and can be trained end-to-end based on examples. Robin Manhaeve, Sebastijan Dumancic, Angelika Kimmig, Thomas Demeester, Luc De Raedt |
NeurIPS | 4 |
| 2018 | An attentive neural architecture for joint segmentation and parsing and its application to real estate ads
Giannis Bekoulis, Johannes Deleu, Thomas Demeester, Chris Develder |
Expert Syst. Appl. | 3 |
| 2018 | Joint entity recognition and relation extraction as a multi-head selection problem
Giannis Bekoulis, Johannes Deleu, Thomas Demeester, Chris Develder |
Expert Syst. Appl. | 3 |
| 2018 | Large-scale user modeling with recurrent neural networks for music discovery on multiple time scales
Cedric De Boom, Rohan Agrawal, Samantha Hansen, Esh Kumar, Romain Yon, Ching-Wei Chen, Thomas Demeester, Bart Dhoedt |
Multim. Tools Appl. | 7 |
| 2018 | Modeling and predicting the popularity of online news based on temporal and content-related features
Steven Van Canneyt, Philip Leroux, Bart Dhoedt, Thomas Demeester |
Multim. Tools Appl. | 4 |
| 2017 | Break it Down for Me: A Study in Automated Lyric AnnotationabstractComprehending lyrics, as found in songs and poems, can pose a challenge to human and machine readers alike. This motivates the need for systems that can understand the ambiguity and jargon found in such creative texts, and provide commentary to aid readers in reaching the correct interpretation. We introduce the task of automated lyric annotation (ALA). Like text simplification, a goal of ALA is to rephrase the original text in a more easily understandable manner. However, in ALA the system must often include additional information to clarify niche terminology and abstract concepts. To stimulate research on this task, we release a large collection of crowdsourced annotations for song lyrics. We analyze the performance of translation and retrieval models on this task, measuring performance with both automated and human evaluation. We find that each model captures a unique type of information important to the task. Lucas Sterckx, Jason Naradowsky, William J. Byrne, Thomas Demeester, Chris Develder |
EMNLP | 4 |
| 2017 | Adversarial Sets for Regularising Neural Link Predictors
Pasquale Minervini, Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001 |
UAI | 2 |
| 2016 | Lifted Rule Injection for Relation EmbeddingsabstractMethods based on representation learning currently hold the state-of-the-art in many natural language processing and knowledge base inference tasks.Yet, a major challenge is how to efficiently incorporate commonsense knowledge into such models.A recent approach regularizes relation and entity representations by propositionalization of first-order logic rules.However, propositionalization does not scale beyond domains with only few entities and rules.In this paper we present a highly efficient method for incorporating implication rules into distributed representations for automated knowledge base construction.We map entity-tuple embeddings into an approximately Boolean space and encourage a partial ordering over relation embeddings based on implication rules mined from WordNet.Surprisingly, we find that the strong restriction of the entity-tuple embedding space does not hurt the expressiveness of the model and even acts as a regularizer that improves generalization.By incorporating few commonsense rules, we achieve an increase of 2 percentage points mean average precision over a matrix factorization baseline, while observing a negligible increase in runtime. Thomas Demeester, Tim Rocktäschel, Sebastian Riedel 0001 |
EMNLP | 1 |
| 2016 | Supervised Keyphrase Extraction as Positive Unlabeled LearningabstractThe problem of noisy and unbalanced training data for supervised keyphrase extraction results from the subjectivity of keyphrase assignment, which we quantify by crowdsourcing keyphrases for news and fashion magazine articles with many annotators per document.We show that annotators exhibit substantial disagreement, meaning that single annotator data could lead to very different training sets for supervised keyphrase extractors.Thus, annotations from single authors or readers lead to noisy training data and poor extraction performance of the resulting supervised extractor.We provide a simple but effective solution to still work with such data by reweighting the importance of unlabeled candidate phrases in a two stage Positive Unlabeled Learning setting.We show that performance of trained keyphrase extractors approximates a classifier trained on articles labeled by multiple annotators, leading to higher average F 1 scores and better rankings of keyphrases.We apply this strategy to a variety of test collections from different backgrounds and show improvements over strong baseline models. Lucas Sterckx, Cornelia Caragea, Thomas Demeester, Chris Develder |
EMNLP | 3 |
| 2016 | An Automated End-To-End Pipeline for Fine-Grained Video Annotation using Deep Neural NetworksabstractThe searchability of video content is often limited to the descriptions authors and/or annotators care to provide. The level of description can range from absolutely nothing to fine-grained annotations at the level of frames. Based on these annotations, certain parts of the video content are more searchable than others. Baptist Vandersmissen, Lucas Sterckx, Thomas Demeester, Azarakhsh Jalalvand, Wesley De Neve, Rik Van de Walle |
ICMR | 3 |
| 2016 | Predicting relevance based on assessor disagreement: analysis and practical applications for search evaluation
Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen 0002, Chris Develder |
Inf. Retr. J. | 1 |
| 2016 | Knowledge base population using semantic label propagation
Lucas Sterckx, Thomas Demeester, Johannes Deleu, Chris Develder |
Knowl. Based Syst. | 2 |
| 2016 | Representation learning for very short texts using weighted word embedding aggregation
Cedric De Boom, Steven Van Canneyt, Thomas Demeester, Bart Dhoedt |
Pattern Recognit. Lett. | 3 |
| 2014 | Aligning Vertical Collection Relevance with User IntentabstractSelecting and aggregating different types of content from multiple vertical search engines is becoming popular in web search. The user vertical intent, the verticals the user expects to be relevant for a particular information need, might not correspond to the vertical collection relevance, the verticals containing the most relevant content. In this work we propose different approaches to define the set of relevant verticals based on document judgments. We correlate the collection-based relevant verticals obtained from these approaches to the real user vertical intent, and show that they can be aligned relatively well. The set of relevant verticals defined by those approaches could therefore serve as an approximate but reliable ground-truth for evaluating vertical selection, avoiding the need for collecting explicit user vertical intent, and vice versa. Ke Zhou 0003, Thomas Demeester, Dong Nguyen 0002, Djoerd Hiemstra, Dolf Trieschnigg |
CIKM | 2 |
| 2014 | Assessing Quality of Unsupervised Topics in Song Lyrics
Lucas Sterckx, Thomas Demeester, Johannes Deleu, Laurent Mertens, Chris Develder |
ECIR | 2 |
| 2014 | Exploiting user disagreement for web search evaluation: an experimental approachabstractTo express a more nuanced notion of relevance as compared to binary judgments, graded relevance levels can be used for the evaluation of search results. Especially in Web search, users strongly prefer top results over less relevant results, and yet they often disagree on which are the top results for a given information need. Whereas previous works have generally considered disagreement as a negative effect, this paper proposes a method to exploit this user disagreement by integrating it into the evaluation procedure. Thomas Demeester, Robin Aly, Djoerd Hiemstra, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder |
WSDM | 1 |
| 2014 | Probabilistic models in IR and their relationships
Robin Aly, Thomas Demeester, Stephen E. Robertson |
Inf. Retr. | 2 |
| 2013 | Snippet-Based Relevance Predictions for Federated Web Search
Thomas Demeester, Dong Nguyen 0002, Dolf Trieschnigg, Chris Develder, Djoerd Hiemstra |
ECIR | 1 |
| 2013 | Taily: shard selection using the tail of score distributionsabstractSearch engines can improve their efficiency by selecting only few promising shards for each query. State-of-the-art shard selection algorithms first query a central index of sampled documents, and their effectiveness is similar to searching all shards. However, the search in the central index also hurts efficiency. Additionally, we show that the effectiveness of these approaches varies substantially with the sampled documents. This paper proposes Taily, a novel shard selection algorithm that models a query's score distribution in each shard as a Gamma distribution and selects shards with highly scored documents in the tail of the distribution. Taily estimates the parameters of score distributions based on the mean and variance of the score function's features in the collections and shards. Because Taily operates on term statistics instead of document samples, it is efficient and has deterministic effectiveness. Experiments on large web collections (Gov2, CluewebA and CluewebB) show that Taily achieves similar effectiveness to sample-based approaches, and improves upon their efficiency by roughly 20% in terms of used resources and response time. Robin Aly, Djoerd Hiemstra, Thomas Demeester |
SIGIR | 3 |
| 2012 | Federated search in the wild: the combined power of over a hundred search enginesabstractFederated search has the potential of improving web search: the user becomes less dependent on a single search provider and parts of the deep web become available through a unified interface, leading to a wider variety in the retrieved search results. However, a publicly available dataset for federated search reflecting an actual web environment has been absent. As a result, it has been difficult to assess whether proposed systems are suitable for the web setting. We introduce a new test collection containing the results from more than a hundred actual search engines, ranging from large general web search engines such as Google and Bing to small domain-specific engines. We discuss the design and analyze the effect of several sampling methods. For a set of test queries, we collected relevance judgements for the top 10 results of each search engine. The dataset is publicly available and is useful for researchers interested in resource selection for web search collections, result merging and size estimation of uncooperative resources. Dong Nguyen 0002, Thomas Demeester, Dolf Trieschnigg, Djoerd Hiemstra |
CIKM | 2 |