EDBT 2026 Demo / reviewers in the wild / expert
Antal van den Bosch
dblp:b/AvdBosch · also Antal P. J. van den Bosch
· DBLP profile ↗
82ranked-venue papers
14as first author
8since 2021 · last 2026
0000-0003-2493-656XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 13 first-author · 8 since 2021Databases, data management, data science and information retrieval · 18 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Memorization or Lucky Guesses: Detecting Short Sequences from Copyrighted Dutch News in LLM Output
Joris Veerbeek, Kas Berendsen, Alessandra Polimeno, Antal van den Bosch |
LREC | 4 |
| 2024 | The Impact of Featuring Comments in Online Discussions
Cedric Waterschoot, Ernst van den Hemel, Antal van den Bosch |
ASONAM (2) | 3 |
| 2024 | Corpus Creation and Automatic Alignment of Historical Dutch Dialect SpeechabstractThe Dutch Dialect Database (also known as the ‘Nederlandse Dialectenbank’) contains dialectal variations of Dutch that were recorded all over the Netherlands in the second half of the twentieth century. A subset of these recordings of about 300 hours were enriched with manual orthographic transcriptions, using non-standard approximations of dialectal speech. In this paper we describe the creation of a corpus containing both the audio recordings and their corresponding transcriptions and focus on our method for aligning the recordings with the transcriptions and the metadata. Martijn Bentum, Eric Sanders, Antal van den Bosch, Douwe Zeldenrust, Henk van den Heuvel |
LREC/COLING | 3 |
| 2024 | Re-evaluating the Tomes for the TimesabstractLiterature is to some degree a snapshot of the time it was written in and the societal attitudes of the time. Not all depictions are pleasant or in-line with modern-day sensibilities; this becomes problematic when the prevalent depictions over a large body of work are negatively biased, leading to their normalisation. Many much-loved and much-read classics are set in periods of heightened social inequality: slavery, pre-womens’ rights movements, colonialism, etc. In this paper, we exploit known text co-occurrence metrics with respect to token-level level contexts to identify prevailing themes associated with known problematic descriptors. We see that prevalent, negative depictions are perpetuated by classic literature. We propose that such a methodology could form the basis of a system for making explicit such problematic associations, for interested parties: such as, sensitivity coordinators of publishing houses, library curators, or organisations concerned with social justice Ryan Brate, Marieke van Erp, Antal van den Bosch |
LREC/COLING | 3 |
| 2023 | Contextual Profiling of Charged Terms in Historical Newspapers
Ryan Brate, Marieke van Erp, Antal van den Bosch |
LDK | 3 |
| 2022 | Detecting Minority Arguments for Mutual Understanding: A Moderation Tool for the Online Climate Change DebateabstractModerating user comments and promoting healthy understanding is a challenging task, especially in the context of polarized topics such as climate change. We propose a moderation tool to assist moderators in promoting mutual understanding in regard to this topic. The approach is twofold. First, we train classifiers to label incoming posts for the arguments they entail, with a specific focus on minority arguments. We apply active learning to further supplement the training data with rare arguments. Second, we dive deeper into singular arguments and extract the lexical patterns that distinguish each argument from the others. Our findings indicate that climate change arguments form clearly separable clusters in the embedding space. These classes are characterized by their own unique lexical patterns that provide a quick insight in an argument’s key concepts. Additionally, supplementing our training data was necessary for our classifiers to be able to adequately recognize rare arguments. We argue that this detailed rundown of each argument provides insight into where others are coming from. These computational approaches can be part of the toolkit for content moderators and researchers struggling with polarized topics. Cedric Waterschoot, Ernst van den Hemel, Antal van den Bosch |
COLING | 3 |
| 2022 | Lexicon or grammar? Using memory-based learning to investigate the syntactic relationship between Belgian and Netherlandic DutchabstractAbstract This article builds on computational tools to investigate the syntactic relationship between the highly related European national varieties of Dutch, viz. Belgian Dutch (BD) and Netherlandic Dutch (ND). It reports on a series of memory-based learning analyses of the post-verbal distribution ofer“there” in adjunct-initial existential constructions likeOp het dak staat (er) een schoorsteen“On the roof (there) is a chimney,’, which has been claimed to be among the most notoriously difficult variables in Dutch. On the basis of balanced datasets extracted from Flemish and Dutch newspaper corpora, it is shown thater’s distribution in both national varieties can be learned to a considerable extent from bare lexical input which is not assigned to higher-level categories. However, whereas this yields good results for ND, BD scores are consistently lower, suggesting that BD cannot do with lexical features alone to attain accuracy scores comparable to ND. This ties in with earlier findings that the more advanced standardization of ND materializes in a higher lexical collocability, whereas Flemish speakers need additional higher-level linguistic information to inserter. Robbert De Troij, Stefan Grondelaers, Dirk Speelman, Antal van den Bosch |
Nat. Lang. Eng. | 4 |
| 2021 | Calculating Argument Diversity in Online ThreadsabstractWe propose a method for estimating argument diversity and interactivity in online discussion threads. Using a case study on the subject of Black Pete ("Zwarte Piet") in the Netherlands, the approach for automatic detection of echo chambers is presented. Dynamic thread scoring calculates the status of the discussion on the thread level, while individual messages receive a contribution score reflecting the extent to which the post contributed to the overall interactivity in the thread. We obtain platform-specific results. Gab hosts only echo chambers, while the majority of Reddit threads are balanced in terms of perspectives. Twitter threads cover the whole spectrum of interactivity. While the results based on the case study mirror previous research, this calculation is only the first step towards better understanding and automatic detection of echo effects in online discussions. Cedric Waterschoot, Antal van den Bosch, Ernst van den Hemel |
LDK | 2 |
| 2020 | Optimising Twitter-based Political Election Prediction with Relevance andSentiment FiltersabstractWe study the relation between the number of mentions of political parties in the last weeks before the elections and the election results. In this paper we focus on the Dutch elections of the parliament in 2012 and for the provinces (and the senate) in 2011 and 2015. With raw counts, without adaptations, we achieve a mean absolute error (MAE) of 2.71% for 2011, 2.02% for 2012 and 2.89% for 2015. A set of over 17,000 tweets containing political party names were annotated by at least three annotators per tweet on ten features denoting communicative intent (including the presence of sarcasm, the message’s polarity, the presence of an explicit voting endorsement or explicit voting advice, etc.). The annotations were used to create oracle (gold-standard) filters. Tweets with or without a certain majority annotation are held out from the tweet counts, with the goal of attaining lower MAEs. With a grid search we tested all combinations of filters and their responding MAE to find the best filter ensemble. It appeared that the filters show markedly different behaviour for the three elections and only a small MAE improvement is possible when optimizing on all three elections. Larger improvements for one election are possible, but result in deterioration of the MAE for the other elections. Eric Sanders, Antal van den Bosch |
LREC | 2 |
| 2020 | Uncovering the language of wine expertsabstractAbstract Talking about odors and flavors is difficult for most people, yet experts appear to be able to convey critical information about wines in their reviews. This seems to be a contradiction, and wine expert descriptions are frequently received with criticism. Here, we propose a method for probing the language of wine reviews, and thus offer a means to enhance current vocabularies, and as a by-product question the general assumption that wine reviews are gibberish. By means of two different quantitative analyses—support vector machines for classification and Termhood analysis—on a corpus of online wine reviews, we tested whether wine reviews are written in a consistent manner, and thus may be considered informative; and whether reviews feature domain-specific language. First, a classification paradigm was trained on wine reviews from one set of authors for which the color, grape variety, and origin of a wine were known, and subsequently tested on data from a new author. This analysis revealed that, regardless of individual differences in vocabulary preferences, color and grape variety were predicted with high accuracy. Second, using Termhood as a measure of how words are used in wine reviews in a domain-specific manner compared to other genres in English, a list of 146 wine-specific terms was uncovered. These words were compared to existing lists of wine vocabulary that are currently used to train experts. Some overlap was observed, but there were also gaps revealed in the extant lists, suggesting these lists could be improved by our automatic analysis. Ilja Croijmans, Iris Hendrickx, Els Lefever, Asifa Majid, Antal van den Bosch |
Nat. Lang. Eng. | 5 |
| 2020 | Query-based summarization of discussion threadsabstractAbstract In this paper, we address query-based summarization of discussion threads. New users can profit from the information shared in the forum, Please check if the inserted city and country names in the affiliations are correct. if they can find back the previously posted information. However, discussion threads on a single topic can easily comprise dozens or hundreds of individual posts. Our aim is to summarize forum threads given real web search queries. We created a data set with search queries from a discussion forum’s search engine log and the discussion threads that were clicked by the user who entered the query. For 120 thread–query combinations, a reference summary was made by five different human raters. We compared two methods for automatic summarization of the threads: a query-independent method based on post features, and Maximum Marginal Relevance (MMR), a method that takes the query into account. We also compared four different word embeddings representations as alternative for standard word vectors in extractive summarization. We find (1) that the agreement between human summarizers does not improve when a query is provided that: (2) the query-independent post features as well as a centroid-based baseline outperform MMR by a large margin; (3) combining the post features with query similarity gives a small improvement over the use of post features alone; and (4) for the word embeddings, a match in domain appears to be more important than corpus size and dimensionality. However, the differences between the models were not reflected by differences in quality of the summaries created with help of these models. We conclude that query-based summarization with web queries is challenging because the queries are short, and a click on a result is not a direct indicator for the relevance of the result. Suzan Verberne, Emiel Krahmer, Sander Wubben, Antal van den Bosch |
Nat. Lang. Eng. | 4 |
| 2019 | Listening with Great Expectations: An Investigation of Word Form Anticipations in Naturalistic SpeechabstractThe event-related potential (ERP) component named phonological mismatch negativity (PMN) arises when listeners hear an unexpected word form in a spoken sentence [1]. The PMN is thought to reflect the mismatch between expected and perceived auditory speech input. In this paper, we use the PMN to test a central premise in the predictive coding framework [2], namely that the mismatch between prior expectations and sensory input is an important mechanism of perception. We test this with natural speech materials containing approximately 50,000 word tokens. The corresponding EEG-signal was recorded while participants (n = 48) listened to these materials. Following [3], we quantify the mismatch with two word probability distributions (WPD): a WPD based on preceding context, and a WPD that is additionally updated based on the incoming audio of the current word. We use the between-WPD cross entropy for each word in the utterances and show that a higher cross entropy correlates with a more negative PMN. Our results show that listeners anticipate auditory input while processing each word in naturalistic speech. Moreover, complementing previous research, we show that predictive language processing occurs across the whole probability spectrum. Martijn Bentum, Louis ten Bosch, Antal van den Bosch, Mirjam Ernestus |
INTERSPEECH | 3 |
| 2019 | Quantifying Expectation Modulation in Human Speech ProcessingabstractThe mismatch between top-down predicted and bottom-up perceptual input is an important mechanism of perception according to the predictive coding framework (Friston, [1]).In this paper we develop and validate a new information-theoretic measure that quantifies the mismatch between expected and observed auditory input during speech processing.We argue that such a mismatch measure is useful for the study of speech processing.To compute the mismatch measure, we use naturalistic speech materials containing approximately 50,000 word tokens.For each word token we first estimate the prior word probability distribution with the aid of statistical language modelling, and next use automatic speech recognition to update this word probability distribution based on the unfolding speech signal.We validate the mismatch measure with multiple analyses, and show that the auditory-based update improves the probability of the correct word and lowers the uncertainty of the word probability distribution.Based on these results, we argue that it is possible to explicitly estimate the mismatch between predicted and perceived speech input with the cross entropy between word expectations computed before and after an auditory update. Martijn Bentum, Louis ten Bosch, Antal van den Bosch, Mirjam Ernestus |
INTERSPEECH | 3 |
| 2018 | Aspect-based summarization of pros and cons in unstructured product reviewsabstractWe developed three systems for generating pros and cons summaries of product reviews. Automating this task eases the writing of product reviews, and offers readers quick access to the most important information. We compared SynPat, a system based on syntactic phrases selected on the basis of valence scores, against a neural-network-based system trained to map bag-of-words representations of reviews directly to pros and cons, and the same neural system trained on clusters of word-embedding encodings of similar pros and cons. We evaluated the systems in two ways: first on held-out reviews with gold-standard pros and cons, and second by asking human annotators to rate the systems’ output on relevance and completeness. In the second evaluation, the gold-standard pros and cons were assessed along with the system output. We find that the human-generated summaries are not deemed as significantly more relevant or complete than the SynPat systems; the latter are scored higher than the human-generated summaries on a precision metric. The neural approaches yield a lower performance in the human assessment, and are outperformed by the baseline. Florian Kunneman, Sander Wubben, Antal van den Bosch, Emiel Krahmer |
COLING | 3 |
| 2018 | A Multilingual Wikified Data Set of Educational Material
Iris Hendrickx, Eirini Takoulidou, Thanasis Naskos, Katia Kermanidis, Vilelmini Sosoni, Hugo De Vos, Maria Stasimioti, Menno van Zaanen, Panayota Georgakopoulou, Valia Kordoni, Maja Popovic, Markus Egg, Antal van den Bosch |
LREC | 13 |
| 2018 | Discovering the Language of Wine Reviews: A Text Mining Account
Els Lefever, Iris Hendrickx, Ilja Croijmans, Antal van den Bosch, Asifa Majid |
LREC | 4 |
| 2017 | Automatic Summarization of Domain-specific Forum Threads: Collecting Reference DataabstractWe create and analyze two sets of reference summaries for discussion threads on a patient support forum: expert summaries and crowdsourced, non-expert summaries. Ideally, reference summaries for discussion forum threads are created by expert members of the forum community. When there are few or no expert members available, crowdsourcing the reference summaries is an alternative. In this paper we investigate whether domain-specific forum data requires the hiring of domain experts for creating reference summaries. We analyze the inter-rater agreement for both data-sets and we train summarization models using the two types of reference summaries. The inter-rater agreement in crowdsourced reference summaries is low, close to random, while domain experts achieve a considerably higher, fair, agreement. The trained models however are similar to each other. We conclude that it is possible to train an extractive summarization model on crowdsourced data that is similar to an expert model, even if the inter-rater agreement for the crowdsourced data is low. Suzan Verberne, Antal van den Bosch, Sander Wubben, Emiel Krahmer |
CHIIR | 2 |
| 2017 | Overview of the 4th HistoInformatics WorkshopabstractIn line with global trends, historical records are increasingly available in forms that computer can process. These ever expanding records (such as scanned books, large-scale corpora, academic papers, maps, photos, audios, videos)---either digitally born or reconstructed through digitization pipelines---are too big to be read or viewed manually. Historians, like other humanities researchers, have a keen interest in computational approaches to process and study digitized historical information for research, writing, and dissemination of historical knowledge. In Computer Science, experimental tools and methods are challenged to be validated regarding their relevance for real-world questions and applications. The HistoInformatics workshop series is focused on the challenges and opportunities of data-driven humanities and brings together scientists and scholars at the forefront of this emerging field, at the interface between History, Anthropology, Archaeology, Computer Science and associated disciplines as well as the cultural heritage sector. The 4th HistoInformatics Workshop was a half day workshop co-located with the 26th ACM International Conference on Information and Knowledge Management (CIKM 2017) in Singapore. Mohammed Hasanuzzaman, Gaël Dias, Adam Jatowt, Marten Düring, Antal van den Bosch |
CIKM | 5 |
| 2017 | Supporting Experts to Handle Tweet Collections About Significant Events
Ali Hurriyetoglu, Nelleke Oostdijk, Mustafa Erkan Basar, Antal van den Bosch |
NLDB | 4 |
| 2016 | Abstractive Compression of Captions with Attentive Recurrent Neural NetworksabstractIn this paper we introduce the task of abstractive caption or scene description compression.We describe a parallel dataset derived from the FLICKR30K and MSCOCO datasets.With this data we train an attention-based bidirectional LSTM recurrent neural network and compare the quality of its output to a Phrasebased Machine Translation (PBMT) model and a human generated short description.An extensive evaluation is done using automatic measures and human judgements.We show that the neural model outperforms the PBMT model.Additionally, we show that automatic measures are not very well suited for evaluating this text-to-text generation task. Sander Wubben, Emiel Krahmer, Antal van den Bosch, Suzan Verberne |
INLG | 3 |
| 2016 | Nederlab: Towards a Single Portal and Research Environment for Diachronic Dutch Text Corpora
Hennie Brugman, Martin Reynaert, Nicoline van der Sijs, René van Stipriaan, Erik F. Tjong Kim Sang, Antal van den Bosch |
LREC | 6 |
| 2016 | Enhancing Access to Online Education: Quality Machine Translation of MOOC Content
Valia Kordoni, Antal van den Bosch, Katia Kermanidis, Vilelmini Sosoni, Kostadin Cholakov, Iris Hendrickx, Matthias Huck, Andy Way |
LREC | 2 |
| 2016 | Can Tweets Predict TV Ratings?
Bridget Sommerdijk, Eric Sanders, Antal van den Bosch |
LREC | 3 |
| 2016 | Open-domain extraction of future events from TwitterabstractAbstract Explicit references on Twitter to future events can be leveraged to feed a fully automatic monitoring system of real-world events. We describe a system that extracts open-domain future events from the Twitter stream. It detects future time expressions and entity mentions in tweets, clusters tweets together that overlap in these mentions above certain thresholds, and summarizes these clusters into event descriptions that can be presented to users of the system. Terms for the event description are selected in an unsupervised fashion.1We evaluated the system on a month of Dutch tweets, by showing the top-250 ranked events found in this month to human annotators. Eighty per cent of the candidate events were indeed assessed as being an event by at least three out of four human annotators, while all four annotators regarded sixty-three per cent as a real event. An added component to complement event descriptions with additional terms was not assessed better than the original system, due to the occasional addition of redundant terms. Comparing the found events to gold-standard events from maintained calendars on the Web mentioned in at least five tweets, the system yields a recall-at-250 of 0.20 and a recall based on all retrieved events of 0.40. Florian Kunneman, Antal van den Bosch |
Nat. Lang. Eng. | 2 |
| 2016 | Human-inspired modulation frequency features for noise-robust ASRabstractThis paper investigates a computational model that combines a frontend based on an auditory model with an exemplar-based sparse coding procedure for estimating the posterior probabilities of sub-word units when processing noisified speech. Envelope modulation spectrogram (EMS) features are extracted using an auditory model which decomposes the envelopes of the outputs of a bank of gammatone filters into one lowpass and multiple bandpass components. Through a systematic analysis of the configuration of the modulation filterbank , we investigate how and why different configurations affect the posterior probabilities of sub-word units by measuring the recognition accuracy on a semantics-free speech recognition task. Our main finding is that representing speech signal dynamics by means of multiple bandpass filters typically improves recognition accuracy. This effect is particularly noticeable in very noisy conditions. In addition we find that to have maximum noise robustness, the bandpass filters should focus on low modulation frequencies . This reenforces our intuition that noise robustness can be increased by exploiting redundancy in those frequency channels which have long enough integration time not to suffer from envelope modulations that are solely due to noise. The ASR system we design based on these findings behaves more similar to human recognition of noisified digit strings than conventional ASR systems. Thanks to the relation between the modulation filterbank and procedures for computing dynamic acoustic features in conventional ASR systems, the finding can be used for improving the frontends in those systems. Sara Ahmadi, Bert Cranen, Lou Boves, Louis ten Bosch, Antal van den Bosch |
Speech Commun. | 5 |
| 2015 | TraMOOC: Translation for Massive Open Online Courses
Valia Kordoni, Kostadin Cholakov, Markus Egg, Andy Way, Lexi Birch, Katia Kermanidis, Vilelmini Sosoni, Dimitrios Tsoumakos, Antal van den Bosch, Iris Hendrickx, Michael Papadopoulos, Panayota Georgakopoulou, Maria Gialama, Menno van Zaanen, Ioana Buliga, Mitja Jermol, Davor Orlic |
EAMT | 9 |
| 2015 | Looking for Books in Social Media: An Analysis of Complex Search Requests
Marijn Koolen, Toine Bogers, Antal van den Bosch, Jaap Kamps |
ECIR | 3 |
| 2015 | Signaling sarcasm: From hyperbole to hashtagabstractTo avoid a sarcastic message being understood in its unintended literal meaning, in microtexts such as messages on Twitter.com sarcasm is often explicitly marked with a hashtag such as ‘#sarcasm’. We collected a training corpus of about 406 thousand Dutch tweets with hashtag synonyms denoting sarcasm. Assuming that the human labeling is correct (annotation of a sample indicates that about 90% of these tweets are indeed sarcastic), we train a machine learning classifier on the harvested examples, and apply it to a sample of a day’s stream of 2.25 million Dutch tweets. Of the 353 explicitly marked tweets on this day, we detect 309 (87%) with the hashtag removed. We annotate the top of the ranked list of tweets most likely to be sarcastic that do not have the explicit hashtag. 35% of the top-250 ranked tweets are indeed sarcastic. Analysis indicates that the use of hashtags reduces the further use of linguistic markers for signaling sarcasm, such as exclamations and intensifiers. We hypothesize that explicit markers such as hashtags are the digital extralinguistic equivalent of non-verbal expressions that people employ in live interaction when conveying sarcasm. Checking the consistency of our finding in a language from another language family, we observe that in French the hashtag ‘#sarcasme’ has a similar polarity switching function, be it to a lesser extent. Florian Kunneman, Christine Liebrecht, Margot van Mulken, Antal van den Bosch |
Inf. Process. Manag. | 4 |
| 2014 | Translation Assistance by Translation of L1 Fragments in an L2 ContextabstractIn this paper we present new research in translation assistance.We describe a system capable of translating native language (L1) fragments to foreign language (L2) fragments in an L2 context.Practical applications of this research can be framed in the context of second language learning.The type of translation assistance system under investigation here encourages language learners to write in their target language while allowing them to fall back to their native language in case the correct word or expression is not known.These code switches are subsequently translated to L2 given the L2 context.We study the feasibility of exploiting cross-lingual context to obtain high-quality translation suggestions that improve over statistical language modelling and word-sense disambiguation baselines.A classificationbased approach is presented that is indeed found to improve significantly over these baselines by making use of a contextual window spanning a small number of neighbouring words. Maarten van Gompel, Antal van den Bosch |
ACL (1) | 2 |
| 2014 | Using idiolects and sociolects to improve word predictionabstractIn this paper the word prediction system Soothsayer 1 is described.This system predicts what a user is going to write as he is keying it in.The main innovation of Soothsayer is that it not only uses idiolects, the language of one individual person, as its source of knowledge, but also sociolects, the language of the social circle around that person.We use Twitter for data collection and experimentation.The idiolect models are based on individual Twitter feeds, the sociolect models are based on the tweets of a particular person and the tweets of the people he often communicates with.The idea behind this is that people who often communicate start to talk alike; therefore the language of the friends of person x can be helpful in trying to predict what person x is going to say.This approach achieved the best results.For a number of users, more than 50% of the keystrokes could have been saved if they had used Soothsayer. Wessel Stoop, Antal van den Bosch |
EACL | 2 |
| 2014 | Creating and using large monolingual parallel corpora for sentential paraphrase generation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
LREC | 2 |
| 2014 | Automatic thematic classification of election manifestos
Suzan Verberne, Eva D'hondt, Antal van den Bosch, Maarten Marx |
Inf. Process. Manag. | 3 |
| 2014 | Peter Spyns and Jan Odijk (eds): Essential speech and language technology for Dutch: results by the STEVIN programme - Springer, 2013, ISBN: 978-3-642-30909-0, xvii + 413 pp
Antal van den Bosch |
Mach. Transl. | 1 |
| 2013 | On the assessment of expertise profilesabstractExpertise retrieval has attracted significant interest in the field of information retrieval. Expert finding has been studied extensively, with less attention going to the complementary task of expert profiling, that is, automatically identifying topics about which a person is knowledgeable. We describe a test collection for expert profiling in which expert users have self‐selected their knowledge areas. Motivated by the sparseness of this set of knowledge areas, we report on an assessment experiment in which academic experts judge a profile that has been automatically generated by state‐of‐the‐art expert‐profiling algorithms; optionally, experts can indicate a level of expertise for relevant areas. Experts may also give feedback on the quality of the system‐generated knowledge areas. We report on a content analysis of these comments and gain insights into what aspects of profiles matter to experts. We provide an error analysis of the system‐generated profiles, identifying factors that help explain why certain experts may be harder to profile than others. We also analyze the impact on evaluating expert‐profiling systems of using self‐selected versus judged system‐generated knowledge areas as ground truth; they rank systems somewhat differently but detect about the same amount of pairwise significant differences despite the fact that the judged system‐generated assessments are more sparse. Richard Berendsen, Maarten de Rijke, Krisztian Balog, Toine Bogers, Antal van den Bosch |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2012 | Sentence Simplification by Monolingual Machine Translation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
ACL (1) | 2 |
| 2012 | The effect of domain and text type on text prediction quality
Suzan Verberne, Antal van den Bosch, Helmer Strik, Lou Boves |
EACL | 2 |
| 2012 | DutchSemCor: Targeting the ideal sense-tagged corpus
Piek Vossen, Attila Görög, Rubén Izquierdo, Antal van den Bosch |
LREC | 4 |
| 2012 | The socialist network
Matje van de Camp, Antal van den Bosch |
Decis. Support Syst. | 2 |
| 2011 | Integrating source-language context into phrase-based statistical machine translation
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
Mach. Transl. | 3 |
| 2010 | Paraphrase Generation as Monolingual Translation: Data and Evaluation
Sander Wubben, Antal van den Bosch, Emiel Krahmer |
INLG | 2 |
| 2009 | A Constraint Satisfaction Approach to Machine Translation
Sander Canisius, Antal van den Bosch |
EAMT | 2 |
| 2009 | Dependency Relations as Source Context in Phrase-Based SMT
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
PACLIC | 3 |
| 2008 | Efficient context-sensitive word completion for mobile devicesabstractWord completion is a basic technology for reducing the effort involved in text entry on mobile devices and in augmentative communication devices, where efficiency and ease of use are needed, but where a low memory footprint is also required. Standard solutions compress a lexicon into a suffix tree with a small memory footprint and high retrieval speed. Keystroke savings, a measurable correlate of text entry effort gain, typically improve when the algorithm would also take into account the previous word; however, this comes at the cost of a large footprint. We develop two word completion algorithms that encode the previous word in the input. The first algorithm utilizes a character buffer that includes a fixed number of recent keystrokes, including those belonging to previous words. The second algorithm includes the complete previous word as an extra input feature. In simulation studies, the first algorithm yields marked improvements in keystroke savings, but has a large memory footprint. The second algorithm can be tuned by frequency thresholding to have a small footprint, and be less than one order of magnitude slower than the baseline system, while its keystroke savings improve over the baseline. Antal van den Bosch, Toine Bogers |
Mobile HCI | 1 |
| 2008 | Recommending scientific articles using citeulikeabstractWe describe the use of the social reference management website CiteULike for recommending scientific articles to users, based on their reference library. We test three different collaborative filtering algorithms, and find that user-based filtering performs best. A temporal analysis of the data indexed by CiteULike shows that it takes about two years for the cold-start problem to disappear and recommendation performance to improve. Toine Bogers, Antal van den Bosch |
RecSys | 2 |
| 2007 | Comparing and evaluating information retrieval algorithms for news recommendationabstractIn this paper, we argue that the performance of content-based news recommender systems has been hampered by using relatively old and simple matching algorithms. Using more current probabilistic retrieval algorithms results in significant performance boosts. We test our ideas on a test collection that we have made publicly available. We perform both binary and graded evaluation of our algorithms and argue for the need for more graded evaluation of content-based recommender systems. Toine Bogers, Antal van den Bosch |
RecSys | 2 |
| 2007 | Broad expertise retrieval in sparse data environmentsabstractExpertise retrieval has been largely unexplored on data other than the W3C collection. At the same time, many intranets of universities and other knowledge-intensive organisations offer examples of relatively small but clean multilingual expertise data, covering broad ranges of expertise areas. We first present two main expertise retrieval tasks, along with a set of baseline approaches based on generative language modeling, aimed at finding expertise relations between topics and people. For our experimental evaluation, we introduce (and release) a new test set based on a crawl of a university site. Using this test set, we conduct two series of experiments. The first is aimed at determining the effectiveness of baseline expertise retrieval methods applied to the new test set. The second is aimed at assessing refined models that exploit characteristic features of the new test set, such as the organizational structure of the university, and the hierarchical structure of the topics in the test set. Expertise retrieval models are shown to be robust with respect to environments smaller than the W3C collection, and current techniques appear to be generalizable to other settings. Krisztian Balog, Toine Bogers, Leif Azzopardi, Maarten de Rijke, Antal van den Bosch |
SIGIR | 5 |
| 2007 | Letter to the Editor
Walter Daelemans, Antal van den Bosch |
Comput. Linguistics | 2 |
| 2006 | Dependency Parsing by Inference over High-recall Dependency Predictions
Sander Canisius, Toine Bogers, Antal van den Bosch, Jeroen Geertzen, Erik F. Tjong Kim Sang |
CoNLL | 3 |
| 2006 | Authoritative Re-ranking of Search Results
Toine Bogers, Antal van den Bosch |
ECIR | 2 |
| 2006 | Transferring PoS-tagging and lemmatization tools from spoken to written Dutch corpus development
Antal van den Bosch, Ineke Schuurman, Vincent Vandeghinste |
LREC | 1 |
| 2006 | Identifying Named Entities in Text Databases from the Natural History Domain
Caroline Sporleder, Marieke van Erp, Tijn Porcelijn, Antal van den Bosch, Pim Arntzen |
LREC | 4 |
| 2006 | A Rule-Based Approach for Process Discovery: Dealing with Noise and Imbalance in Process Logs
Laura Maruster, A. J. M. M. Weijters, Wil M. P. van der Aalst, Antal van den Bosch |
Data Min. Knowl. Discov. | 4 |
| 2005 | Improving Sequence Segmentation Learning by Predicting Trigrams
Antal van den Bosch, Walter Daelemans |
CoNLL | 1 |
| 2005 | Applying Spelling Error Correction Techniques for Improving Semantic Role Labelling
Erik F. Tjong Kim Sang, Sander Canisius, Antal van den Bosch, Toine Bogers |
CoNLL | 3 |
| 2005 | Hybrid Algorithms with Instance-Based Classification
Iris Hendrickx, Antal van den Bosch |
ECML | 2 |
| 2004 | Memory-based semantic role labeling: Optimizing features, algorithm, and output
Antal van den Bosch, Sander Canisius, Walter Daelemans, Iris Hendrickx, Erik F. Tjong Kim Sang |
CoNLL | 1 |
| 2003 | Learning to Predict Pitch Accents and Prosodic Boundaries in DutchabstractWe train a decision tree inducer (CART) and a memory-based classifier (MBL) on predicting prosodic pitch accents and breaks in Dutch text, on the basis of shallow, easy-to-compute features. We train the algorithms on both tasks individually and on the two tasks simultaneously. The parameters of both algorithms and the selection of features are optimized per task with iterative deepening, an efficient wrapper procedure that uses progressive sampling of training data. Results show a consistent significant advantage of MBL over CART, and also indicate that task combination can be done at the cost of little generalization score loss. Tests on cross-validated data and on held-out data yield F-scores of MBL on accent placement of 84 and 87, respectively, and on breaks of 88 and 91, respectively. Accent placement is shown to outperform an informed baseline rule; reliably predicting breaks other than those already indicated by intra-sentential punctuation, however, appears to be more challenging. Erwin Marsi, Martin Reynaert, Antal van den Bosch, Walter Daelemans, Véronique Hoste |
ACL | 3 |
| 2003 | Memory-based one-step named-entity recognition: Effects of seed list features, classifier stacking, and unannotated data
Iris Hendrickx, Antal van den Bosch |
CoNLL | 2 |
| 2003 | Learning PP attachment for filtering prosodic phrasing
Olga van Herwijnen, Antal van den Bosch, Jacques M. B. Terken, Erwin Marsi |
EACL | 2 |
| 2002 | Shallow Parsing on the Basis of Words Only: A Case StudyabstractWe describe a case study in which a memory-based learning algorithm is trained to simultaneously chunk sentences and assign grammatical function tags to these chunks. We compare the algorithm's performance on this parsing task with varying training set sizes (yielding learning curves) and different input representations. In particular we compare input consisting of words only, a variant that includes word form information for low-frequency words, gold-standard POS only, and combinations of these. The word-based shallow parser displays an apparently log-linear increase in performance, and surpasses the flatter POS-based curve at about 50,000 sentences of training data. The low-frequency variant performs even better, and the combinations is best. Comparative experiments with a real POS tagger produce lower results. We argue that we might not need an explicit intermediate POS-tagging step for parsing when a sufficient amount of training material is available and word form information is used for low-frequency words. Antal van den Bosch, Sabine Buchholz |
ACL | 1 |
| 2002 | Process Mining: Discovering Direct Successors in Process Logs
Laura Maruster, A. J. M. M. Weijters, Wil M. P. van der Aalst, Antal van den Bosch |
Discovery Science | 4 |
| 2002 | Combining information sources for memory-based pitch accent placementabstractWe describe results on pitch accent placement in Dutch text obtained with a memory-based learning approach. The training material consists of newspaper texts that have been prosodically annotated by humans, and subsequently enriched with linguistic features and informational metrics using generally available, lowcost, shallow, knowledge-poor tools. We report on the effects of context-modelling and the nearest neighbours parameter (k), and show the advantage of combining features of a different nature, where the best performance yields a cross-validated F-score of 82. Evaluation on an independent test corpus shows that our approach outperforms existing TTS systems for Dutch. 1. Erwin Marsi, Bertjan Busser, Walter Daelemans, Véronique Hoste, Martin Reynaert, Antal van den Bosch |
INTERSPEECH | 6 |
| 2002 | Logistic-based patient grouping for multi-disciplinary treatment
Laura Maruster, A. J. M. M. Weijters, Geerhard de Vries, Antal van den Bosch, Walter Daelemans |
Artif. Intell. Medicine | 4 |
| 2002 | Parameter optimization for machine-learning of word sense disambiguationabstractVarious Machine Learning (ML) approaches have been demonstrated to produce relatively successful Word Sense Disambiguation (WSD) systems. There are still unexplained differences among the performance measurements of different algorithms, hence it is warranted to deepen the investigation into which algorithm has the right ‘bias’ for this task. In this paper, we show that this is not easy to accomplish, due to intricate interactions between information sources, parameter settings, and properties of the training data. We investigate the impact of parameter optimization on generalization accuracy in a memory-based learning approach to English and Dutch WSD. A ‘word-expert’ architecture was adopted, yielding a set of classifiers, each specialized in one single wordform. The experts consist of multiple memory-based learning classifiers, each taking different information sources as input, combined in a voting scheme. We optimized the architectural and parametric settings for each individual word-expert by performing cross-validation experiments on the learning material. The results of these experiments show that the variation of both the algorithmic parameters and the information sources available to the classifiers leads to large fluctuations in accuracy. We demonstrate that optimization per word-expert leads to an overall significant improvement in the generalization accuracies of the produced WSD systems. Véronique Hoste, Iris Hendrickx, Walter Daelemans, Antal van den Bosch |
Nat. Lang. Eng. | 4 |
| 2001 | Detecting Problematic Turns in Human-Machine Interactions: Rule-induction Versus Memory-based Learning ApproachesabstractWe address the issue of on-line detection of communication problems in spoken dialogue systems. The usefulness is investigated of the sequence of system question types and the word graphs corresponding to the respective user utterances. By applying both rule-induction and memory-based learning techniques to data obtained with a Dutch train time-table information system, the current paper demonstrates that the aforementioned features indeed lead to a method for problem detection that performs significantly above baseline. The results are interesting from a dialogue perspective since they employ features that are present in the majority of spoken dialogue systems and can be obtained with little or no computational overhead. The results are interesting from a machine learning perspective, since they show that the rule-based method performs significantly better than the memory-based method, because the former is better capable of representing interactions between features. Antal van den Bosch, Emiel Krahmer, Marc Swerts |
ACL | 1 |
| 2000 | Systematic design of a 14-bit 150-MS/s CMOS current-steering D/A converterabstractThis paper presents a D/A converter with a 14-bit intrinsic linearity in 0.5µm CMOS technology, which has been designed using a systematic design methodology for current-steering D/A converters. A flexible architecture is proposed for which the design parameters are calculated using a performance-driven top-down design methodology. The layout of the regular structure typical for D/A converters is automatically generated. Measurement results are reported. Due to the systematic design methodology, the design was realized in less than one month total accumulated person effort. Geert Van der Plas, Jan Vandenbussche, Walter Daems, Antal van den Bosch, Georges Gielen, Willy M. C. Sansen |
DAC | 4 |
| 2000 | Unpacking Multi-valued Symbolic Features and Classes in Memory-Based Language Learning
Antal van den Bosch, Jakub Zavrel |
ICML | 1 |
| 2000 | Integrating Seed Names and ngrams for a Named Entity List and Classifier
Sabine Buchholz, Antal van den Bosch |
LREC | 2 |
| 1999 | Memory-Based Morphological AnalysisabstractWe present a general architecture for efficient and deterministic morphological analysis based on memory-based learning, and apply it to morphological analysis of Dutch.The system makes direct mappings from letters in context to rich categories that encode morphological boundaries, syntactic class labels, and spelling changes.Both precision and recall of labeled morphemes are over 84% on held-out dictionary test words and estimated to be over 93% in free text. Antal van den Bosch, Walter Daelemans |
ACL | 1 |
| 1999 | Instance-Family Abstraction in Memory-Based Language Learning
Antal van den Bosch |
ICML | 1 |
| 1999 | Machine learning of word pronunciation: the case against abstractionabstractAn adequate approach to speech translation for small to medium sized tasks is the use of subsequential trans-ducers —a finite state model — as language model for a speech recognizer. These transducers can be automati-cally trained from sample corpora. The use of manually defined categories improves the training of the subsequential transducers when the avail-able data are scarce. These categories depend on the source and target languages we want to translate. We introduce an automatic approach to derive cate-gories that can be used in training subsequential transduc-ers. This approach extends monolingual word clustering methods to the bilingual case using alignments obtained from statistical models. Experimental results indicate that the models trained with these categories have lower trans-lation errors. 1 Bertjan Busser, Walter Daelemans, Antal van den Bosch |
EUROSPEECH | 3 |
| 1999 | Learning Statistically Neutral Tasks without Expert Guidance
A. J. M. M. Weijters, Antal van den Bosch, Eric O. Postma |
NIPS | 2 |
| 1999 | Careful abstraction from instance families in memory-based language learning
Antal van den Bosch |
J. Exp. Theor. Artif. Intell. | 1 |
| 1999 | Forgetting Exceptions is Harmful in Language Learning
Walter Daelemans, Antal van den Bosch, Jakub Zavrel |
Mach. Learn. | 2 |
| 1998 | Do Not Forget: Full Memory in Memory-Based Learning of Word Pronunciation
Antal van den Bosch, Walter Daelemans |
CoNLL | 1 |
| 1998 | Modularity in Inductively-Learned Word Pronunciation Systems
Antal van den Bosch, A. J. M. M. Weijters, Walter Daelemans |
CoNLL | 1 |
| 1998 | Interpretable Neural Networks with BP-SOM
A. J. M. M. Weijters, Antal van den Bosch, H. Jaap van den Herik |
ECML | 2 |
| 1997 | Empirical Learning of Natural Language Processing Task
Walter Daelemans, Antal van den Bosch, A. J. M. M. Weijters |
ECML | 2 |
| 1997 | Avoiding Overfitting with BP-SOM
A. J. M. M. Weijters, H. Jaap van den Herik, Antal van den Bosch, Eric O. Postma |
IJCAI | 3 |
| 1997 | Behavioural Aspects of Combining Backpropagation Learning and Self-organizing MapsabstractBackpropagation learning (BP) is known for its serious limitations in generalizing knowledge from certain types of learning material. In this paper, we describe a new learning algorithm, BP-SOM, which overcomes some of these limitations as is shown by its application to four benchmark tasks. BP-SOM is a combination of a multi-layered feedforward network (MFN) trained with BP and Kohonen's self-organizing maps (SOMs). During the learning process, hidden-unit activations of the MFN are presented as learning vectors to SOMs trained in parallel. The SOM information is used when updating the connection weights of the MFN in addition to standard error backpropagation. The effect of the augmented error signal is that, during learning, clusters of hiddenunit activation patterns of instances associated with the same class tend to become highly similar. In a number of experiments, BP-SOM is shown (i) to improve generalization performance (i.e. avoid overfitting); (ii) to increase the amount of hidden units that can be pruned without loss of generalization performance and (iii) to provide a means for automatic rule extraction from trained networks. The results are compared with results achieved by two other learning algorithms for MFNs: conventional BP and BP augmented with weight decay. From the experiments and the comparisons, we conclude that the hybrid BP-SOM architecture, in which supervised and unsupervised and learning co-operate in finding adequate hidden-layer representations, successfully combines the advantages of supervised and unsupervised learning. A. J. M. M. Weijters, Antal van den Bosch, H. Jaap van den Herik |
Connect. Sci. | 2 |
| 1993 | Data-Oriented Methods for Grapheme-to-Phoneme Conversion
Antal van den Bosch, Walter Daelemans |
EACL | 1 |
| 1993 | Tabtalk: reusability in data-oriented grapheme-to-phoneme conversionabstractIn the traditional (knowledge-based) approach to the design of grapheme-to-phoneme modules in text-to-speech systems, it is claimed that various explicitly coded, language-specific, linguistic knowledge sources are necessary for a good performance. Due to knowledge acquisition bottlenecks, this implies long development cycles. As an alternative, we propose to use inductive methods from machine learning in a simple combined Trie Search and Similarity-Based Reasoning approach and show that, for Dutch, its performance is better than that of the knowledge-based approach and backpropagation learning. Furthermore, we show that our approach is reusable for any language for which a training corpus exists. Keywords: grapheme-to-phoneme conversion, text-tospeech, trie search, similarity-based reasoning, machine learning INTRODUCTION The larger part of research on grapheme-to-phoneme conversion focuses on developing systems that implement various levels of language-specific linguistic knowledge.... Walter Daelemans, Antal van den Bosch |
EUROSPEECH | 2 |