VLDB 2026 Research / reviewers in the wild / expert
Mike Kestemont
dblp:42/10071
· DBLP profile ↗
10ranked-venue papers
2as first author
4since 2021 · last 2024
0000-0003-3590-693XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CATMuS Medieval: A Multilingual Large-Scale Cross-Century Dataset in Latin Script for Handwritten Text Recognition and Beyond
Thibault Clérice, Ariane Pinche, Malamatenia Vlachou-Efstathiou, Alix Chagué, Jean-Baptiste Camps, Matthias Gille Levenson, Olivier Brisville-Fertin, Federico Boschetti, Franz Fischer, Michael Gervers, Agnès Boutreux, Avery Manton, Simon Gabay, Patricia O'Connor, Wouter Haverals, Mike Kestemont, Caroline Vandyck, Benjamin Kiessling |
ICDAR (3) | 16 |
| 2022 | Overview of PAN 2022: Authorship Verification, Profiling Irony and Stereotype Spreaders, Style Change Detection, and Trigger Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reyner Ortega-Bueno, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 5 |
| 2021 | Overview of PAN 2021: Authorship Verification, Profiling Hate Speech Spreaders on Twitter, and Style Change Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Gretel Liz De la Peña Sarracén, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle |
ECIR (2) | 4 |
| 2021 | Multi-modal Label Retrieval for the Visual Arts: The Case of IconclassabstractIconclass is an iconographic classification system from the domain of cultural heritage which is used to annotate subjects represented in the visual arts.In this work, we investigate the feasibility of automatically assigning Iconclass codes to visual artworks using a cross-modal retrieval set-up.We explore the text and image branches of the cross-modal network.In addition, we describe a multi-modal architecture that can jointly capitalize on multiple feature sources: textual features, coming from the titles for these artworks (in multiple languages) and visual features, extracted from photographic reproductions of the artworks.We utilize Iconclass definitions in English as matching labels.We evaluate our approach on a publicly available dataset of artworks (containing English and Dutch titles).Our results demonstrate that, in isolation, textual features strongly outperform visual features, although visual features can still offer a useful complement to purely linguistic features.Moreover, we show the cross-lingual (Dutch-English) strategy to be on par with the monolingual approach (English-English), which opens important perspectives for applications of this approach beyond resource-rich languages. Nikolay Banar, Walter Daelemans, Mike Kestemont |
ICAART (1) | 3 |
| 2020 | Shared Tasks on Authorship Analysis at PAN 2020
Janek Bevendorff, Bilal Ghanem, Anastasia Giahanou, Mike Kestemont, Enrique Manjavacas, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Günther Specht, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Eva Zangerle |
ECIR (2) | 4 |
| 2020 | Transfer Learning for Digital Heritage Collections: Comparing Neural Machine Translation at the Subword-level and Character-levelabstractTransfer learning via pre-training has become an important strategy for the efficient application of NLP methods in domains where only limited training data is available.This paper reports on a focused case study in which we apply transfer learning in the context of neural machine translation (French-Dutch) for cultural heritage metadata (i.e.titles of artistic works).Nowadays, neural machine translation (NMT) is commonly applied at the subword level using byte-pair encoding (BPE), because word-level models struggle with rare and out-of-vocabulary words.Because unseen vocabulary is a significant issue in domain adaptation, BPE seems a better fit for transfer learning across text varieties.We discuss an experiment in which we compare a subword-level to a character-level NMT approach.We pre-trained models on a large, generic corpus and fine-tuned them in a two-stage process: first, on a domain-specific dataset extracted from Wikipedia, and then on our metadata.While our experiments show comparable performance for character-level and BPEbased models on the general dataset, we demonstrate that the character-level approach nevertheless yields major downstream performance gains during the subsequent stages of fine-tuning.We therefore conclude that character-level translation can be beneficial compared to the popular subword-level approach in the cultural heritage domain. Nikolay Banar, Karine Lasaracina, Walter Daelemans, Mike Kestemont |
ICAART (1) | 4 |
| 2019 | Generation of Hip-Hop Lyrics with Hierarchical Modeling and Conditional TemplatesabstractThis paper addresses Hip-Hop lyric generation with conditional Neural Language Models.We develop a simple yet effective mechanism to extract and apply conditional templates from text snippets, and show-on the basis of a large-scale crowd-sourced manual evaluation-that these templates significantly improve the quality and realism of the generated snippets.Importantly, the proposed approach enables end-to-end training, targeting formal properties of text such as rhythm and rhyme, which are central characteristics of rap texts.Additionally, we explore how generating text at different scales (e.g.character-level or word-level) affects the quality of the output.We find that a hybrid form-a hierarchical model that aims to integrate Language Modeling at both word and character-level scalesyields significant improvements in text quality, yet surprisingly, cannot exploit conditional templates to their fullest extent.Our findings highlight that text generation models based on Recurrent Neural Networks (RNN) are sensitive to the modeling scale and call for further research on the observed differences in effectiveness of the conditioning mechanism at different scales. Enrique Manjavacas, Mike Kestemont, Folgert Karsdorp |
INLG | 2 |
| 2016 | Authenticating the writings of Julius Caesar
Mike Kestemont, Justin Anthony Stover, Moshe Koppel, Folgert Karsdorp, Walter Daelemans |
Expert Syst. Appl. | 1 |
| 2016 | Computational authorship verification method attributes a new work to a major 2nd century African authorabstractWe discuss a real‐world application of a recently proposed machine learning method for authorship verification. Authorship verification is considered an extremely difficult task in computational text classification, because it does not assume that the correct author of an anonymous text is included in the candidate authors available. To determine whether 2 documents have been written by the same author, the verification method discussed uses repeated feature subsampling and a pool of impostor authors. We use this technique to attribute a newly discovered Latin text from antiquity (the Compendiosa expositio) to Apuleius. This North African writer was one of the most important authors of the Roman Empire in the 2nd century and authored one of the world's first novels. This attribution has profound and wide‐reaching cultural value, because it has been over a century since a new text by a major author from antiquity was discovered. This research therefore illustrates the rapidly growing potential of computational methods for studying the global textual heritage. Justin Anthony Stover, Yaron Winter, Moshe Koppel, Mike Kestemont |
J. Assoc. Inf. Sci. Technol. | 4 |
| 2012 | The Netlog Corpus. A Resource for the Study of Flemish Dutch Internet Language
Mike Kestemont, Claudia Peersman, Benny De Decker, Guy De Pauw, Kim Luyckx, Roser Morante, Frederik Vaassen, Janneke van de Loo, Walter Daelemans |
LREC | 1 |