VLDB 2026 Research / reviewers in the wild / expert
Simon Clematide
dblp:02/2249
· DBLP profile ↗
31ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0003-1365-0662ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6Databases, data management, data science and information retrieval · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
Juri Opitz, Corina Julia Raclé, Emanuela Boros, Andrianos Michail, Matteo Romanello, Maud Ehrmann, Simon Clematide |
ECIR (4) | 7 |
| 2026 | MaritimEmails: A Synthetic Dataset for Maritime Chartering Correspondence
Kevin Bründler, Simon Clematide |
LREC | 2 |
| 2026 | A Recipe for Adapting Multilingual Embedders to OCR-Error Robustness and Historical Texts
Andrianos Michail, Stylianos Psychias, Juri Opitz, Simon Clematide |
LREC | 4 |
| 2026 | A Scalable Pipeline for Novelty Detection in Skill Extraction Using Large Language Models
Gian Seifert, Simon Clematide |
LREC | 2 |
| 2025 | PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection ModelsabstractThe task of determining whether two texts are paraphrases has long been a challenge in NLP. However, the prevailing notion of paraphrase is often quite simplistic, offering only a limited view of the vast spectrum of paraphrase phenomena. Indeed, we find that evaluating models in a paraphrase dataset can leave uncertainty about their true semantic understanding. To alleviate this, we create PARAPHRASUS, a benchmark designed for multi-dimensional assessment, benchmarking and selection of paraphrase detection models. We find that paraphrase detection models under our fine-grained evaluation lens exhibit trade-offs that cannot be captured through a single classification dataset. Furthermore, PARAPHRASUS allows prompt calibration for different use cases, tailoring LLM models to specific strictness levels. PARAPHRASUS includes 3 challenges spanning over 10 datasets, including 8 repurposed and 2 newly annotated; we release it along with a benchmarking library at https://github.com/impresso/paraphrasus Andrianos Michail, Simon Clematide, Juri Opitz |
COLING | 2 |
| 2025 | Sentence Smith: Controllable Edits for Evaluating Text EmbeddingsabstractControllable and transparent text generation has been a long-standing goal in NLP.Almost as long-standing is a general idea for addressing this challenge: Parsing text to a symbolic representation, and generating from it.However, earlier approaches were hindered by parsing and generation insufficiencies.Using modern parsers and a safety supervision mechanism, we show how close current methods come to this goal.Concretely, we propose the SENTENCE-SMITH framework for English, which has three steps: 1. Parsing a sentence into a semantic graph.2. Applying human-designed semantic manipulation rules.3. Generating text from the manipulated graph.A final entailment check (4.) verifies the validity of the applied transformation.To demonstrate our framework's utility, we use it to induce hard negative text pairs that challenge text embedding models.Since the controllable generation makes it possible to clearly isolate different types of semantic shifts, we can evaluate text embedding models in a fine-grained way, also addressing an issue in current benchmarking where linguistic phenomena remain opaque.Human validation confirms that our transparent generation process produces texts of good quality.Notably, our way of generation is very resource-efficient, since it relies only on smaller neural networks. Andrianos Michail, Reto Gubelmann, Simon Clematide, Juri Opitz |
EMNLP | 4 |
| 2025 | Interpretable Text Embeddings and Text Similarity Explanation: A SurveyabstractText embeddings are a fundamental component in many NLP tasks, including classification, regression, clustering, and semantic search.However, despite their ubiquitous application, challenges persist in interpreting embeddings and explaining similarities between them.In this work, we provide a structured overview of methods specializing in inherently interpretable text embeddings and text similarity explanation, an underexplored research area.We characterize the main ideas, approaches, and tradeoffs.We compare means of evaluation, discuss overarching lessons learned and finally identify opportunities and open challenges for future research. Juri Opitz, Lucas Möller, Andrianos Michail, Sebastian Padó, Simon Clematide |
EMNLP | 5 |
| 2024 | Mapping Work Task Descriptions from German Job Ads on the O*NET Work Activities OntologyabstractThis work addresses the challenge of extracting job tasks from German job postings and mapping them to the fine-grained work activities classification in the O*NET labor market ontology. By utilizing ontological data with a Multiple Negatives Ranking loss and integrating a modest volume of labeled job advertisement data into the training process, our top configuration achieved a notable precision of 70% for the best mapping on the test set, representing a substantial improvement compared to the 33% baseline delivered by a general-domain SBERT. In our experiments the following factors proved to be most effective for improving SBERT models: First, the incorporation of subspan markup, both during training and inference, supports accurate classification, by streamlining varied job ad task formats with structured, uniform ontological work activities. Second, the inclusion of additional occupational information from O*NET into training supported learning by contextualizing hierarchical ontological relationships. Third, the most significant performance improvement was achieved by updating SBERT models with labeled job ad data specifically addressing challenging cases encountered during pre-finetuning, effectively bridging the semantic gap between O*NET and job ad data. Ann-Sophie Gnehm, Simon Clematide |
LREC/COLING | 2 |
| 2024 | New "ArchAIval" Practices: Using GPT for OCR and Historical Narration of Index Cards
Phillip Ströbel, Simon Clematide, Pascal Werner |
TPDL (2) | 2 |
| 2022 | Introducing the HIPE 2022 Shared Task: Named Entity Recognition and Linking in Multilingual Historical Documents
Maud Ehrmann, Matteo Romanello, Antoine Doucet, Simon Clematide |
ECIR (2) | 4 |
| 2022 | Evaluation of Transfer Learning and Domain Adaptation for Analyzing German-Speaking Job AdvertisementsabstractThis paper presents text mining approaches on German-speaking job advertisements to enable social science research on the development of the labour market over the last 30 years. In order to build text mining applications providing information about profession and main task of a job, as well as experience and ICT skills needed, we experiment with transfer learning and domain adaptation. Our main contribution consists in building language models which are adapted to the domain of job advertisements, and their assessment on a broad range of machine learning problems. Our findings show the large value of domain adaptation in several respects. First, it boosts the performance of fine-tuned task-specific models consistently over all evaluation experiments. Second, it helps to mitigate rapid data shift over time in our special domain, and enhances the ability to learn from small updates with new, labeled task data. Third, domain-adaptation of language models is efficient: With continued in-domain pre-training we are able to outperform general-domain language models pre-trained on ten times more data. We share our domain-adapted language models and data with the research community. Ann-Sophie Gnehm, Eva Bühlmann, Simon Clematide |
LREC | 3 |
| 2022 | Evaluation of HTR models without Ground Truth MaterialabstractThe evaluation of Handwritten Text Recognition (HTR) models during their development is straightforward: because HTR is a supervised problem, the usual data split into training, validation, and test data sets allows the evaluation of models in terms of accuracy or error rates. However, the evaluation process becomes tricky as soon as we switch from development to application. A compilation of a new (and forcibly smaller) ground truth (GT) from a sample of the data that we want to apply the model on and the subsequent evaluation of models thereon only provides hints about the quality of the recognised text, as do confidence scores (if available) the models return. Moreover, if we have several models at hand, we face a model selection problem since we want to obtain the best possible result during the application phase. This calls for GT-free metrics to select the best model, which is why we (re-)introduce and compare different metrics, from simple, lexicon-based to more elaborate ones using standard language models and masked language models (MLM). We show that MLM-based evaluation can compete with lexicon-based methods, with the advantage that large and multilingual transformers are readily available, thus making compiling lexical resources for other metrics superfluous. Phillip Ströbel, Martin Volk 0001, Simon Clematide, Raphael Schwitter, Tobias Hodel, David Schoch |
LREC | 3 |
| 2020 | Semi-supervised Contextual Historical Text NormalizationabstractHistorical text normalization, the task of mapping historical word forms to their modern counterparts, has recently attracted a lot of interest (Bollmann, 2019;Tang et al., 2018;Lusetti et al., 2018;Bollmann et al., 2018;Robertson and Goldwater, 2018;Bollmann et al., 2017;Korchagina, 2017).Yet, virtually all approaches suffer from the two limitations: 1) They consider a fully supervised setup, often with impractically large manually normalized datasets; 2) Normalization happens on words in isolation.By utilizing a simple generative normalization model and obtaining powerful contextualization from the target-side language model, we train accurate models with unlabeled historical data.In realistic training scenarios, our approach often leads to reduction in manually normalized data at the same accuracy levels. Peter Makarov, Simon Clematide |
ACL | 2 |
| 2020 | Introducing the CLEF 2020 HIPE Shared Task: Named Entity Recognition and Linking on Historical NewspapersabstractSince its introduction some twenty years ago, named entity (NE) processing has become an essential component of virtually any text mining application and has undergone major changes. Recently, two main trends characterise its developments: the adoption of deep learning architectures and the consideration of textual material originating from historical and cultural heritage collections. While the former opens up new opportunities, the latter introduces new challenges with heterogeneous, historical and noisy inputs. If NE processing tools are increasingly being used in the context of historical documents, performance values are below the ones on contemporary data and are hardly comparable. In this context, this paper introduces the CLEF 2020 Evaluation Lab HIPE (Identifying Historical People, Places and other Entities) on named entity recognition and linking on diachronic historical newspaper material in French, German and English. Our objective is threefold: strengthening the robustness of existing approaches on non-standard inputs, enabling performance comparison of NE processing on historical texts, and, in the long run, fostering efficient semantic indexing of historical documents in order to support scholarship on digital cultural heritage collections. Maud Ehrmann, Matteo Romanello, Stefan Bircher, Simon Clematide |
ECIR (2) | 4 |
| 2020 | Language Resources for Historical Newspapers: the Impresso CollectionabstractFollowing decades of massive digitization, an unprecedented amount of historical document facsimiles can now be retrieved and accessed via cultural heritage online portals. If this represents a huge step forward in terms of preservation and accessibility, the next fundamental challenge– and real promise of digitization– is to exploit the contents of these digital assets, and therefore to adapt and develop appropriate language technologies to search and retrieve information from this ‘Big Data of the Past’. Yet, the application of text processing tools on historical documents in general, and historical newspapers in particular, poses new challenges, and crucially requires appropriate language resources. In this context, this paper presents a collection of historical newspaper data sets composed of text and image resources, curated and published within the context of the ‘impresso - Media Monitoring of the Past’ project. With corpora, benchmarks, semantic annotations and language models in French, German and Luxembourgish covering ca. 200 years, the objective of the impresso resource collection is to contribute to historical language resources, and thereby strengthen the robustness of approaches to non-standard inputs and foster efficient processing of historical documents. Maud Ehrmann, Matteo Romanello, Simon Clematide, Phillip Ströbel, Raphaël Barman |
LREC | 3 |
| 2020 | How Much Data Do You Need? About the Creation of a Ground Truth for Black Letter and the Effectiveness of Neural OCRabstractRecent advances in Optical Character Recognition (OCR) and Handwritten Text Recognition (HTR) have led to more accurate textrecognition of historical documents. The Digital Humanities heavily profit from these developments, but they still struggle whenchoosing from the plethora of OCR systems available on the one hand and when defining workflows for their projects on the other hand. In this work, we present our approach to build a ground truth for a historical German-language newspaper published in black letter. Wealso report how we used it to systematically evaluate the performance of different OCR engines. Additionally, we used this ground truthto make an informed estimate as to how much data is necessary to achieve high-quality OCR results. The outcomes of our experimentsshow that HTR architectures can successfully recognise black letter text and that a ground truth size of 50 newspaper pages suffices toachieve good OCR accuracy. Moreover, our models perform equally well on data they have not seen during training, which means thatadditional manual correction for diverging data is superfluous. Phillip Ströbel, Simon Clematide, Martin Volk 0001 |
LREC | 2 |
| 2018 | Neural Transition-based String Transduction for Limited-Resource Setting in MorphologyabstractWe present a neural transition-based model that uses a simple set of edit actions (copy, delete, insert) for morphological transduction tasks such as inflection generation, lemmatization, and reinflection. In a large-scale evaluation on four datasets and dozens of languages, our approach consistently outperforms state-of-the-art systems on low and medium training-set sizes and is competitive in the high-resource setting. Learning to apply a generic copy action enables our approach to generalize quickly from a few data points. We successfully leverage minimum risk training to compensate for the weaknesses of MLE parameter learning and neutralize the negative effects of training a pipeline with a separate character aligner. Peter Makarov, Simon Clematide |
COLING | 2 |
| 2018 | Imitation Learning for Neural Morphological String TransductionabstractWe employ imitation learning to train a neural transition-based string transducer for morphological tasks such as inflection generation and lemmatization.Previous approaches to training this type of model either rely on an external character aligner for the production of gold action sequences, which results in a suboptimal model due to the unwarranted dependence on a single gold action sequence despite spurious ambiguity, or require warm starting with an MLE model.Our approach only requires a simple expert policy, eliminating the need for a character aligner or warm start.It also addresses familiar MLE training biases and leads to strong and state-of-the-art performance on several benchmarks. Peter Makarov, Simon Clematide |
EMNLP | 2 |
| 2018 | Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French
Jean-Philippe Goldman, Simon Clematide, Mathieu Avanzi, Raphaël Tandler |
LREC | 2 |
| 2017 | Verb-Mediated Composition of Attitude Relations Comprising Reader and Writer Perspective
Manfred Klenner, Simon Clematide, Don Tuggener |
CICLing (2) | 2 |
| 2016 | Crowdsourcing an OCR Gold Standard for a German and French Heritage Corpus
Simon Clematide, Lenz Furrer, Martin Volk 0001 |
LREC | 1 |
| 2015 | A multilingual gold-standard corpus for biomedical concept recognition: the Mantra GSCabstractOBJECTIVE: To create a multilingual gold-standard corpus for biomedical concept recognition. MATERIALS AND METHODS: We selected text units from different parallel corpora (Medline abstract titles, drug labels, biomedical patent claims) in English, French, German, Spanish, and Dutch. Three annotators per language independently annotated the biomedical concepts, based on a subset of the Unified Medical Language System and covering a wide range of semantic groups. To reduce the annotation workload, automatically generated preannotations were provided. Individual annotations were automatically harmonized and then adjudicated, and cross-language consistency checks were carried out to arrive at the final annotations. RESULTS: The number of final annotations was 5530. Inter-annotator agreement scores indicate good agreement (median F-score 0.79), and are similar to those between individual annotators and the gold standard. The automatically generated harmonized annotation set for each language performed equally well as the best annotator for that language. DISCUSSION: The use of automatic preannotations, harmonized annotations, and parallel corpora helped to keep the manual annotation efforts manageable. The inter-annotator agreement scores provide a reference standard for gauging the performance of automatic annotation techniques. CONCLUSION: To our knowledge, this is the first gold-standard corpus for biomedical concept recognition in languages other than English. Other distinguishing features are the wide variety of semantic groups that are being covered, and the diversity of text genres that were annotated. Jan A. Kors, Simon Clematide, Saber A. Akhondi, Erik M. van Mulligen, Dietrich Rebholz-Schuhmann |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Using Large Biomedical Databases as Gold Annotations for Automatic Relation Extraction
Tilia Ellendorff, Fabio Rinaldi 0001, Simon Clematide |
LREC | 3 |
| 2014 | Collaboratively Annotating Multilingual Parallel Corpora in the Biomedical Domain―some MANTRAs
Johannes Hellrich, Simon Clematide, Udo Hahn, Dietrich Rebholz-Schuhmann |
LREC | 2 |
| 2014 | OntoGene web services for biomedical text miningabstractText mining services are rapidly becoming a crucial component of various knowledge management pipelines, for example in the process of database curation, or for exploration and enrichment of biomedical data within the pharmaceutical industry. Traditional architectures, based on monolithic applications, do not offer sufficient flexibility for a wide range of use case scenarios, and therefore open architectures, as provided by web services, are attracting increased interest. We present an approach towards providing advanced text mining capabilities through web services, using a recently proposed standard for textual data interchange (BioC). The web services leverage a state-of-the-art platform for text mining (OntoGene) which has been tested in several community-organized evaluation challenges,with top ranked results in several of them. Fabio Rinaldi 0001, Simon Clematide, Hernani Marques-Madeira, Tilia Ellendorff, Martin Romacker, Raul Rodriguez-Esteban |
BMC Bioinform. | 2 |
| 2012 | MLSA - A Multi-layered Reference Corpus for German Sentiment Analysis
Simon Clematide, Stefan Gindl, Manfred Klenner, Stefanos Petrakis, Robert Remus, Josef Ruppenhofer, Ulli Waltinger, Michael Wiegand |
LREC | 1 |
| 2012 | Dependency parsing for interaction detection in pharmacogenomics
Gerold Schneider, Fabio Rinaldi 0001, Simon Clematide |
LREC | 3 |
| 2012 | Relation mining experiments in the pharmacogenomics domain
Fabio Rinaldi 0001, Gerold Schneider, Simon Clematide |
J. Biomed. Informatics | 3 |
| 2011 | BioCreative III interactive task: an overviewabstractBACKGROUND: The BioCreative challenge evaluation is a community-wide effort for evaluating text mining and information extraction systems applied to the biological domain. The biocurator community, as an active user of biomedical literature, provides a diverse and engaged end user group for text mining tools. Earlier BioCreative challenges involved many text mining teams in developing basic capabilities relevant to biological curation, but they did not address the issues of system usage, insertion into the workflow and adoption by curators. Thus in BioCreative III (BC-III), the InterActive Task (IAT) was introduced to address the utility and usability of text mining tools for real-life biocuration tasks. To support the aims of the IAT in BC-III, involvement of both developers and end users was solicited, and the development of a user interface to address the tasks interactively was requested. RESULTS: A User Advisory Group (UAG) actively participated in the IAT design and assessment. The task focused on gene normalization (identifying gene mentions in the article and linking these genes to standard database identifiers), gene ranking based on the overall importance of each gene mentioned in the article, and gene-oriented document retrieval (identifying full text papers relevant to a selected gene). Six systems participated and all processed and displayed the same set of articles. The articles were selected based on content known to be problematic for curation, such as ambiguity of gene names, coverage of multiple genes and species, or introduction of a new gene name. Members of the UAG curated three articles for training and assessment purposes, and each member was assigned a system to review. A questionnaire related to the interface usability and task performance (as measured by precision and recall) was answered after systems were used to curate articles. Although the limited number of articles analyzed and users involved in the IAT experiment precluded rigorous quantitative analysis of the results, a qualitative analysis provided valuable insight into some of the problems encountered by users when using the systems. The overall assessment indicates that the system usability features appealed to most users, but the system performance was suboptimal (mainly due to low accuracy in gene normalization). Some of the issues included failure of species identification and gene name ambiguity in the gene normalization task leading to an extensive list of gene identifiers to review, which, in some cases, did not contain the relevant genes. The document retrieval suffered from the same shortfalls. The UAG favored achieving high performance (measured by precision and recall), but strongly recommended the addition of features that facilitate the identification of correct gene and its identifier, such as contextual information to assist in disambiguation. DISCUSSION: The IAT was an informative exercise that advanced the dialog between curators and developers and increased the appreciation of challenges faced by each group. A major conclusion was that the intended users should be actively involved in every phase of software development, and this will be strongly encouraged in future tasks. The IAT Task provides the first steps toward the definition of metrics and functional requirements that are necessary for designing a formal evaluation of interactive curation systems in the BioCreative IV challenge. Cecilia N. Arighi, Phoebe M. Roberts, Shashank Agarwal, Sanmitra Bhattacharya, Gianni Cesareni, Andrew Chatr-aryamontri, Simon Clematide, Pascale Gaudet, Michelle G. Giglio, Ian Harrow, Eva Huala, Martin Krallinger, Ulf Leser, Zhiyong Lu, Lois J. Maltais, Naoaki Okazaki, Livia Perfetto, Fabio Rinaldi 0001, Rune Sætre, David Salgado, Padmini Srinivasan, Philippe Thomas 0002, Luca Toldo, Lynette Hirschman, Cathy H. Wu |
BMC Bioinform. | 7 |
| 2011 | Detection of interaction articles and experimental methods in biomedical literatureabstractBACKGROUND: This article describes the approaches taken by the OntoGene group at the University of Zurich in dealing with two tasks of the BioCreative III competition: classification of articles which contain curatable protein-protein interactions (PPI-ACT) and extraction of experimental methods (PPI-IMT). RESULTS: Two main achievements are described in this paper: (a) a system for document classification which crucially relies on the results of an advanced pipeline of natural language processing tools; (b) a system which is capable of detecting all experimental methods mentioned in scientific literature, and listing them with a competitive ranking (AUC iP/R > 0.5). CONCLUSIONS: The results of the BioCreative III shared evaluation clearly demonstrate that significant progress has been achieved in the domain of biomedical text mining in the past few years. Our own contribution, together with the results of other participants, provides evidence that natural language processing techniques have become by now an integral part of advanced text mining approaches. Gerold Schneider, Simon Clematide, Fabio Rinaldi 0001 |
BMC Bioinform. | 2 |
| 2010 | OntoGene in BioCreative II.5abstractWe describe a system for the detection of mentions of protein-protein interactions in the biomedical scientific literature. The original system was developed as a part of the OntoGene project, which focuses on using advanced computational linguistic techniques for text mining applications in the biomedical domain. In this paper, we focus in particular on the participation to the BioCreative II.5 challenge, where the OntoGene system achieved best-ranked results. Additionally, we describe a feature-analysis experiment performed after the challenge, which shows the unexpected result that one single feature alone performs better than the combination of features used in the challenge. Fabio Rinaldi 0001, Gerold Schneider, Kaarel Kaljurand, Simon Clematide, Thérèse Vachon, Martin Romacker |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |