VLDB 2026 Research / reviewers in the wild / expert
Alexander Mehler
dblp:m/AlexanderMehler
· DBLP profile ↗
50ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-2567-7539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards the Generation and Application of Dynamic Web-Based Visualization of UIMA-based Annotations for Big-Data Corpora with the Help of Unified Dynamic Annotation VisualizerabstractThe automatic and manual annotation of unstructured corpora is a routine task in many scientific fields and is supported by a variety of existing software solutions.Despite this variety, few solutions currently support annotation visualization, especially for dynamic generation and interaction.To bridge this gap and visualize annotated corpora based on user-, project-, or corpus-specific aspects, we developed Unified Dynamic Annotation Visualizer (UDAV).UDAV is a web-based solution that implements features not supported by comparable tools, enabling a customizable and extensible toolbox for interacting with annotations and allowing integration into existing big-data frameworks.We exemplify UDAV through a range of visualizations and also provide an evaluation of corpus import and processing performance. Thiemo Dahmann, Julian Schneider, Philipp Stephan, Giuseppe Abrami, Alexander Mehler |
LREC | 5 |
| 2026 | GhostWriter: Hidden AI-Generated Texts over Multiple Languages, Domains and Generators
Manuel Schaaf, Kevin Bönisch, Alexander Mehler |
LREC | 3 |
| 2026 | Predicting Topic (Co-)Occurrence Using Topic Networks Built from the Project Gutenberg Corpus
Bhuvanesh Verma, Alexander Mehler |
LREC | 2 |
| 2025 | MedLinkDE - MedDRA Entity Linking for German with Guided Chain of Thought ReasoningabstractRoman Christof, Farnaz Zeidi, Manuela Messelhäußer, Dirk Mentzer, Renate Koenig, Liam Childs, Alexander Mehler. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Roman Christof, Farnaz Zeidi, Manuela Messelhäußer, Dirk Mentzer, Renate König, Liam Harold Childs, Alexander Mehler |
EMNLP | 7 |
| 2024 | German Parliamentary Corpus (GerParCor) ReloadedabstractIn 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level from the countries of Germany, Austria, Switzerland and Liechtenstein, was collected and published - GerParCor. Through GerParCor, it became possible to provide for the first time various parliamentary protocols which were not available digitally and, moreover, could not be retrieved and processed in a uniform manner. Furthermore, GerParCor was additionally preprocessed using NLP methods and made available in XMI format. In this paper, GerParCor is significantly updated by including all new parliamentary protocols in the corpus, as well as adding and preprocessing further parliamentary protocols previously not covered, so that a period up to 1797 is now covered. Besides the integration of a new, state-of-the-art and appropriate NLP preprocessing for the handling of large text corpora, this update also provides an overview of the further reuse of GerParCor by presenting various provisioning capabilities such as API’s, among others. Giuseppe Abrami, Mevlüt Bagci, Alexander Mehler |
LREC/COLING | 3 |
| 2024 | German SRL: Corpus Construction and Model TrainingabstractA useful semantic role-annotated resource for training semantic role models for the German language is missing. We point out some problems of previous resources and provide a new one due to a combined translation and alignment process: The gold standard CoNLL-2012 semantic role annotations are translated into German. Semantic role labels are transferred due to alignment models. The resulting dataset is used to train a German semantic role model. With F1-scores around 0.7, the major roles achieve competitive evaluation scores, but avoid limitations of previous approaches. The described procedure can be applied to other languages as well. Maxim Konca, Andy Lücking, Alexander Mehler |
LREC/COLING | 3 |
| 2024 | Dependencies over Times and Tools (DoTT)abstractPurpose: Based on the examples of English and German, we investigate to what extent parsers trained on modern variants of these languages can be transferred to older language levels without loss. Methods: We developed a treebank called DoTT (https://github.com/texttechnologylab/DoTT) which covers, roughly, the time period from 1800 until today, in conjunction with the further development of the annotation tool DependencyAnnotator. DoTT consists of a collection of diachronic corpora enriched with dependency annotations using 3 parsers, 6 pre-trained language models, 5 newly trained models for German, and two tag sets (TIGER and Universal Dependencies). To assess how the different parsers perform on texts from different time periods, we created a gold standard sample as a benchmark. Results: We found that the parsers/models perform quite well on modern texts (document-level LAS ranging from 82.89 to 88.54) and slightly worse on older texts, as expected (average document-level LAS 84.60 vs. 86.14), but not significantly. For German texts, the (German) TIGER scheme achieved slightly better results than UD. Conclusion: Overall, this result speaks for the transferability of parsers to past language levels, at least dating back until around 1800. This very transferability, it is however argued, means that studies of language change in the field of dependency syntax can draw on dependency distance but miss out on some grammatical phenomena. Andy Lücking, Giuseppe Abrami, Leon Hammerla, Marc Rahn, Daniel Baumartz, Steffen Eger, Alexander Mehler |
LREC/COLING | 7 |
| 2023 | BUNDESTAG-MINE: Natural Language Processing for Extracting Key Information from Government DocumentsabstractAs governments worldwide continue to release vast amounts of textual information, the need for efficient and insightful tools to extract, interpret and present this data has become increasingly critical. Towards solving this issue, we present the BUNDESTAG-MINE: an environment that periodically retrieves pertinent data from the German parliament, parses and analyzes it using pipelines for natural language processing, and then displays the results in a web application that is publicly accessible. BUNDESTAG-MINE helps to extract key information from parliamentary documents in a visually appealing matter for many use cases. For instance, the tool can be leveraged by journalists for news detection, lawyers for compliance checking, linguists for discourse analysis, and the broad public to inform themselves about the positions of political party members on a topic. Kevin Bönisch, Giuseppe Abrami, Sabine Wehnert, Alexander Mehler |
JURIX | 4 |
| 2022 | Tafsir Dataset: A Novel Multi-Task Benchmark for Named Entity Recognition and Topic Modeling in Classical Arabic LiteratureabstractVarious historical languages, which used to be lingua franca of science and arts, deserve the attention of current NLP research. In this work, we take the first data-driven steps towards this research line for Classical Arabic (CA) by addressing named entity recognition (NER) and topic modeling (TM) on the example of CA literature. We manually annotate the encyclopedic work of Tafsir Al-Tabari with span-based NEs, sentence-based topics, and span-based subtopics, thus creating the Tafsir Dataset with over 51,000 sentences, the first large-scale multi-task benchmark for CA. Next, we analyze our newly generated dataset, which we make open-source available, with current language models (lightweight BiLSTM, transformer-based MaChAmP) along a novel script compression method, thereby achieving state-of-the-art performance for our target task CA-NER. We also show that CA-TM from the perspective of historical topic models, which are central to Arabic studies, is very challenging. With this interdisciplinary work, we lay the foundations for future research on automatic analysis of CA literature. Sajawel Ahmed, Rob van der Goot, Misbahur Rehman, Carl Kruse, Ömer Özsoy, Alexander Mehler, Gemma Roig |
COLING | 6 |
| 2022 | German Parliamentary Corpus (GerParCor)abstractParliamentary debates represent a large and partly unexploited treasure trove of publicly accessible texts. In the German-speaking area, there is a certain deficit of uniformly accessible and annotated corpora covering all German-speaking parliaments at the national and federal level. To address this gap, we introduce the German Parliamentary Corpus (GerParCor). GerParCor is a genre-specific corpus of (predominantly historical) German-language parliamentary protocols from three centuries and four countries, including state and federal level data. In addition, GerParCor contains conversions of scanned protocols and, in particular, of protocols in Fraktur converted via an OCR process based on Tesseract. All protocols were preprocessed by means of the NLP pipeline of spaCy3 and automatically annotated with metadata regarding their session date. GerParCor is made available in the XMI format of the UIMA project. In this way, GerParCor can be used as a large corpus of historical texts in the field of political communication for various tasks in NLP. Giuseppe Abrami, Mevlüt Bagci, Leon Hammerla, Alexander Mehler |
LREC | 4 |
| 2022 | I still have Time(s): Extending HeidelTime for German TextsabstractHeidelTime is one of the most widespread and successful tools for detecting temporal expressions in texts. Since HeidelTime’s pattern matching system is based on regular expression, it can be extended in a convenient way. We present such an extension for the German resources of HeidelTime: HeidelTimeExt. The extension has been brought about by means of observing false negatives within real world texts and various time banks. The gain in coverage is 2.7 % or 8.5 %, depending on the admitted degree of potential overgeneralization. We describe the development of HeidelTimeExt, its evaluation on text samples from various genres, and share some linguistic observations. HeidelTimeExt can be obtained from https://github.com/texttechnologylab/heideltime. Andy Lücking, Manuel Stoeckel, Giuseppe Abrami, Alexander Mehler |
LREC | 4 |
| 2022 | What do Toothbrushes do in the Kitchen? How Transformers Think our World is StructuredabstractTransformer-based models are now predominant in NLP.They outperform approaches based on static models in many respects.This success has in turn prompted research that reveals a number of biases in the language models generated by transformers.In this paper we utilize this research on biases to investigate to what extent transformer-based language models allow for extracting knowledge about object relations (X occurs in Y ; X consists of Z; action A involves using X).To this end, we compare contextualized models with their static counterparts.We make this comparison dependent on the application of a number of similarity measures and classifiers.Our results are threefold: Firstly, we show that the models combined with the different similarity measures differ greatly in terms of the amount of knowledge they allow for extracting.Secondly, our results suggest that similarity measures perform much worse than classifier-based approaches.Thirdly, we show that, surprisingly, static models perform almost as well as contextualized models -in some cases even better. Alexander Henlein, Alexander Mehler |
NAACL-HLT | 2 |
| 2020 | TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of TextsabstractThe annotation of texts and other material in the field of digital humanities and Natural Language Processing (NLP) is a common task of research projects. At the same time, the annotation of corpora is certainly the most time- and cost-intensive component in research projects and often requires a high level of expertise according to the research interest. However, for the annotation of texts, a wide range of tools is available, both for automatic and manual annotation. Since the automatic pre-processing methods are not error-free and there is an increasing demand for the generation of training data, also with regard to machine learning, suitable annotation tools are required. This paper defines criteria of flexibility and efficiency of complex annotations for the assessment of existing annotation tools. To extend this list of tools, the paper describes TextAnnotator, a browser-based, multi-annotation system, which has been developed to perform platform-independent multimodal annotations and annotate complex textual structures. The paper illustrates the current state of development of TextAnnotator and demonstrates its ability to evaluate annotation quality (inter-annotator agreement) at runtime. In addition, it will be shown how annotations of different users can be performed simultaneously and collaboratively on the same document from different platforms using UIMA as the basis for annotation. Giuseppe Abrami, Manuel Stoeckel, Alexander Mehler |
LREC | 3 |
| 2020 | On the Influence of Coreference Resolution on Word Embeddings in Lexical-semantic Evaluation TasksabstractCoreference resolution (CR) aims to find all spans of a text that refer to the same entity. The F1-Scores on these task have been greatly improved by new developed End2End-approaches and transformer networks. The inclusion of CR as a pre-processing step is expected to lead to improvements in downstream tasks. The paper examines this effect with respect to word embeddings. That is, we analyze the effects of CR on six different embedding methods and evaluate them in the context of seven lexical-semantic evaluation tasks and instantiation/hypernymy detection. Especially in the last tasks we hoped for a significant increase in performance. We show that all word embedding approaches do not benefit significantly from pronoun substitution. The measurable improvements are only marginal (around 0.5% in most test cases). We explain this result with the loss of contextual information, reduction of the relative occurrence of rare words and the lack of pronouns to be replaced. Alexander Henlein, Alexander Mehler |
LREC | 2 |
| 2020 | Recognizing Sentence-level Logical Document Structures with the Help of Context-free GrammarsabstractCurrent sentence boundary detectors split documents into sequentially ordered sentences by detecting their beginnings and ends. Sentences, however, are more deeply structured even on this side of constituent and dependency structure: they can consist of a main sentence and several subordinate clauses as well as further segments (e.g. inserts in parentheses); they can even recursively embed whole sentences and then contain multiple sentence beginnings and ends. In this paper, we introduce a tool that segments sentences into tree structures to detect this type of recursive structure. To this end, we retrain different constituency parsers with the help of modified training data to transform them into sentence segmenters. With these segmenters, documents are mapped to sequences of sentence-related “logical document structures”. The resulting segmenters aim to improve downstream tasks by providing additional structural information. In this context, we experiment with German dependency parsing. We show that for certain sentence categories, which can be determined automatically, improvements in German dependency parsing can be achieved using our segmenter for preprocessing. The assumption suggests that improvements in other languages and tasks can be achieved. Jonathan Hildebrand, Wahed Hemati, Alexander Mehler |
LREC | 3 |
| 2019 | Computing Classifier-Based Embeddings with the Help of Text2ddc
Tolga Uslu, Alexander Mehler, Daniel Baumartz |
CICLing (2) | 2 |
| 2019 | BIOfid Dataset: Publishing a German Gold Standard for Named Entity Recognition in Historical Biodiversity LiteratureabstractThe Specialized Information Service Biodiversity Research (BIOfid) has been launched to mobilize valuable biological data from printed literature hidden in German libraries for over the past 250 years.In this project, we annotate German texts converted by OCR from historical scientific literature on the biodiversity of plants, birds, moths and butterflies.Our work enables the automatic extraction of biological information previously buried in the mass of papers and volumes.For this purpose, we generated training data for the tasks of Named Entity Recognition (NER) and Taxa Recognition (TR) in biological documents.We use this data to train a number of leading machine learning tools and create a gold standard for TR in biodiversity literature.More specifically, we perform a practical analysis of our newly generated BIOfid dataset through various downstream-task evaluations and establish a new state of the art for TR with 80.23% Fscore.In this sense, our paper lays the foundations for future work in the field of information extraction in biology texts. Sajawel Ahmed, Manuel Stoeckel, Christine Driller, Adrian Pachzelt, Alexander Mehler |
CoNLL | 5 |
| 2018 | Resource-Size Matters: Improving Neural Named Entity Recognition with Optimized Large CorporaabstractThis study improves the performance of neural named entity recognition by a margin of up to 11% in terms of F-score on the example of a low-resource language like German, thereby outperforming existing baselines and establishing a new state-of-the-art on each single open-source dataset (CoNLL 2003, GermEval 2014 and Tübingen Treebank 2018). Rather than designing deeper and wider hybrid neural architectures, we gather all available resources and perform a detailed optimization and grammar-dependent morphological processing consisting of lemmatization and part-of-speech tagging prior to exposing the raw data to any training process. We test our approach in a threefold monolingual experimental setup of a) single, b) joint, and c) optimized training and shed light on the dependency of downstream-tasks on the size of corpora used to compute word embeddings. Sajawel Ahmed, Alexander Mehler |
ICMLA | 2 |
| 2018 | A UIMA Database Interface for Managing NLP-related Text Annotations
Giuseppe Abrami, Alexander Mehler |
LREC | 2 |
| 2018 | WikiDragon: A Java Framework For Diachronic Content And Network Analysis Of MediaWikis
Rüdiger Gleim, Alexander Mehler, Sung Y. Song |
LREC | 2 |
| 2018 | TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations
Philipp Helfrich, Elias Rieb, Giuseppe Abrami, Andy Lücking, Alexander Mehler |
LREC | 5 |
| 2018 | FastSense: An Efficient Word Sense Disambiguation Classifier
Tolga Uslu, Alexander Mehler, Daniel Baumartz, Alexander Henlein, Wahed Hemati |
LREC | 2 |
| 2016 | Language classification from bilingual word embedding graphsabstractWe study the role of the second language in bilingual word embeddings in monolingual semantic evaluation tasks. We find strongly and weakly positive correlations between down-stream task performance and second language similarity to the target language. Additionally, we show how bilingual word embeddings can be employed for the task of semantic language classification and that joint semantic spaces vary in meaningful ways across second languages. Our results support the hypothesis that semantic language similarity is influenced by both structural similarity as well as geography/contact. Steffen Eger, Armin Hoenen, Alexander Mehler |
COLING | 3 |
| 2016 | TLT-CRF: A Lexicon-supported Morphological Tagger for Latin Based on Conditional Random Fields
Tim vor der Brück, Alexander Mehler |
LREC | 2 |
| 2016 | Lemmatization and Morphological Tagging in German and Latin: A Comparison and a Survey of the State-of-the-art
Steffen Eger, Rüdiger Gleim, Alexander Mehler |
LREC | 3 |
| 2016 | TGermaCorp - A (Digital) Humanities Resource for (Computational) Linguistics
Andy Lücking, Armin Hoenen, Alexander Mehler |
LREC | 3 |
| 2016 | Finding Recurrent Features of Image Schema Gestures: the FIGURE corpus
Andy Lücking, Alexander Mehler, Désirée Walther, Marcel Mauri, Dennis Kurfürst |
LREC | 2 |
| 2015 | Complex Decomposition of the Negative Distance KernelabstractA Support Vector Machine (SVM) has become a very popular machine learning method for text classification. One reason for this relates to the range of existing kernels which allow for classifying data that is not linearly separable. The linear, polynomial and RBF (Gaussian Radial Basis Function) kernel are commonly used and serve as a basis of comparison in our study. We show how to derive the primal form of the quadratic Power Kernel (PK) -- also called the Negative Euclidean Distance Kernel (NDK) -- by means of complex numbers. We exemplify the NDK in the framework of text categorization using the Dewey Document Classification (DDC) as the target scheme. Our evaluation shows that the power kernel produces F-scores that are comparable to the reference kernels, but is -- except for the linear kernel -- faster to compute. Finally, we show how to extend the NDK-approach by including the Mahalanobis distance. Tim vor der Brück, Steffen Eger, Alexander Mehler |
ICMLA | 3 |
| 2014 | Readability Classification of Bangla Texts
Zahurul Islam, Md. Rashedur Rahman, Alexander Mehler |
CICLing (2) | 3 |
| 2014 | ColLex.en: Automatically Generating and Evaluating a Full-form Lexicon for English
Tim vor der Brück, Alexander Mehler, Zahurul Islam |
LREC | 2 |
| 2013 | Using Complex Network Analysis in the Cognitive Sciences
Nicole Beckage, Michael S. Vitevitch, Alexander Mehler, Eliana Colunga |
CogSci | 3 |
| 2012 | Customization of the Europarl Corpus for Translation Studies
Zahurul Islam, Alexander Mehler |
LREC | 2 |
| 2012 | Text Readability Classification of Textbooks of a Low-Resource Language
Zahurul Islam, Alexander Mehler, Md. Rashedur Rahman |
PACLIC | 2 |
| 2012 | Assessing cognitive alignment in different types of dialog by means of a network model
Alexander Mehler, Andy Lücking, Peter Menke |
Neural Networks | 1 |
| 2011 | Assessing Lexical Alignment in Spontaneous Direction Dialogue Data by Means of a Lexicon Network Model
Alexander Mehler, Andy Lücking, Peter Menke |
CICLing (1) | 1 |
| 2011 | From neural activation to symbolic alignment: A network-based approach to the formation of dialogue lexicaabstractWe present a lexical network model, called TiTAN, that captures the formation and the structure of natural language dialogue lexica. The model creates a bridge between neural connectionist networks and symbolic architectures: On the one hand, TiTAN is driven by the neural motor of lexical alignment, namely priming. On the other hand, TiTAN accounts for observed symbolic output of interlocutors, namely uttered words. The TiTAN series update is driven by the dialogue inherent dynamics of turns and incorporates a measure of the structural similarity of graphs. This allows to apply and evaluate the model: TiTAN is tested classifying 55 experimental dialogue data according to their alignment status. The trade-off between precision and recall of the classification results in an F-score of 0.92. Alexander Mehler, Andy Lücking, Peter Menke |
IJCNN | 1 |
| 2011 | An Online Platform for Visualizing Lexical NetworksabstractThis demo paper outlines www.linguistic-networks.net -- an online platform for displaying and browsing linguistic networks as resources in text and web mining. The paper focuses on the main features of this platform, including the visualization of time series of co-occurrence patterns, and features its underlying technology. Markus Lux, Jan Laußmann, Alexander Mehler, Christian Menßen |
Web Intelligence | 3 |
| 2011 | Geography of social ontologies: Testing a variant of the Sapir-Whorf Hypothesis in the context of Wikipedia
Alexander Mehler, Olga Pustylnikov, Nils Diewald |
Comput. Speech Lang. | 1 |
| 2010 | Computational Linguistics for Mere Mortals - Powerful but Easy-to-use Linguistic Processing for Scientists in the Humanities
Rüdiger Gleim, Alexander Mehler |
LREC | 2 |
| 2010 | The Ariadne System: A Flexible and Extensible Framework for the Modeling and Storage of Experimental Data in the Humanities
Peter Menke, Alexander Mehler |
LREC | 2 |
| 2010 | eHumanities Desktop - An Architecture for Flexible Annotation in Iconographic Research
Rüdiger Gleim, Paul Warner, Alexander Mehler |
WEBIST (2) | 3 |
| 2009 | Social Semantics and Its Evaluation by Means of Semantic Relatedness and Open Topic ModelsabstractThis paper presents an approach using social semantics for the task of topic labelling by means of Open Topic Models. Our approach utilizes a social ontology to create an alignment of documents within a social network. Comprised category information is used to compute a topic generalization. We propose a feature-frequency-based method for measuring semantic relatedness which is needed in order to reduce the number of document features for the task of topic labelling. This method is evaluated against multiple human judgement experiments comprising two languages and three different resources. Overall the results show that social ontologies provide a rich source of terminological knowledge. The performance of the semantic relatedness measure with correlation values of up to .77 are quite promising. Results on the topic labelling experiment show, with an accuracy of up to .79, that our approach can be a valuable method for various NLP applications. Ulli Waltinger, Alexander Mehler |
Web Intelligence | 2 |
| 2009 | A Two-level Approach to Web Genre Classification
Ulli Waltinger, Alexander Mehler, Armin Wegner |
WEBIST | 2 |
| 2008 | A Unified Database of Dependency Treebanks: Integrating, Quantifying & Evaluating Dependency Data
Olga Pustylnikov, Alexander Mehler, Rüdiger Gleim |
LREC | 2 |
| 2008 | Towards a Reference Corpus of Web Genres for the Evaluation of Genre Identification Systems
Georg Rehm, Marina Santini, Alexander Mehler, Pavel Braslavski 0001, Rüdiger Gleim, Andrea Stubbe, Svetlana Symonenko, Mirko Tavosanis, Vedrana Vidulin |
LREC | 3 |
| 2008 | Who Is It? Context Sensitive Named Entity and Instance Recognition by Means of WikipediaabstractThis paper presents an approach for predicting context sensitive entities exemplified in the domain of person names. Our approach is based on building a weighted context but also a weighted people graph and predicting the context entity by extracting the best fitting sub graph using a spreading activation technique. The results of the experiments show a quite promising F-Measure of 0.99. Ulli Waltinger, Alexander Mehler |
Web Intelligence | 2 |
| 2008 | Towards Automatic Content Tagging - Enhanced Web Services in Digital Libraries using Lexical Chaining
Ulli Waltinger, Alexander Mehler, Gerhard Heyer |
WEBIST (2) | 2 |
| 2007 | Classification of Documents Based on the Structure of Their DOM Trees
Peter Geibel, Olga Pustylnikov, Alexander Mehler, Helmar Gust, Kai-Uwe Kühnberger |
ICONIP (2) | 3 |
| 2007 | Aisles through the Category Forest - Utilising the Wikipedia Category System for Corpus Building in Machine Learning
Rüdiger Gleim, Alexander Mehler, Matthias Dehmer, Olga Pustylnikov |
WEBIST (2) | 2 |
| 2002 | Hierarchical Orderings of Textual Units
Alexander Mehler |
COLING | 1 |