VLDB 2026 Research / reviewers in the wild / expert
Alan Akbik
dblp:127/0198
· DBLP profile ↗
27ranked-venue papers
10as first author
13since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 10 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pre-Training Curriculum for Multi-Token Prediction in Language ModelsabstractMulti-token prediction (MTP) is a recently proposed pre-training objective for language models.Rather than predicting only the next token (NTP), MTP predicts the next k tokens at each prediction step, using multiple prediction heads.MTP has shown promise in improving downstream performance, inference speed, and training efficiency, particularly for large models.However, prior work has shown that smaller language models (SLMs) struggle with the MTP objective.To address this, we propose a curriculum learning strategy for MTP training, exploring two variants: a forward curriculum, which gradually increases the complexity of the pre-training objective from NTP to MTP, and a reverse curriculum, which does the opposite.Our experiments show that the forward curriculum enables SLMs to better leverage the MTP objective during pre-training, improving downstream NTP performance and generative output quality, while retaining the benefits of self-speculative decoding.The reverse curriculum achieves stronger NTP performance and output quality, but fails to provide any selfspeculative decoding benefits. Ansar Aynetdinov, Alan Akbik |
ACL (1) | 2 |
| 2025 | Evaluating Design Decisions for Dual Encoder-based Entity DisambiguationabstractEntity disambiguation (ED) is the task of linking mentions in text to corresponding entries in a knowledge base.Dual Encoders address this by embedding mentions and label candidates in a shared embedding space and applying a similarity metric to predict the correct label.In this work, we focus on evaluating key design decisions for Dual Encoder-based ED, such as its loss function, similarity metric, label verbalization format, and negative sampling strategy.We present the resulting model VERBALIZED, a document-level Dual Encoder model that includes contextual label verbalizations and efficient hard negative sampling.Additionally, we explore an iterative prediction variant that aims to improve the disambiguation of challenging data points.To support our analysis, we first conduct comprehensive ablation experiments on specific design decisions using AIDA-Yago, followed by large-scale, multi-domain evaluation on the ZELDA benchmark. Susanna Rücker, Alan Akbik |
ACL (1) | 2 |
| 2025 | Improving Online Job Advertisement Analysis via Compositional Entity ExtractionabstractWe propose a compositional entity modeling framework for requirement extraction from Online Job Advertisements (OJAs).To more accurately capture the structure of requirements in OJAs, we reframe the task from identifying single-span annotations to modeling complex, tree-like structures that connect atomic entity types via typed relationships.Based on this schema, we introduce GOJA, a high-quality dataset of 500 German job ads.GOJA captures the internal semantics of job requirements, including roles, tools, experience levels, attitudes, and their functional context.We describe the annotation process, report strong inter-annotator agreement, and benchmark transformer models to demonstrate the feasibility of training on this structure.To illustrate the analytical potential of our approach, we present a focused case study on AI-related job requirements.We show how our proposed compositional representation enables new types of labor market analyses. Kai Krüger, Johanna Binnewitt, Kathrin Ehmann, Stefan Winnige, Alan Akbik |
EMNLP | 5 |
| 2025 | Familarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training DataabstractJonas Golde, Patrick Haller, Max Ploner, Fabio Barth, Nicolaas Jedema, Alan Akbik. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jonas Golde, Patrick Haller 0002, Max Ploner, Fabio Barth, Nicolaas Paul Jedema, Alan Akbik |
NAACL (Long Papers) | 6 |
| 2024 | PECC: Problem Extraction and Coding ChallengesabstractRecent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning. Existing benchmarks evaluate tasks in isolation, yet the extent to which LLMs can understand prose-style tasks, identify the underlying problems, and then generate appropriate code solutions is still unexplored. Addressing this gap, we introduce PECC, a novel benchmark derived from Advent Of Code (AoC) challenges and Project Euler, including 2396 problems. Unlike conventional benchmarks, PECC requires LLMs to interpret narrative-embedded problems, extract requirements, and generate executable code. A key feature of our dataset is the complexity added by natural language prompting in chat-based evaluations, mirroring real-world instruction ambiguities. Results show varying model performance between narrative and neutral problems, with specific challenges in the Euler math-based subset with GPT-3.5-Turbo passing 50% of the AoC challenges and only 8% on the Euler problems. By probing the limits of LLMs’ capabilities, our benchmark provides a framework to monitor and assess the subsequent progress of LLMs as a universal problem solver. Patrick Haller 0002, Jonas Golde, Alan Akbik |
LREC/COLING | 3 |
| 2024 | Large-Scale Label Interpretation Learning for Few-Shot Named Entity RecognitionabstractFew-shot named entity recognition (NER) detects named entities within text using only a few annotated examples.One promising line of research is to leverage natural language descriptions of each entity type: the common label PER might, for example, be verbalized as "person entity."In an initial label interpretation learning phase, the model learns to interpret such verbalized descriptions of entity types.In a subsequent few-shot tagset extension phase, this model is then given a description of a previously unseen entity type (such as "music album") and optionally a few training examples to perform few-shot NER for this type.In this paper, we systematically explore the impact of a strong semantic prior to interpret verbalizations of new entity types by massively scaling up the number and granularity of entity types used for label interpretation learning.To this end, we leverage an entity linking benchmark to create a dataset with orders of magnitude of more distinct entity types and descriptions as currently used datasets.We find that this increased signal yields strong results in zeroand few-shot NER in in-domain, cross-domain, and even cross-lingual settings.Our findings indicate significant potential for improving fewshot NER through heuristical data-based optimization. Jonas Golde, Felix Hamborg, Alan Akbik |
EACL (1) | 3 |
| 2024 | NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionabstractAvailable training data for named entity recognition (NER) often contains a significant percentage of incorrect labels for entity types and entity boundaries.Such label noise poses challenges for supervised learning and may significantly deteriorate model quality.To address this, prior work proposed various noise-robust learning approaches capable of learning from data with partially incorrect labels.These approaches are typically evaluated using simulated noise where the labels in a clean dataset are automatically corrupted.However, as we show in this paper, this leads to unrealistic noise that is far easier to handle than real noise caused by human error or semi-automatic annotation.To enable the study of the impact of various types of real noise, we introduce NOISEBENCH, an NER benchmark consisting of clean training data corrupted with 6 types of real noise, including expert errors, crowdsourcing errors, automatic annotation errors and LLM errors.We present an analysis that shows that real noise is significantly more challenging than simulated noise, and show that current state-of-the-art models for noise-robust learning fall far short of their achievable upper bound.We release NOISEBENCH for both English and German to the research community 1 . Elena Merdjanovska, Ansar Aynetdinov, Alan Akbik |
EMNLP | 3 |
| 2024 | Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer LearningabstractIntermediate task transfer learning can greatly improve model performance.If, for example, one has little training data for emotion detection, first fine-tuning a language model on a sentiment classification dataset may improve performance strongly.But which task to choose for transfer learning?Prior methods producing useful task rankings are infeasible for large source pools, as they require forward passes through all source language models.We overcome this by introducing Embedding Space Maps (ESMs), light-weight neural networks that approximate the effect of fine-tuning a language model.We conduct the largest study on NLP task transferability and task selection with 12k source-target pairs.We find that applying ESMs on a prior method reduces execution time and disk space usage by factors of 10 and 278, respectively, while retaining high selection performance (avg.regret@5 score of 2.95). David Schulte, Felix Hamborg, Alan Akbik |
EMNLP | 3 |
| 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization toolsabstractMOTIVATION: With the exponential growth of the life sciences literature, biomedical text mining (BTM) has become an essential technology for accelerating the extraction of insights from publications. The identification of entities in texts, such as diseases or genes, and their normalization, i.e. grounding them in knowledge base, are crucial steps in any BTM pipeline to enable information aggregation from multiple documents. However, tools for these two steps are rarely applied in the same context in which they were developed. Instead, they are applied "in the wild," i.e. on application-dependent text collections from moderately to extremely different from those used for training, varying, e.g. in focus, genre or text type. This raises the question whether the reported performance, usually obtained by training and evaluating on different partitions of the same corpus, can be trusted for downstream applications. RESULTS: Here, we report on the results of a carefully designed cross-corpus benchmark for entity recognition and normalization, where tools were applied systematically to corpora not used during their training. Based on a survey of 28 published systems, we selected five, based on predefined criteria like feature richness and availability, for an in-depth analysis on three publicly available corpora covering four entity types. Our results present a mixed picture and show that cross-corpus performance is significantly lower than the in-corpus performance. HunFlair2, the redesigned and extended successor of the HunFlair tool, showed the best performance on average, being closely followed by PubTator Central. Our results indicate that users of BTM tools should expect a lower performance than the original published one when applying tools in "the wild" and show that further research is necessary for more robust BTM tools. AVAILABILITY AND IMPLEMENTATION: All our models are integrated into the Natural Language Processing (NLP) framework flair: https://github.com/flairNLP/flair. Code to reproduce our results is available at: https://github.com/hu-ner/hunflair2-experiments. Mario Sänger, Samuele Garda, Xing David Wang, Leon Weber-Genzel, Pia Droop, Benedikt Fuchs, Alan Akbik, Ulf Leser |
Bioinform. | 7 |
| 2023 | ZELDA: A Comprehensive Benchmark for Supervised Entity DisambiguationabstractEntity disambiguation (ED) is the task of disambiguating named entity mentions in text to unique entries in a knowledge base.Due to its industrial relevance, as well as current progress in leveraging pre-trained language models, a multitude of ED approaches have been proposed in recent years.However, we observe a severe lack of uniformity across experimental setups in current ED work, rendering a direct comparison of approaches based solely on reported numbers impossible: Current approaches widely differ in the data set used to train, the size of the covered entity vocabulary, and the usage of additional signals such as candidate lists.To address this issue, we present ZELDA, a novel entity disambiguation benchmark that includes a unified training data set, entity vocabulary, candidate lists, as well as challenging evaluation splits covering 8 different domains.We illustrate its design and construction, and present experiments in which we train and compare current state-of-the-art approaches on our benchmark.To encourage greater direct comparability in the entity disambiguation domain, we open source our benchmark at https: //github.com/flairNLP/zelda. Marcel Milich, Alan Akbik |
EACL | 2 |
| 2023 | CleanCoNLL: A Nearly Noise-Free Named Entity Recognition DatasetabstractThe CoNLL-03 corpus is arguably the most well-known and utilized benchmark dataset for named entity recognition (NER).However, prior works found significant numbers of annotation errors, incompleteness, and inconsistencies in the data.This poses challenges to objectively comparing NER approaches and analyzing their errors, as current state-of-the-art models achieve F1-scores that are comparable to or even exceed the estimated noise level in CoNLL-03.To address this issue, we present a comprehensive relabeling effort assisted by automatic consistency checking that corrects 7.0% of all labels in the English CoNLL-03.Our effort adds a layer of entity linking annotation both for better explainability of NER labels and as additional safeguard of annotation quality.Our experimental evaluation finds not only that state-of-the-art approaches reach significantly higher F1-scores (97.1%) on our data, but crucially that the share of correct predictions falsely counted as errors due to annotation noise drops from 47% to 6%.This indicates that our resource is well suited to analyze the remaining errors made by state-of-the-art models, and that the theoretical upper bound even on high resource, coarse-grained NER is not yet reached.To facilitate such analysis, we make CLEANCONLL publicly available to the research community 1 . Susanna Rücker, Alan Akbik |
EMNLP | 2 |
| 2021 | Early Detection of Sexual Predators in ChatsabstractMatthias Vogt, Ulf Leser, Alan Akbik. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Matthias Vogt, Ulf Leser, Alan Akbik |
ACL/IJCNLP (1) | 3 |
| 2021 | HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognitionabstractSUMMARY: Named entity recognition (NER) is an important step in biomedical information extraction pipelines. Tools for NER should be easy to use, cover multiple entity types, be highly accurate and be robust toward variations in text genre and style. We present HunFlair, a NER tagger fulfilling these requirements. HunFlair is integrated into the widely used NLP framework Flair, recognizes five biomedical entity types, reaches or overcomes state-of-the-art performance on a wide set of evaluation corpora, and is trained in a cross-corpus setting to avoid corpus-specific bias. Technically, it uses a character-level language model pretrained on roughly 24 million biomedical abstracts and three million full texts. It outperforms other off-the-shelf biomedical NER tools with an average gain of 7.26 pp over the next best tool in a cross-corpus setting and achieves on-par results with state-of-the-art research prototypes in in-corpus experiments. HunFlair can be installed with a single command and is applied with only four lines of code. Furthermore, it is accompanied by harmonized versions of 23 biomedical NER corpora. AVAILABILITY AND IMPLEMENTATION: HunFlair ist freely available through the Flair NLP framework (https://github.com/flairNLP/flair) under an MIT license and is compatible with all major operating systems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leon Weber-Genzel, Mario Sänger, Jannes Münchmeyer, Maryam Habibi, Ulf Leser, Alan Akbik |
Bioinform. | 6 |
| 2020 | Task-Aware Representation of Sentences for Generic Text ClassificationabstractState-of-the-art approaches for text classification leverage a transformer architecture with a linear layer on top that outputs a class distribution for a given prediction problem.While effective, this approach suffers from conceptual limitations that affect its utility in few-shot or zero-shot transfer learning scenarios.First, the number of classes to predict needs to be pre-defined.In a transfer learning setting, in which new classes are added to an already trained classifier, all information contained in a linear layer is therefore discarded, and a new layer is trained from scratch.Second, this approach only learns the semantics of classes implicitly from training examples, as opposed to leveraging the explicit semantic information provided by the natural language names of the classes.For instance, a classifier trained to predict the topics of news articles might have classes like "business" or "sports" that themselves carry semantic information.Extending a classifier to predict a new class named "politics" with only a handful of training examples would benefit from both leveraging the semantic information in the name of a new class and using the information contained in the already trained linear layer.This paper presents a novel formulation of text classification that addresses these limitations.It imbues the notion of the task at hand into the transformer model itself by factorizing arbitrary classification problems into a generic binary classification problem.We present experiments in few-shot and zero-shot transfer learning that show that our approach significantly outperforms previous approaches on small training data and can even learn to predict new classes with no training examples at all.The implementation of our model is publicly available at Kishaloy Halder, Alan Akbik, Josip Krapac, Roland Vollgraf |
COLING | 2 |
| 2018 | Contextual String Embeddings for Sequence LabelingabstractRecent advances in language modeling using recurrent neural networks have made it viable to model language as distributions over characters. By learning to predict the next character on the basis of previous characters, such models have been shown to automatically internalize linguistic concepts such as words, sentences, subclauses and even sentiment. In this paper, we propose to leverage the internal states of a trained character language model to produce a novel type of word embedding which we refer to as contextual string embeddings. Our proposed embeddings have the distinct properties that they (a) are trained without any explicit notion of words and thus fundamentally model words as sequences of characters, and (b) are contextualized by their surrounding text, meaning that the same word will have different embeddings depending on its contextual use. We conduct a comparative evaluation against previous embeddings and find that our embeddings are highly useful for downstream tasks: across four classic sequence labeling tasks we consistently outperform the previous state-of-the-art. In particular, we significantly outperform previous work on English and German named entity recognition (NER), allowing us to report new state-of-the-art F1-scores on the CoNLL03 shared task. We release all code and pre-trained language models in a simple-to-use framework to the research community, to enable reproduction of these experiments and application of our proposed embeddings to other tasks: https://github.com/zalandoresearch/flair Alan Akbik, Duncan Blythe, Roland Vollgraf |
COLING | 1 |
| 2018 | ZAP: An Open-Source Multilingual Annotation Projection Framework
Alan Akbik, Roland Vollgraf |
LREC | 1 |
| 2018 | FEIDEGGER: A Multi-modal Corpus of Fashion Images and Descriptions in German
Leonidas Lefakis, Alan Akbik, Roland Vollgraf |
LREC | 2 |
| 2017 | CROWD-IN-THE-LOOP: A Hybrid Approach for Annotating Semantic RolesabstractCrowdsourcing has proven to be an effective method for generating labeled data for a range of NLP tasks.However, multiple recent attempts of using crowdsourcing to generate gold-labeled training data for semantic role labeling (SRL) reported only modest results, indicating that SRL is perhaps too difficult a task to be effectively crowdsourced.In this paper, we postulate that while producing SRL annotation does require expert involvement in general, a large subset of SRL labeling tasks is in fact appropriate for the crowd.We present a novel workflow in which we employ a classifier to identify difficult annotation tasks and route each task either to experts or crowd workers according to their difficulties.Our experimental evaluation shows that the proposed approach reduces the workload for experts by over two-thirds, and thus significantly reduces the cost of producing SRL annotation at little loss in quality. Chenguang Wang 0001, Alan Akbik, Laura Chiticariu, Yunyao Li 0001, Anbang Xu |
EMNLP | 2 |
| 2016 | Multilingual Aliasing for Auto-Generating Proposition BanksabstractSemantic Role Labeling (SRL) is the task of identifying the predicate-argument structure in sentences with semantic frame and role labels. For the English language, the Proposition Bank provides both a lexicon of all possible semantic frames and large amounts of labeled training data. In order to expand SRL beyond English, previous work investigated automatic approaches based on parallel corpora to automatically generate Proposition Banks for new target languages (TLs). However, this approach heuristically produces the frame lexicon from word alignments, leading to a range of lexicon-level errors and inconsistencies. To address these issues, we propose to manually alias TL verbs to existing English frames. For instance, the German verb drehen may evoke several meanings, including “turn something” and “film something”. Accordingly, we alias the former to the frame TURN.01 and the latter to a group of frames that includes FILM.01 and SHOOT.03. We execute a large-scale manual aliasing effort for three target languages and apply the new lexicons to automatically generate large Proposition Banks for Chinese, French and German with manually curated frames. We present a detailed evaluation in which we find that our proposed approach significantly increases the quality and consistency of the generated Proposition Banks. We release these resources to the research community. Alan Akbik, Yunyao Li 0001 |
COLING | 1 |
| 2016 | K-SRL: Instance-based Learning for Semantic Role LabelingabstractSemantic role labeling (SRL) is the task of identifying and labeling predicate-argument structures in sentences with semantic frame and role labels. A known challenge in SRL is the large number of low-frequency exceptions in training data, which are highly context-specific and difficult to generalize. To overcome this challenge, we propose the use of instance-based learning that performs no explicit generalization, but rather extrapolates predictions from the most similar instances in the training data. We present a variant of k-nearest neighbors (kNN) classification with composite features to identify nearest neighbors for SRL. We show that high-quality predictions can be derived from a very small number of similar instances. In a comparative evaluation we experimentally demonstrate that our instance-based learning approach significantly outperforms current state-of-the-art systems on both in-domain and out-of-domain data, reaching F1-scores of 89,28% and 79.91% respectively. Alan Akbik, Yunyao Li 0001 |
COLING | 1 |
| 2016 | Towards Semi-Automatic Generation of Proposition Banks for Low-Resource LanguagesabstractAnnotation projection based on parallel corpora has shown great promise in inexpensively creating Proposition Banks for languages for which high-quality parallel corpora and syntactic parsers are available.In this paper, we present an experimental study where we apply this approach to three languages that lack such resources: Tamil, Bengali and Malayalam.We find an average quality difference of 6 to 20 absolute F-measure points vis-avis high-resource languages, which indicates that annotation projection alone is insufficient in low-resource scenarios.Based on these results, we explore the possibility of using annotation projection as a starting point for inexpensive data curation involving both experts and non-experts.We give an outline of what such a process may look like and present an initial study to discuss its potential and challenges. Alan Akbik, Vishwajeet Kumar, Yunyao Li 0001 |
EMNLP | 1 |
| 2015 | Generating High Quality Proposition Banks for Multilingual Semantic Role LabelingabstractAlan Akbik, Laura Chiticariu, Marina Danilevsky, Yunyao Li, Shivakumar Vaithyanathan, Huaiyu Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Alan Akbik, Laura Chiticariu, Marina Danilevsky, Yunyao Li 0001, Shivakumar Vaithyanathan, Huaiyu Zhu 0001 |
ACL (1) | 1 |
| 2014 | Exploratory Relation Extraction in Large Text Corpora
Alan Akbik, Thilo Michael, Christoph Boden |
COLING | 1 |
| 2014 | The Weltmodell: A Data-Driven Commonsense Knowledge Base
Alan Akbik, Thilo Michael |
LREC | 1 |
| 2014 | Freepal: A Large Collection of Deep Lexico-Syntactic Patterns for Relation Extraction
Johannes Kirschnick, Alan Akbik, Holmer Hemsen |
LREC | 2 |
| 2013 | Effective Selectional Restrictions for Unsupervised Relation Extraction
Alan Akbik, Larysa Visengeriyeva, Johannes Kirschnick, Alexander Löser |
IJCNLP | 1 |
| 2012 | Unsupervised Discovery of Relations and Discriminative Extraction Patterns
Alan Akbik, Larysa Visengeriyeva, Priska Herger, Holmer Hemsen, Alexander Löser |
COLING | 1 |