VLDB 2026 Research / reviewers in the wild / expert
Leon Weber-Genzel
dblp:209/7969 · also Leon Weber
· DBLP profile ↗
13ranked-venue papers
7as first author
9since 2021 · last 2024
0000-0002-2499-472XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 32% Information extraction and text analysis · 25% Knowledge representation and reasoning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 75% Computational science and engineering · 14% Medical and health informatics · 11% |
Topics — the 25 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
biomedical text mining |
1.3 | 2 | 2023 | PEDL+: protein-centered relation extraction from PubMed at your fingertip · Bioinform. 2023 BELB: a biomedical entity linking benchmark · Bioinform. 2023 |
Bioinformatics and computational biology › biomedical text mining
biomedical named entity recognition |
1.2 | 3 | 2021 | HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognition · Bioinform. 2021 HUNER: improving biomedical NER with pretraining · Bioinform. 2020 Deep learning with word embeddings improves biomedical named entity recognition · Bioinform. 2017 |
Natural language and speech › Information extraction and text analysis › data annotation
annotation error detection |
0.8 | 1 | 2024 | VariErr NLI: Separating Annotation Error from Human Label Variation · ACL (1) 2024 |
Natural language and speech › Information extraction and text analysis › named entity recognition
biomedical named entity recognition |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Natural language and speech › Information extraction and text analysis
entity normalization |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base grounding |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.8 | 1 | 2024 | VariErr NLI: Separating Annotation Error from Human Label Variation · ACL (1) 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.7 | 1 | 2023 | Establishing Trustworthiness: Rethinking Tasks and Model Evaluation · EMNLP 2023 |
Bioinformatics and computational biology › biomedical text mining
biomedical entity linking |
0.7 | 1 | 2023 | BELB: a biomedical entity linking benchmark · Bioinform. 2023 |
Natural language and speech › Language models and text generation
multilingual dataset |
0.6 | 1 | 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Natural language and speech › Language models and text generation
multilingual language models |
0.6 | 1 | 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › data curation
training data curation |
0.6 | 1 | 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset · NeurIPS 2022 |
Medical and health informatics
biomedical natural language processing |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Computational science and engineering
multi-task learning |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Bioinformatics and computational biology › biomedical text mining
relation extraction |
0.4 | 1 | 2020 | PEDL: extracting protein-protein associations using deep language models and distant supervision · Bioinform. 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic programming |
0.4 | 1 | 2019 | NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language · ACL (1) 2019 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.4 | 1 | 2019 | NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.4 | 1 | 2019 | NLProlog: Reasoning with Weak Unification for Question Answering in Natural Language · ACL (1) 2019 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.3 | 2 | 2021 | HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognition · Bioinform. 2021 HUNER: improving biomedical NER with pretraining · Bioinform. 2020 |
Computational science and engineering › scientific data management
database curation |
0.2 | 1 | 2023 | PEDL+: protein-centered relation extraction from PubMed at your fingertip · Bioinform. 2023 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway reconstruction |
0.2 | 1 | 2023 | PEDL+: protein-centered relation extraction from PubMed at your fingertip · Bioinform. 2023 |
Natural language and speech › Language models and text generation › language modeling
character-level language modeling |
0.1 | 1 | 2021 | HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognition · Bioinform. 2021 |
Natural language and speech › Language models and text generation › pre-trained language model
biomedical language model |
0.1 | 1 | 2020 | HUNER: improving biomedical NER with pretraining · Bioinform. 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
biomedical relation extraction |
0.1 | 1 | 2020 | PEDL: extracting protein-protein associations using deep language models and distant supervision · Bioinform. 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.1 | 1 | 2020 | PEDL: extracting protein-protein associations using deep language models and distant supervision · Bioinform. 2020 |
Methods — techniques the papers use, named apart from their topics
LSTM-CRF · 1.2cross-corpus training · 1.0character-level language model · 1.0automatic error detection · 0.8GPT-4 · 0.8rule-based entity linking · 0.7ranking · 0.7pre-trained language model · 0.7natural language processing · 0.7holistic evaluation · 0.7filtering · 0.7language prompting · 0.6instruction tuning · 0.6autoregressive sequence labelling · 0.5supervised pretraining · 0.4semi-supervised pretraining · 0.4deep language model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | VariErr NLI: Separating Annotation Error from Human Label VariationabstractHuman label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons.These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation.To the best of our knowledge, there exists no prior work that focuses on teasing apart error from signal, especially in cases where signal is beyond black-and-white.To fill this gap, we introduce a systematic methodology and a new dataset, VARIERR (variation versus error), focusing on the NLI task in English.We propose a 2-round annotation procedure with annotators explaining each label and subsequently judging the validity of label-explanation pairs.VARIERR contains 7,732 validity judgments on 1,933 explanations for 500 re-annotated MNLI items.We assess the effectiveness of various automatic error detection (AED) methods and GPTs in uncovering errors versus human label variation.We find that state-of-the-art AED methods significantly underperform GPTs and humans.While GPT-4 is the best system, it still falls short of human performance.Our methodology is applicable beyond NLI, offering fertile ground for future research on error versus plausible variation, which in turn can yield better and more trustworthy NLP systems. Leon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe, Barbara Plank |
ACL (1) | 1 |
| 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization toolsabstractMOTIVATION: With the exponential growth of the life sciences literature, biomedical text mining (BTM) has become an essential technology for accelerating the extraction of insights from publications. The identification of entities in texts, such as diseases or genes, and their normalization, i.e. grounding them in knowledge base, are crucial steps in any BTM pipeline to enable information aggregation from multiple documents. However, tools for these two steps are rarely applied in the same context in which they were developed. Instead, they are applied "in the wild," i.e. on application-dependent text collections from moderately to extremely different from those used for training, varying, e.g. in focus, genre or text type. This raises the question whether the reported performance, usually obtained by training and evaluating on different partitions of the same corpus, can be trusted for downstream applications. RESULTS: Here, we report on the results of a carefully designed cross-corpus benchmark for entity recognition and normalization, where tools were applied systematically to corpora not used during their training. Based on a survey of 28 published systems, we selected five, based on predefined criteria like feature richness and availability, for an in-depth analysis on three publicly available corpora covering four entity types. Our results present a mixed picture and show that cross-corpus performance is significantly lower than the in-corpus performance. HunFlair2, the redesigned and extended successor of the HunFlair tool, showed the best performance on average, being closely followed by PubTator Central. Our results indicate that users of BTM tools should expect a lower performance than the original published one when applying tools in "the wild" and show that further research is necessary for more robust BTM tools. AVAILABILITY AND IMPLEMENTATION: All our models are integrated into the Natural Language Processing (NLP) framework flair: https://github.com/flairNLP/flair. Code to reproduce our results is available at: https://github.com/hu-ner/hunflair2-experiments. Mario Sänger, Samuele Garda, Xing David Wang, Leon Weber-Genzel, Pia Droop, Benedikt Fuchs, Alan Akbik, Ulf Leser |
Bioinform. | 4 |
| 2023 | Establishing Trustworthiness: Rethinking Tasks and Model EvaluationabstractLanguage understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades.Traditionally, facets of linguistic intelligence have been compartmentalized into tasks with specialized model architectures and corresponding evaluation protocols.With the advent of large language models (LLMs) the community has witnessed a dramatic shift towards general purpose, task-agnostic approaches powered by generative models.As a consequence, the traditional compartmentalized notion of language tasks is breaking down, followed by an increasing challenge for evaluation and analysis.At the same time, LLMs are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups, increasing the need for trustworthy and reliable systems.Therefore, we argue that it is time to rethink what constitutes tasks and model evaluation in NLP, and pursue a more holistic view on language, placing trustworthiness at the center.Towards this goal, we review existing compartmentalized approaches for understanding the origins of a model's functional capacity, and provide recommendations for more multifaceted evaluation protocols."Trust arises from knowledge of origin as well as from knowledge of functional capacity." Robert Litschko, Max Müller-Eberstein, Rob van der Goot, Leon Weber-Genzel, Barbara Plank |
EMNLP | 4 |
| 2023 | BELB: a biomedical entity linking benchmarkabstractMOTIVATION: Biomedical entity linking (BEL) is the task of grounding entity mentions to a knowledge base (KB). It plays a vital role in information extraction pipelines for the life sciences literature. We review recent work in the field and find that, as the task is absent from existing benchmarks for biomedical text mining, different studies adopt different experimental setups making comparisons based on published numbers problematic. Furthermore, neural systems are tested primarily on instances linked to the broad coverage KB UMLS, leaving their performance to more specialized ones, e.g. genes or variants, understudied. RESULTS: We therefore developed BELB, a biomedical entity linking benchmark, providing access in a unified format to 11 corpora linked to 7 KBs and spanning six entity types: gene, disease, chemical, species, cell line, and variant. BELB greatly reduces preprocessing overhead in testing BEL systems on multiple corpora offering a standardized testbed for reproducible experiments. Using BELB, we perform an extensive evaluation of six rule-based entity-specific systems and three recent neural approaches leveraging pre-trained language models. Our results reveal a mixed picture showing that neural approaches fail to perform consistently across entity types, highlighting the need of further studies towards entity-agnostic models. AVAILABILITY AND IMPLEMENTATION: The source code of BELB is available at: https://github.com/sg-wbi/belb. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belb-exp. Samuele Garda, Leon Weber-Genzel, Robert Martin, Ulf Leser |
Bioinform. | 2 |
| 2023 | PEDL+: protein-centered relation extraction from PubMed at your fingertipabstractSUMMARY: Relation extraction (RE) from large text collections is an important tool for database curation, pathway reconstruction, or functional omics data analysis. In practice, RE often is part of a complex data analysis pipeline requiring specific adaptations like restricting the types of relations or the set of proteins to be considered. However, current systems are either non-programmable web sites or research code with fixed functionality. We present PEDL+, a user-friendly tool for extracting protein-protein and protein-chemical associations from PubMed articles. PEDL+ combines state-of-the-art NLP technology with adaptable ranking and filtering options and can easily be integrated into analysis pipelines. We evaluated PEDL+ in two pathway curation projects and found that 59% to 80% of its extractions were helpful. AVAILABILITY AND IMPLEMENTATION: PEDL+ is freely available at https://github.com/leonweber/pedl. Leon Weber-Genzel, Fabio Barth, Leonie J. Lorenz, Fabian Konrath, Kirsten Huska, Jana Wolf, Ulf Leser |
Bioinform. | 1 |
| 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language ProcessingabstractTraining and evaluating language models increasingly requires the construction of meta-datasets -- diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-shot generalization by transforming existing, supervised datasets into a variety of novel instruction tuning tasks, highlighting the benefits of meta-dataset curation. While successful in general-domain text, translating these data-centric approaches to biomedical language modeling remains challenging, as labeled biomedical datasets are significantly underrepresented in popular data hubs. To address this challenge, we introduce BigBio a community library of 126+ biomedical NLP datasets, currently covering 13 task categories and 10+ languages. BigBio facilitates reproducible meta-dataset curation via programmatic access to datasets and their metadata, and is compatible with current platforms for prompt engineering and end-to-end few/zero shot language model evaluation. We discuss our process for task schema harmonization, data auditing, contribution guidelines, and outline two illustrative use cases: zero-shot evaluation of biomedical prompts and large-scale, multi-task learning. BigBio is an ongoing community effort and is available at https://github.com/bigscience-workshop/biomedical Jason Alan Fries, Leon Weber-Genzel, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen H. Bach, Stella Biderman, Mario Sänger, Bo Wang 0044, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller 0002, Jenny Chim, José D. Posada, John M. Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, Benjamin Beilharz |
NeurIPS | 2 |
| 2022 | The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual DatasetabstractAs language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the foreground. This paper documents the data creation and curation efforts undertaken by BigScience to assemble the Responsible Open-science Open-collaboration Text Sources (ROOTS) corpus, a 1.6TB dataset spanning 59 languages that was used to train the 176-billion-parameter BigScience Large Open-science Open-access Multilingual (BLOOM) language model. We further release a large initial subset of the corpus and analyses thereof, and hope to empower large-scale monolingual and multilingual modeling projects with both the data and the processing tools, as well as stimulate research around this large multilingual corpus. Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo G. Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa 0001, Paulo Villegas, Tristan Thrush, Shayne Longpre, Sebastian Nagel 0005, Leon Weber-Genzel, Manuel Muñoz, Daniel van Strien, Zaid Alyafeai, Khalid Almubarak, Minh Chien Vu, Itziar Gonzalez-Dios, Aitor Soroa, Kyle Lo, Manan Dey, Pedro Ortiz Suarez, Aaron Gokaslan, Shamik Bose, David Ifeoluwa Adelani, Long Phan, Hieu Tran, Ian Yu, Suhas Pai, Jenny Chim, Violette Lepercq, Suzana Ilic, Margaret Mitchell, Sasha Luccioni, Yacine Jernite |
NeurIPS | 30 |
| 2021 | Extend, don't rebuild: Phrasing conditional graph modification as autoregressive sequence labellingabstractPhrasing conditional graph modification as autoregressive sequence labelling. Leon Weber-Genzel, Jannes Münchmeyer, Samuele Garda, Ulf Leser |
EMNLP (1) | 1 |
| 2021 | HunFlair: an easy-to-use tool for state-of-the-art biomedical named entity recognitionabstractSUMMARY: Named entity recognition (NER) is an important step in biomedical information extraction pipelines. Tools for NER should be easy to use, cover multiple entity types, be highly accurate and be robust toward variations in text genre and style. We present HunFlair, a NER tagger fulfilling these requirements. HunFlair is integrated into the widely used NLP framework Flair, recognizes five biomedical entity types, reaches or overcomes state-of-the-art performance on a wide set of evaluation corpora, and is trained in a cross-corpus setting to avoid corpus-specific bias. Technically, it uses a character-level language model pretrained on roughly 24 million biomedical abstracts and three million full texts. It outperforms other off-the-shelf biomedical NER tools with an average gain of 7.26 pp over the next best tool in a cross-corpus setting and achieves on-par results with state-of-the-art research prototypes in in-corpus experiments. HunFlair can be installed with a single command and is applied with only four lines of code. Furthermore, it is accompanied by harmonized versions of 23 biomedical NER corpora. AVAILABILITY AND IMPLEMENTATION: HunFlair ist freely available through the Flair NLP framework (https://github.com/flairNLP/flair) under an MIT license and is compatible with all major operating systems. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leon Weber-Genzel, Mario Sänger, Jannes Münchmeyer, Maryam Habibi, Ulf Leser, Alan Akbik |
Bioinform. | 1 |
| 2020 | HUNER: improving biomedical NER with pretrainingabstractMOTIVATION: Several recent studies showed that the application of deep neural networks advanced the state-of-the-art in named entity recognition (NER), including biomedical NER. However, the impact on performance and the robustness of improvements crucially depends on the availability of sufficiently large training corpora, which is a problem in the biomedical domain with its often rather small gold standard corpora. RESULTS: We evaluate different methods for alleviating the data sparsity problem by pretraining a deep neural network (LSTM-CRF), followed by a rather short fine-tuning phase focusing on a particular corpus. Experiments were performed using 34 different corpora covering five different biomedical entity types, yielding an average increase in F1-score of ∼2 pp compared to learning without pretraining. We experimented both with supervised and semi-supervised pretraining, leading to interesting insights into the precision/recall trade-off. Based on our results, we created the stand-alone NER tool HUNER incorporating fully trained models for five entity types. On the independent CRAFT corpus, which was not used for creating HUNER, it outperforms the state-of-the-art tools GNormPlus and tmChem by 5-13 pp on the entity types chemicals, species and genes. AVAILABILITY AND IMPLEMENTATION: HUNER is freely available at https://hu-ner.github.io. HUNER comes in containers, making it easy to install and use, and it can be applied off-the-shelf to arbitrary texts. We also provide an integrated tool for obtaining and converting all 34 corpora used in our evaluation, including fixed training, development and test splits to enable fair comparisons in the future. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leon Weber-Genzel, Jannes Münchmeyer, Tim Rocktäschel, Maryam Habibi, Ulf Leser |
Bioinform. | 1 |
| 2020 | PEDL: extracting protein-protein associations using deep language models and distant supervisionabstractMOTIVATION: A significant portion of molecular biology investigates signalling pathways and thus depends on an up-to-date and complete resource of functional protein-protein associations (PPAs) that constitute such pathways. Despite extensive curation efforts, major pathway databases are still notoriously incomplete. Relation extraction can help to gather such pathway information from biomedical publications. Current methods for extracting PPAs typically rely exclusively on rare manually labelled data which severely limits their performance. RESULTS: We propose PPA Extraction with Deep Language (PEDL), a method for predicting PPAs from text that combines deep language models and distant supervision. Due to the reliance on distant supervision, PEDL has access to an order of magnitude more training data than methods solely relying on manually labelled annotations. We introduce three different datasets for PPA prediction and evaluate PEDL for the two subtasks of predicting PPAs between two proteins, as well as identifying the text spans stating the PPA. We compared PEDL with a recently published state-of-the-art model and found that on average PEDL performs better in both tasks on all three datasets. An expert evaluation demonstrates that PEDL can be used to predict PPAs that are missing from major pathway databases and that it correctly identifies the text spans supporting the PPA. AVAILABILITY AND IMPLEMENTATION: PEDL is freely available at https://github.com/leonweber/pedl. The repository also includes scripts to generate the used datasets and to reproduce the experiments from this article. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Leon Weber-Genzel, Kirsten Thobe, Oscar Arturo Migueles Lozano, Jana Wolf, Ulf Leser |
Bioinform. | 1 |
| 2019 | NLProlog: Reasoning with Weak Unification for Question Answering in Natural LanguageabstractRule-based models are attractive for various tasks because they inherently lead to interpretable and explainable decisions and can easily incorporate prior knowledge.However, such systems are difficult to apply to problems involving natural language, due to its linguistic variability.In contrast, neural models can cope very well with ambiguity by learning distributed representations of words and their composition from data, but lead to models that are difficult to interpret.In this paper, we describe a model combining neural networks with logic programming in a novel manner for solving multi-hop reasoning tasks over natural language.Specifically, we propose to use a Prolog prover which we extend to utilize a similarity function over pretrained sentence encoders.We fine-tune the representations for the similarity function via backpropagation.This leads to a system that can apply rulebased reasoning to natural language, and induce domain-specific rules from training data.We evaluate the proposed system on two different question answering tasks, showing that it outperforms two baselines -BIDAF (Seo et al., 2016a) and FASTQA (Weissenborn et al., 2017b) on a subset of the WIKIHOP corpus and achieves competitive results on the MEDHOP data set (Welbl et al., 2017). Leon Weber-Genzel, Pasquale Minervini, Jannes Münchmeyer, Ulf Leser, Tim Rocktäschel |
ACL (1) | 1 |
| 2017 | Deep learning with word embeddings improves biomedical named entity recognitionabstractMOTIVATION: Text mining has become an important tool for biomedical research. The most fundamental text-mining task is the recognition of biomedical named entities (NER), such as genes, chemicals and diseases. Current NER methods rely on pre-defined features which try to capture the specific surface properties of entity types, properties of the typical local context, background knowledge, and linguistic information. State-of-the-art tools are entity-specific, as dictionaries and empirically optimal feature sets differ between entity types, which makes their development costly. Furthermore, features are often optimized for a specific gold standard corpus, which makes extrapolation of quality measures difficult. RESULTS: We show that a completely generic method based on deep learning and statistical word embeddings [called long short-term memory network-conditional random field (LSTM-CRF)] outperforms state-of-the-art entity-specific NER tools, and often by a large margin. To this end, we compared the performance of LSTM-CRF on 33 data sets covering five different entity classes with that of best-of-class NER tools and an entity-agnostic CRF implementation. On average, F1-score of LSTM-CRF is 5% above that of the baselines, mostly due to a sharp increase in recall. AVAILABILITY AND IMPLEMENTATION: The source code for LSTM-CRF is available at https://github.com/glample/tagger and the links to the corpora are available at https://corposaurus.github.io/corpora/ . CONTACT: [email protected]. Maryam Habibi, Leon Weber-Genzel, Mariana L. Neves, David Luis Wiegandt, Ulf Leser |
Bioinform. | 2 |