VLDB 2026 Research / reviewers in the wild / expert
Samuele Garda
dblp:229/9210
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2024
0009-0002-8234-8299ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 74% Medical and health informatics · 13% Computational science and engineering · 13% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 55% Knowledge representation and reasoning · 45% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › biomedical text mining
biomedical entity linking |
1.4 | 2 | 2024 | BELHD: improving biomedical entity linking with homonym disambiguation · Bioinform. 2024 BELB: a biomedical entity linking benchmark · Bioinform. 2023 |
Bioinformatics and computational biology
biomedical text mining |
1.4 | 2 | 2024 | BELHD: improving biomedical entity linking with homonym disambiguation · Bioinform. 2024 BELB: a biomedical entity linking benchmark · Bioinform. 2023 |
Natural language and speech › Information extraction and text analysis › named entity recognition
biomedical named entity recognition |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Natural language and speech › Information extraction and text analysis
entity normalization |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
knowledge base grounding |
0.8 | 1 | 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization tools · Bioinform. 2024 |
Medical and health informatics
biomedical natural language processing |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Computational science and engineering
multi-task learning |
0.6 | 1 | 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language Processing · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
dense retrieval · 0.8candidate sharing · 0.8autoregressive modeling · 0.8rule-based entity linking · 0.7pre-trained language model · 0.7language prompting · 0.6instruction tuning · 0.6autoregressive sequence labelling · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | BELHD: improving biomedical entity linking with homonym disambiguationabstractMOTIVATION: Biomedical entity linking (BEL) is the task of grounding entity mentions to a given knowledge base (KB). Recently, neural name-based methods, system identifying the most appropriate name in the KB for a given mention using neural network (either via dense retrieval or autoregressive modeling), achieved remarkable results for the task, without requiring manual tuning or definition of domain/entity-specific rules. However, as name-based methods directly return KB names, they cannot cope with homonyms, i.e. different KB entities sharing the exact same name. This significantly affects their performance for KBs where homonyms account for a large amount of entity mentions (e.g. UMLS and NCBI Gene). RESULTS: We present BELHD (Biomedical Entity Linking with Homonym Disambiguation), a new name-based method that copes with this challenge. BELHD builds upon the BioSyn model with two crucial extensions. First, it performs pre-processing of the KB, during which it expands homonyms with a specifically constructed disambiguating string, thus enforcing unique linking decisions. Second, it introduces candidate sharing, a novel strategy that strengthens the overall training signal by including similar mentions from the same document as positive or negative examples, according to their corresponding KB identifier. Experiments with 10 corpora and 5 entity types show that BELHD improves upon current neural state-of-the-art approaches, achieving the best results in 6 out of 10 corpora with an average improvement of 4.55pp recall@1. Furthermore, the KB preprocessing is orthogonal to the prediction model and thus can also improve other neural methods, which we exemplify for GenBioEL, a generative name-based BEL approach. AVAILABILITY AND IMPLEMENTATION: The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belhd. Samuele Garda, Ulf Leser |
Bioinform. | 1 |
| 2024 | HunFlair2 in a cross-corpus evaluation of biomedical named entity recognition and normalization toolsabstractMOTIVATION: With the exponential growth of the life sciences literature, biomedical text mining (BTM) has become an essential technology for accelerating the extraction of insights from publications. The identification of entities in texts, such as diseases or genes, and their normalization, i.e. grounding them in knowledge base, are crucial steps in any BTM pipeline to enable information aggregation from multiple documents. However, tools for these two steps are rarely applied in the same context in which they were developed. Instead, they are applied "in the wild," i.e. on application-dependent text collections from moderately to extremely different from those used for training, varying, e.g. in focus, genre or text type. This raises the question whether the reported performance, usually obtained by training and evaluating on different partitions of the same corpus, can be trusted for downstream applications. RESULTS: Here, we report on the results of a carefully designed cross-corpus benchmark for entity recognition and normalization, where tools were applied systematically to corpora not used during their training. Based on a survey of 28 published systems, we selected five, based on predefined criteria like feature richness and availability, for an in-depth analysis on three publicly available corpora covering four entity types. Our results present a mixed picture and show that cross-corpus performance is significantly lower than the in-corpus performance. HunFlair2, the redesigned and extended successor of the HunFlair tool, showed the best performance on average, being closely followed by PubTator Central. Our results indicate that users of BTM tools should expect a lower performance than the original published one when applying tools in "the wild" and show that further research is necessary for more robust BTM tools. AVAILABILITY AND IMPLEMENTATION: All our models are integrated into the Natural Language Processing (NLP) framework flair: https://github.com/flairNLP/flair. Code to reproduce our results is available at: https://github.com/hu-ner/hunflair2-experiments. Mario Sänger, Samuele Garda, Xing David Wang, Leon Weber-Genzel, Pia Droop, Benedikt Fuchs, Alan Akbik, Ulf Leser |
Bioinform. | 2 |
| 2023 | BELB: a biomedical entity linking benchmarkabstractMOTIVATION: Biomedical entity linking (BEL) is the task of grounding entity mentions to a knowledge base (KB). It plays a vital role in information extraction pipelines for the life sciences literature. We review recent work in the field and find that, as the task is absent from existing benchmarks for biomedical text mining, different studies adopt different experimental setups making comparisons based on published numbers problematic. Furthermore, neural systems are tested primarily on instances linked to the broad coverage KB UMLS, leaving their performance to more specialized ones, e.g. genes or variants, understudied. RESULTS: We therefore developed BELB, a biomedical entity linking benchmark, providing access in a unified format to 11 corpora linked to 7 KBs and spanning six entity types: gene, disease, chemical, species, cell line, and variant. BELB greatly reduces preprocessing overhead in testing BEL systems on multiple corpora offering a standardized testbed for reproducible experiments. Using BELB, we perform an extensive evaluation of six rule-based entity-specific systems and three recent neural approaches leveraging pre-trained language models. Our results reveal a mixed picture showing that neural approaches fail to perform consistently across entity types, highlighting the need of further studies towards entity-agnostic models. AVAILABILITY AND IMPLEMENTATION: The source code of BELB is available at: https://github.com/sg-wbi/belb. The code to reproduce our experiments can be found at: https://github.com/sg-wbi/belb-exp. Samuele Garda, Leon Weber-Genzel, Robert Martin, Ulf Leser |
Bioinform. | 1 |
| 2022 | BigBio: A Framework for Data-Centric Biomedical Natural Language ProcessingabstractTraining and evaluating language models increasingly requires the construction of meta-datasets -- diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-shot generalization by transforming existing, supervised datasets into a variety of novel instruction tuning tasks, highlighting the benefits of meta-dataset curation. While successful in general-domain text, translating these data-centric approaches to biomedical language modeling remains challenging, as labeled biomedical datasets are significantly underrepresented in popular data hubs. To address this challenge, we introduce BigBio a community library of 126+ biomedical NLP datasets, currently covering 13 task categories and 10+ languages. BigBio facilitates reproducible meta-dataset curation via programmatic access to datasets and their metadata, and is compatible with current platforms for prompt engineering and end-to-end few/zero shot language model evaluation. We discuss our process for task schema harmonization, data auditing, contribution guidelines, and outline two illustrative use cases: zero-shot evaluation of biomedical prompts and large-scale, multi-task learning. BigBio is an ongoing community effort and is available at https://github.com/bigscience-workshop/biomedical Jason Alan Fries, Leon Weber-Genzel, Natasha Seelam, Gabriel Altay, Debajyoti Datta, Samuele Garda, Sunny Kang, Rosaline Su, Wojciech Kusa, Samuel Cahyawijaya, Fabio Barth, Simon Ott, Matthias Samwald, Stephen H. Bach, Stella Biderman, Mario Sänger, Bo Wang 0044, Alison Callahan, Daniel León Periñán, Théo Gigant, Patrick Haller 0002, Jenny Chim, José D. Posada, John M. Giorgi, Karthik Rangasai Sivaraman, Marc Pàmies, Marianna Nezhurina, Robert Martin, Michael Cullan, Moritz Freidank, Nathan Dahlberg, Shubhanshu Mishra, Shamik Bose, Nicholas Broad, Yanis Labrak, Shlok Deshmukh, Sid Kiblawi, Ayush Singh, Minh Chien Vu, Trishala Neeraj, Jonas Golde, Albert Villanova del Moral, Benjamin Beilharz |
NeurIPS | 6 |
| 2021 | Extend, don't rebuild: Phrasing conditional graph modification as autoregressive sequence labellingabstractPhrasing conditional graph modification as autoregressive sequence labelling. Leon Weber-Genzel, Jannes Münchmeyer, Samuele Garda, Ulf Leser |
EMNLP (1) | 3 |