VLDB 2026 Research / reviewers in the wild / expert
Robert Bossy
dblp:12/2412
· DBLP profile ↗
11ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-6652-9319ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LifeCLEF 2026 Teaser: AI Challenges for Biodiversity Understanding and Ecosystem Management
Alexis Joly, Lukás Picek, Stefan Kahl, Hervé Goëau, Lukás Adam, Robert Bossy, Kostas Papafitsoros, Vojtech Cermák, Holger Klinck, Willem-Pier Vellinga, Robert Planqué, Tom Denton, Laura Chrobak, Kevin Barnard, Claire Nedellec, Louise Deléger, Marine Courtin, Giulio Martellucci, Fabrice Vinatier, Pierre Bonnet |
ECIR (4) | 6 |
| 2026 | EPOP: A Benchmark Corpus for Assessing NLP Models on Structured Information Extraction in Plant Health
Claire Nedellec, Marine Courtin, Xinzhi Yao, Marie Grosdidier, Isabelle Pieretti, Sandy Duperier, Robert Bossy |
LREC | 7 |
| 2024 | Semantically-Informed Domain Adaptation for Named Entity Recognition
Mariya Borovikova, Arnaud Ferré, Robert Bossy, Mathieu Roche, Claire Nedellec |
ISMIS | 3 |
| 2024 | Exploiting Graph Embeddings from Knowledge Bases for Neural Biomedical Relation Extraction
Anfu Tang, Louise Deléger, Robert Bossy, Pierre Zweigenbaum, Claire Nedellec |
NLDB (1) | 3 |
| 2023 | Could KeyWord Masking Strategy Improve Language Model?
Mariya Borovikova, Arnaud Ferré, Robert Bossy, Mathieu Roche, Claire Nedellec |
NLDB | 3 |
| 2020 | Handling Entity Normalization with no Annotated Corpus: Weakly Supervised Methods Based on Distributional Representation and Ontological InformationabstractEntity normalization (or entity linking) is an important subtask of information extraction that links entity mentions in text to categories or concepts in a reference vocabulary. Machine learning based normalization methods have good adaptability as long as they have enough training data per reference with a sufficient quality. Distributional representations are commonly used because of their capacity to handle different expressions with similar meanings. However, in specific technical and scientific domains, the small amount of training data and the relatively small size of specialized corpora remain major challenges. Recently, the machine learning-based CONTES method has addressed these challenges for reference vocabularies that are ontologies, as is often the case in life sciences and biomedical domains. And yet, its performance is dependent on manually annotated corpus. Furthermore, like other machine learning based methods, parametrization remains tricky. We propose a new approach to address the scarcity of training data that extends the CONTES method by corpus selection, pre-processing and weak supervision strategies, which can yield high-performance results without any manually annotated examples. We also study which hyperparameters are most influential, with sometimes different patterns compared to previous work. The results show that our approach significantly improves accuracy and outperforms previous state-of-the-art algorithms. Arnaud Ferré, Robert Bossy, Mouhamadou Ba, Louise Deléger, Thomas Lavergne, Pierre Zweigenbaum, Claire Nedellec |
LREC | 2 |
| 2020 | C-Norm: a neural approach to few-shot entity normalizationabstractBACKGROUND: Entity normalization is an important information extraction task which has gained renewed attention in the last decade, particularly in the biomedical and life science domains. In these domains, and more generally in all specialized domains, this task is still challenging for the latest machine learning-based approaches, which have difficulty handling highly multi-class and few-shot learning problems. To address this issue, we propose C-Norm, a new neural approach which synergistically combines standard and weak supervision, ontological knowledge integration and distributional semantics. RESULTS: Our approach greatly outperforms all methods evaluated on the Bacteria Biotope datasets of BioNLP Open Shared Tasks 2019, without integrating any manually-designed domain-specific rules. CONCLUSIONS: Our results show that relatively shallow neural network methods can perform well in domains that present highly multi-class and few-shot learning problems. Arnaud Ferré, Louise Deléger, Robert Bossy, Pierre Zweigenbaum, Claire Nedellec |
BMC Bioinform. | 3 |
| 2015 | Overview of the gene regulation network and the bacteria biotope tasks in BioNLP'13 shared taskabstractWe present the two Bacteria Track tasks of BioNLP 2013 Shared Task (ST): Gene Regulation Network (GRN) and Bacteria Biotope (BB). These tasks were previously introduced in the 2011 BioNLP-ST Bacteria Track as Bacteria Gene Interaction (BI) and Bacteria Biotope (BB). The Bacteria Track was motivated by a need to develop specific BioNLP tools for fine-grained event extraction in bacteria biology. The 2013 tasks expand on the 2011 version by better addressing the biological knowledge modeling needs. New evaluation metrics were designed for the new goals. Moving beyond a list of gene interactions, the goal of the GRN task is to build a gene regulation network from the extracted gene interactions. BB'13 is dedicated to the extraction of bacteria biotopes, i.e . bacterial environmental information, as was BB'11. BB'13 extends the typology of BB'11 to a large diversity of biotopes, as defined by the OntoBiotope ontology. The detection of entities and events is tackled by distinct subtasks in order to measure the progress achieved by the participant systems since 2011. This paper details the corpus preparations and the evaluation metrics, as well as summarizing and discussing the participant results. Five groups participated in each of the two tasks. The high diversity of the participant methods reflects the dynamism of the BioNLP research community. The highest scores for the GRN and BB'13 tasks are similar to those obtained by the participants in 2011, despite of the increase in difficulty. The high density of events in short text segments (multi-event extraction) was a difficult issue for the participating systems for both tasks. The analysis of the BB'13 results also shows that co-reference resolution and entity boundary detection remain major hindrances. The evaluation results suggest new research directions for the improvement and development of Information Extraction for molecular and environmental biology. The Bacteria Track tasks remain publicly open; the BioNLP-ST website provides an online evaluation service, the reference corpora and the evaluation tools. Robert Bossy, Wiktoria Golik, Zorana Ratkovic, Dialekti Valsamou, Philippe Bessières, Claire Nedellec |
BMC Bioinform. | 1 |
| 2012 | BioNLP Shared Task - The Bacteria TrackabstractBACKGROUND: We present the BioNLP 2011 Shared Task Bacteria Track, the first Information Extraction challenge entirely dedicated to bacteria. It includes three tasks that cover different levels of biological knowledge. The Bacteria Gene Renaming supporting task is aimed at extracting gene renaming and gene name synonymy in PubMed abstracts. The Bacteria Gene Interaction is a gene/protein interaction extraction task from individual sentences. The interactions have been categorized into ten different sub-types, thus giving a detailed account of genetic regulations at the molecular level. Finally, the Bacteria Biotopes task focuses on the localization and environment of bacteria mentioned in textbook articles. We describe the process of creation for the three corpora, including document acquisition and manual annotation, as well as the metrics used to evaluate the participants' submissions. RESULTS: Three teams submitted to the Bacteria Gene Renaming task; the best team achieved an F-score of 87%. For the Bacteria Gene Interaction task, the only participant's score had reached a global F-score of 77%, although the system efficiency varies significantly from one sub-type to another. Three teams submitted to the Bacteria Biotopes task with very different approaches; the best team achieved an F-score of 45%. However, the detailed study of the participating systems efficiency reveals the strengths and weaknesses of each participating system. CONCLUSIONS: The three tasks of the Bacteria Track offer participants a chance to address a wide range of issues in Information Extraction, including entity recognition, semantic typing and coreference resolution. We found common trends in the most efficient systems: the systematic use of syntactic dependencies and machine learning. Nevertheless, the originality of the Bacteria Biotopes task encouraged the use of interesting novel methods and techniques, such as term compositionality, scopes wider than the sentence. Robert Bossy, Julien Jourde, Alain-Pierre Manine, Philippe Veber, Érick Alphonse, Maarten van de Guchte, Philippe Bessières, Claire Nedellec |
BMC Bioinform. | 1 |
| 2010 | Building Large Lexicalized Ontologies from Text: A Use Case in Automatic Indexing of Biotechnology Patents
Claire Nedellec, Wiktoria Golik, Sophie Aubin, Robert Bossy |
EKAW | 4 |
| 2001 | An Edition Control Policy Model for Scientific Collaborative Databases
Robert Bossy |
WISE (1) | 1 |