VLDB 2026 Research / reviewers in the wild / expert
Xavier Tannier
dblp:83/4811
· DBLP profile ↗
53ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health CorpusabstractInternational audience Aidan Mannion, Cécile Macaire, Armand Violle, Stéphane Ohayon, Xavier Tannier, Didier Schwab, Lorraine Goeuriot, François Portet |
LREC | 5 |
| 2026 | Hierarchical supervision in DINOv2 training improves generalizability on white blood cell imagesabstract• Hierarchical supervision in DINOv2 improves the latent space for WBC classification. • Biologically informed hierarchy reduces error severity and aligns features. • The clustering is improved even for non-leukocytic out-of-domain datasets. • Hierarchy supports labels with different degrees of precision. • Hierarchical supervision in DINOv2 training improves generalizability on white blood cell images. The microscopic observation of blood cells is a crucial step in diagnosing pathologies such as leukemia. DINOv2 models have been employed to extract features from blood cell images, but they do not include biological knowledge, nor do they allow multi-granular labels. To enhance the representation of these cells, we propose leveraging a biologically informed hierarchy of white blood cell types. We train a DINOv2-based foundation model with a semi-supervised framework that uses hierarchical supervision. It enables using datasets with varying levels of label precision within a structure that represents the process of cell differentiation. To support multi-level label precision, we modify the original hierarchical loss function, allowing any hierarchy level to serve as a ground truth class. We evaluate our model on three external datasets, including an out-of-domain set of cervical cells. Our approach improves generalization of the model to new datasets, improving by 1 percentage point the balanced accuracy on the two blood cell external datasets, and by 2.5 percentage point the balanced accuracy on the out-of-domain dataset. In addition the proposed strategy better aligns the model’s latent space with biological properties, leading to more acceptable misclassifications Manon Chossegros, Sophia J. Wagner, Christian Matek, Daniel Stockholm, Xavier Tannier, Carsten Marr |
Expert Syst. Appl. | 5 |
| 2024 | A Benchmark Evaluation of Clinical Named Entity Recognition in FrenchabstractBackground: Transformer-based language models have shown strong performance on many Natural Language Processing (NLP) tasks. Masked Language Models (MLMs) attract sustained interest because they can be adapted to different languages and sub-domains through training or fine-tuning on specific corpora while remaining lighter than modern Large Language Models (MLMs). Recently, several MLMs have been released for the biomedical domain in French, and experiments suggest that they outperform standard French counterparts. However, no systematic evaluation comparing all models on the same corpora is available. Objective: This paper presents an evaluation of masked language models for biomedical French on the task of clinical named entity recognition. Material and methods: We evaluate biomedical models CamemBERT-bio and DrBERT and compare them to standard French models CamemBERT, FlauBERT and FrAlBERT as well as multilingual mBERT using three publically available corpora for clinical named entity recognition in French. The evaluation set-up relies on gold-standard corpora as released by the corpus developers. Results: Results suggest that CamemBERT-bio outperforms DrBERT consistently while FlauBERT offers competitive performance and FrAlBERT achieves the lowest carbon footprint. Conclusion: This is the first benchmark evaluation of biomedical masked language models for French clinical entity recognition that compares model performance consistently on nested entity recognition using metrics covering performance and environmental impact. Nesrine Bannour, Christophe Servan, Aurélie Névéol, Xavier Tannier |
LREC/COLING | 4 |
| 2024 | Leveraging Information Redundancy of Real-World Data through Distant SupervisionabstractWe explore the task of event extraction and classification by harnessing the power of distant supervision. We present a novel text labeling method that leverages the redundancy of temporal information in a data lake. This method enables the creation of a large programmatically annotated corpus, allowing the training of transformer models using distant supervision. This aims to reduce expert annotation time, a scarce and expensive resource. Our approach utilizes temporal redundancy between structured sources and text, enabling the design of a replicable framework applicable to diverse real-world databases and use cases. We employ this method to create multiple silver datasets to reconstruct key events in cancer patients’ pathways, using clinical notes from a cohort of 380,000 oncological patients. By employing various noise label management techniques, we validate our end-to-end approach and compare it with a baseline classifier built on expert-annotated data. The implications of our work extend to accelerating downstream applications, such as patient recruitment for clinical trials, treatment effectiveness studies, survival analysis, and epidemiology research. While our study showcases the potential of the method, there remain avenues for further exploration, including advanced noise management techniques, semi-supervised approaches, and a deeper understanding of biases in the generated datasets and models. Ariel Cohen 0005, Alexandrine Lanson, Emmanuelle Kempf, Xavier Tannier |
LREC/COLING | 4 |
| 2024 | Towards Semantic Interoperability Among Heterogeneous Cancer Data Models Using a Layered Modular Hyper-OntologyabstractSemantic interoperability is a growing and challenging subject in the healthcare domain. It aims to ensure a coherent and unambiguous exchange, use, and reuse of health information among different systems and applications. In the context of the EUCAIM (Cancer Image Europe) project, semantic interoperability among various heterogeneous cancer image data models is required to support the communication, integration, and sharing of data in a standardized and structured way. For this purpose, hyper-ontology is developed as a common semantic meta-model that bridges the disparate imaging and clinical knowledge of the various repositories in EUCAIM and supports their integration. EUCAIM’s hyper-ontology is also an application-based ontology targeted for federated semantic querying and image annotation. To facilitate the hyper-ontology building process and ensure the extensibility of the ontology model, an iterative hybrid well-founded approach that divides the ontology structure into layers and modules is established. Mirna El Ghosh, Varvara Kalokyri, Mélanie Sambres, Morgan Vaterkowski, Catherine Duclos, Xavier Tannier, Gianna Tsakou, Manolis Tsiknakis, Christel Daniel-Le Bozec, Ferdinand Dhombres |
FOIS | 6 |
| 2024 | Collaborative and privacy-enhancing workflows on a clinical data warehouse: an example developing natural language processing pipelines to detect medical conditionsabstractOBJECTIVE: To develop and validate a natural language processing (NLP) pipeline that detects 18 conditions in French clinical notes, including 16 comorbidities of the Charlson index, while exploring a collaborative and privacy-enhancing workflow. MATERIALS AND METHODS: The detection pipeline relied both on rule-based and machine learning algorithms, respectively, for named entity recognition and entity qualification, respectively. We used a large language model pre-trained on millions of clinical notes along with annotated clinical notes in the context of 3 cohort studies related to oncology, cardiology, and rheumatology. The overall workflow was conceived to foster collaboration between studies while respecting the privacy constraints of the data warehouse. We estimated the added values of the advanced technologies and of the collaborative setting. RESULTS: The pipeline reached macro-averaged F1-score positive predictive value, sensitivity, and specificity of 95.7 (95%CI 94.5-96.3), 95.4 (95%CI 94.0-96.3), 96.0 (95%CI 94.0-96.7), and 99.2 (95%CI 99.0-99.4), respectively. F1-scores were superior to those observed using alternative technologies or non-collaborative settings. The models were shared through a secured registry. CONCLUSIONS: We demonstrated that a community of investigators working on a common clinical data warehouse could efficiently and securely collaborate to develop, validate and use sensitive artificial intelligence models. In particular, we provided an efficient and robust NLP pipeline that detects conditions mentioned in clinical notes. Thomas Petit-Jean, Christel Gérardin, Emmanuelle Berthelot, Gilles Chatellier, Marie Frank, Xavier Tannier, Emmanuelle Kempf, Romain Bey |
J. Am. Medical Informatics Assoc. | 6 |
| 2022 | How to improve cancer Patients ENrollment within clinical trials from rEal Life databases using the OMOP oncology Extension: the French PENELOPE initiative
Emmanuelle Kempf, Morgan Vaterkowski, Nicolas Griffon, Damien Leprovost, Stéphane Bréant, Patricia Serre, Alexandre Mouchet, Rafael Gozlan, Bastien Rance, Gilles Chatellier, Ali Bellamine, Marie Frank, Martin Hilka, Julien Guerin, Xavier Tannier, Alain Livartowski, Christel Daniel-Le Bozec |
AMIA | 15 |
| 2022 | Multilabel classification of medical concepts for patient clinical profile identification
Christel Gérardin, Perceval Wajsbürt, Pascal Vaillant, Ali Bellamine, Fabrice Carrat, Xavier Tannier |
Artif. Intell. Medicine | 6 |
| 2022 | Privacy-preserving mimic models for clinical named entity recognition in French
Nesrine Bannour, Perceval Wajsbürt, Bastien Rance, Xavier Tannier, Aurélie Névéol |
J. Biomed. Informatics | 4 |
| 2021 | Effect of Depth Order on Iterative Nested Named Entity Recognition Models
Perceval Wajsbürt, Yoann Taillé, Xavier Tannier |
AIME | 3 |
| 2021 | Medical concept normalization in French using multilingual terminologies and contextual embeddings
Perceval Wajsbürt, Arnaud Sarfati, Xavier Tannier |
J. Biomed. Informatics | 3 |
| 2020 | Terminologies augmented recurrent neural network model for clinical named entity recognition
Ivan Lerner, Nicolas Paris, Xavier Tannier |
J. Biomed. Informatics | 3 |
| 2019 | BeLink: Querying Networks of Facts, Statements and BeliefsabstractAn important class of journalistic fact-checking scenarios involves verifying the claims and knowledge of different actors at different moments in time. Claims may be about facts, or about other claims, leading to chains of hearsay. We have recently proposed a data model for (time-anchored) facts, statements and beliefs. It builds upon the W3C's RDF standard for Linked Open Data to describe connections between agents and their statements, and to trace information propagation as agents communicate. We propose to demonstrate BeLink, a prototype capable of storing such interconnected corpora, and answer powerful queries over them relying on SPARQL 1.1. The demo will showcase the exploration of a rich real-data corpus built from Twitter and mainstream media, and interconnected through extraction of statements with their sources, time, and topics. Tien Duc Cao, Ludivine Duroyon, François Goasdoué, Ioana Manolescu, Xavier Tannier |
CIKM | 5 |
| 2019 | Extracting Statistical Mentions from Textual Claims to Provide Trusted Content
Tien Duc Cao, Ioana Manolescu, Xavier Tannier |
NLDB | 3 |
| 2018 | Searching for Truth in a Database of StatisticsabstractThe proliferation of falsehood and misinformation, in particular through the Web, has lead to increasing energy being invested into journalistic fact-checking. Fact-checking journalists typically check the accuracy of a claim against some trusted data source. Statistic databases such as those compiled by state agencies are often used as trusted data sources, as they contain valuable, high-quality information. However, their usability is limited when they are shared in a format such as HTML or spreadsheets: this makes it hard to find the most relevant dataset for checking a specific claim, or to quickly extract from a dataset the best answer to a given query. Tien Duc Cao, Ioana Manolescu, Xavier Tannier |
WebDB | 3 |
| 2018 | Computational fact-checking: a content management perspectiveabstractData journalism designates journalistic work inspired by digital data sources. A particularly popular and active area of data journalism is concerned with fact-checking. The term was born in the journalist community and referred the process of verifying and ensuring the accuracy of published media content; since 2012, however, it has increasingly focused on the analysis of politics, economy, science, and news content shared in any form, but first and foremost on the Web (social and otherwise). These trends have been noticed by computer scientists working in the industry and academia. Thus, a very lively area of digital content management research has taken up these problems and works to propose foundations (models), algorithms, and implement them through concrete tools. Our tutorial: (i) Outlines the current state of affairs in the area of digital (or computational) fact-checking in newsrooms, by journalists, NGO workers, scientists and IT companies; (ii) Shows which areas of digital content management research, in particular those relying on the Web, can be leveraged to help fact-checking, and gives a comprehensive survey of efforts in this area; (iii) Highlights ongoing trends, unsolved problems, and areas where we envision future scientific and practical advances. Sylvie Cazalens, Julien Leblay, Ioana Manolescu, Philippe Lamarre, Xavier Tannier |
Proc. VLDB Endow. | 5 |
| 2017 | Combining Word and Entity Embeddings for Entity Linking
José G. Moreno 0001, Romaric Besançon, Romain Beaumont, Eva D'hondt, Anne-Laure Ligozat, Sophie Rosset, Xavier Tannier, Brigitte Grau |
ESWC (1) | 7 |
| 2016 | Datasets for Aspect-Based Sentiment Analysis in French
Marianna Apidianaki, Xavier Tannier, Cécile Richart |
LREC | 2 |
| 2016 | A Dataset for Open Event Extraction in English
Kiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon |
LREC | 2 |
| 2016 | INEX Tweet Contextualization task: Evaluation, results and lesson learned
Patrice Bellot, Véronique Moriceau, Josiane Mothe, Eric SanJuan, Xavier Tannier |
Inf. Process. Manag. | 5 |
| 2016 | Mixed-instance querying: a lightweight integration architecture for data journalismabstractAs the world's affairs get increasingly more digital, timely production and consumption of news require to efficiently and quickly exploit heterogeneous data sources. Discussions with journalists revealed that content management tools currently at their disposal fall very short of expectations. We demonstrate T atooine , a lightweight data integration prototype, which allows to quickly set up integration queries across (very) heterogeneous data sources, capitalizing on the many data links (joins) available in this application domain. Our demonstration is based on scenarios we study in collaboration with Le Monde, France's major newspaper. Raphaël Bonaque, Tien Duc Cao, Bogdan Cautis, François Goasdoué, Javier Letelier, Ioana Manolescu, Oscar Mendoza, Swen Ribeiro, Xavier Tannier, Michaël Thomazo |
Proc. VLDB Endow. | 9 |
| 2015 | Generative Event Schema Induction with Entity DisambiguationabstractKiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Kiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon |
ACL (1) | 2 |
| 2015 | Automatic Extraction of Time Expressions Accross Domains in French NarrativesabstractThe prevalence of temporal references across all types of natural language utterances makes temporal analysis a key issue in Natural Language Processing.This work adresses three research questions: 1/is temporal expression recognition specific to a particular domain?2/if so, can we characterize domain specificity?and 3/how can subdomain specificity be integrated in a single tool for unified temporal expression extraction?Herein, we assess temporal expression recognition from documents written in French covering three domains.We present a new corpus of clinical narratives annotated for temporal expressions, and also use existing corpora in the newswire and historical domains.We show that temporal expressions can be extracted with high performance across domains (best F-measure 0.96 obtained with a CRF model on clinical narratives).We argue that domain adaptation for the extraction of temporal expressions can be done with limited efforts and should cover pre-processing as well as temporal specific tasks. Mike Donald Tapi-Nzali, Xavier Tannier, Aurélie Névéol |
EMNLP | 2 |
| 2015 | Supervised Machine Learning Techniques to Detect TimeML Events in French and English
Béatrice Arnulphy, Vincent Claveau, Xavier Tannier, Anne Vilnat |
NLDB | 3 |
| 2014 | How to de-identify a large clinical corpus in 10 days
Cyril Grouin, Louise Deléger, Jean-Baptiste Escudié, Gregory Groisy, Anne-Sophie Jannot, Bastien Rance, Xavier Tannier, Aurélie Névéol |
AMIA | 7 |
| 2014 | Ranking Multidocument Event Descriptions for Building Thematic Timelines
Kiem-Hieu Nguyen, Xavier Tannier, Véronique Moriceau |
COLING | 2 |
| 2014 | Evaluating Web-as-corpus Topical Document Retrieval with an Index of the OpenDirectory
Clément de Groc, Xavier Tannier |
LREC | 2 |
| 2014 | Thematic Cohesion: measuring terms discriminatory power toward themes
Clément de Groc, Xavier Tannier, Claude de Loupy |
LREC | 2 |
| 2014 | French Resources for Extraction and Normalization of Temporal Expressions with HeidelTime
Véronique Moriceau, Xavier Tannier |
LREC | 2 |
| 2014 | Extracting News Web Page Creation Time with DCTFinder
Xavier Tannier |
LREC | 1 |
| 2013 | Building Event Threads out of Multiple News ArticlesabstractWe present an approach for building multidocument event threads from a large corpus of newswire articles.An event thread is basically a succession of events belonging to the same story.It helps the reader to contextualize the information contained in a single article, by navigating backward or forward in the thread from this article.A specific effort is also made on the detection of reactions to a particular event.In order to build these event threads, we use a cascade of classifiers and other modules, taking advantage of the redundancy of information in the newswire corpus.We also share interesting comments concerning our manual annotation procedure for building a training and testing set 1 . Xavier Tannier, Véronique Moriceau |
EMNLP | 1 |
| 2013 | Eventual situations for timeline extraction from clinical reportsabstractOBJECTIVE: To identify the temporal relations between clinical events and temporal expressions in clinical reports, as defined in the i2b2/VA 2012 challenge. DESIGN: To detect clinical events, we used rules and Conditional Random Fields. We built Random Forest models to identify event modality and polarity. To identify temporal expressions we built on the HeidelTime system. To detect temporal relations, we systematically studied their breakdown into distinct situations; we designed an oracle method to determine the most prominent situations and the most suitable associated classifiers, and combined their results. RESULTS: We achieved F-measures of 0.8307 for event identification, based on rules, and 0.8385 for temporal expression identification. In the temporal relation task, we identified nine main situations in three groups, experimentally confirming shared intuitions: within-sentence relations, section-related time, and across-sentence relations. Logistic regression and Naïve Bayes performed best on the first and third groups, and decision trees on the second. We reached a 0.6231 global F-measure, improving by 7.5 points our official submission. CONCLUSIONS: Carefully hand-crafted rules obtained good results for the detection of events and temporal expressions, while a combination of classifiers improved temporal link prediction. The characterization of the oracle recall of situations allowed us to point at directions where further work would be most useful for temporal relation detection: within-sentence relations and linking History of Present Illness events to the admission date. We suggest that the systematic situation breakdown proposed in this paper could also help improve other systems addressing this task. Cyril Grouin, Natalia Grabar, Thierry Hamon, Sophie Rosset, Xavier Tannier, Pierre Zweigenbaum |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Finding Salient Dates for Building Thematic Timelines
Rémy Kessler, Xavier Tannier, Caroline Hagège, Véronique Moriceau, André Bittar |
ACL (1) | 2 |
| 2012 | Automatically Generated Noun Lexicons for Event Extraction
Béatrice Arnulphy, Xavier Tannier, Anne Vilnat |
CICLing (2) | 2 |
| 2012 | Event Nominals: Annotation Guidelines and a Manually Annotated Corpus in French
Béatrice Arnulphy, Xavier Tannier, Anne Vilnat |
LREC | 2 |
| 2012 | Temporal Annotation: A Proposal for Guidelines and an Experiment with Inter-annotator Agreement
André Bittar, Caroline Hagège, Véronique Moriceau, Xavier Tannier, Charles Teissèdre |
LREC | 4 |
| 2012 | A Rough Set Formalization of Quantitative Evaluation with Ambiguity
Patrick Paroubek, Xavier Tannier |
LREC | 2 |
| 2012 | WebAnnotator, an Annotation Tool for Web Pages
Xavier Tannier |
LREC | 1 |
| 2012 | Evolution of Event Designation in Media: Preliminary Study
Xavier Tannier, Véronique Moriceau, Béatrice Arnulphy, Ruixin He |
LREC | 1 |
| 2012 | Experiments on Pseudo Relevance Feedback Using Graph Random Walks
Clément de Groc, Xavier Tannier |
SPIRE | 2 |
| 2011 | Evaluating Temporal Graphs Built from Texts via Transitive ReductionabstractTemporal information has been the focus of recent attention in information extraction, leading to some standardization effort, in particular for the task of relating events in a text. This task raises the problem of comparing two annotations of a given text, because relations between events in a story are intrinsically interdependent and cannot be evaluated separately. A proper evaluation measure is also crucial in the context of a machine learning approach to the problem. Finding a common comparison referent at the text level is not obvious, and we argue here in favor of a shift from event-based measures to measures on a unique textual object, a minimal underlying temporal graph, or more formally the transitive reduction of the graph of relations between event boundaries. We support it by an investigation of its properties on synthetic data and on a well-know temporal corpus. Xavier Tannier, Philippe Muller |
J. Artif. Intell. Res. | 1 |
| 2010 | Named and Specific Entity Detection in Varied Data: The Quæro Named Entity Baseline Evaluation
Olivier Galibert, Ludovic Quintard, Sophie Rosset, Pierre Zweigenbaum, Claire Nedellec, Sophie Aubin, Laurent Gillard, Jean-Pierre Raysz, Delphine Pois, Xavier Tannier, Louise Deléger, Dominique Laurent 0003 |
LREC | 10 |
| 2010 | Hybrid Citation Extraction from Patents
Olivier Galibert, Sophie Rosset, Xavier Tannier, Fanny Grandry |
LREC | 3 |
| 2010 | A Corpus for Studying Full Answer Justification
Arnaud Grappy, Brigitte Grau, Olivier Ferret, Cyril Grouin, Véronique Moriceau, Isabelle Robba, Xavier Tannier, Anne Vilnat, Vincent Barbier |
LREC | 7 |
| 2010 | Question Answering on Web Data: The QA Evaluation in Quæro
Ludovic Quintard, Olivier Galibert, Gilles Adda, Brigitte Grau, Dominique Laurent 0003, Véronique Moriceau, Sophie Rosset, Xavier Tannier, Anne Vilnat |
LREC | 8 |
| 2010 | FIDJI: Web Question-Answering at Quaero 2009
Xavier Tannier, Véronique Moriceau |
LREC | 1 |
| 2010 | FIDJI: using syntax for validating answers in multiple documents
Véronique Moriceau, Xavier Tannier |
Inf. Retr. | 2 |
| 2008 | XTM: A Robust Temporal Text Processor
Caroline Hagège, Xavier Tannier |
CICLing | 2 |
| 2008 | Evaluation Metrics for Automatic Temporal Annotation of Texts
Xavier Tannier, Philippe Muller |
LREC | 1 |
| 2005 | Classifying XML tags through "reading contexts"abstractSome tags used in XML documents create arbitrary breaks in the natural flow of the text. This may constitute an impediment to the application of some methods of document engineering. This article introduces the concept of ``reading contexts'', and gives clues to handle it theorically and in practice. This work should notably allow to recognize emphasis tags in a text, to define a new concept of term proximity in structured documents, to improve indexing techniques, and also to open up the way to advanced linguistic analyses of XML corpora. Xavier Tannier, Jean-Jacques Girardot, Mihaela Juganaru-Mathieu |
ACM Symposium on Document Engineering | 1 |
| 2005 | Retrieval Status Values in Information Retrieval Evaluation
Amélie Imafouo, Xavier Tannier |
SPIRE | 2 |
| 2005 | XML Retrieval with a Natural Language Interface
Xavier Tannier, Shlomo Geva |
SPIRE | 1 |
| 2004 | Annotating and measuring temporal relations in texts
Philippe Muller, Xavier Tannier |
COLING | 2 |