Xavier Tannier

dblp:83/4811 · DBLP profile ↗
← Back
53ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Is Biomedical Specialization Still Worth It? Insights from Domain-Adaptive Language Modelling with a New French Health Corpus
abstract
International audience
Aidan Mannion, Cécile Macaire, Armand Violle, Stéphane Ohayon, Xavier Tannier, Didier Schwab, Lorraine Goeuriot, François Portet
LREC5
2026 Hierarchical supervision in DINOv2 training improves generalizability on white blood cell images
abstract
• Hierarchical supervision in DINOv2 improves the latent space for WBC classification. • Biologically informed hierarchy reduces error severity and aligns features. • The clustering is improved even for non-leukocytic out-of-domain datasets. • Hierarchy supports labels with different degrees of precision. • Hierarchical supervision in DINOv2 training improves generalizability on white blood cell images. The microscopic observation of blood cells is a crucial step in diagnosing pathologies such as leukemia. DINOv2 models have been employed to extract features from blood cell images, but they do not include biological knowledge, nor do they allow multi-granular labels. To enhance the representation of these cells, we propose leveraging a biologically informed hierarchy of white blood cell types. We train a DINOv2-based foundation model with a semi-supervised framework that uses hierarchical supervision. It enables using datasets with varying levels of label precision within a structure that represents the process of cell differentiation. To support multi-level label precision, we modify the original hierarchical loss function, allowing any hierarchy level to serve as a ground truth class. We evaluate our model on three external datasets, including an out-of-domain set of cervical cells. Our approach improves generalization of the model to new datasets, improving by 1 percentage point the balanced accuracy on the two blood cell external datasets, and by 2.5 percentage point the balanced accuracy on the out-of-domain dataset. In addition the proposed strategy better aligns the model’s latent space with biological properties, leading to more acceptable misclassifications
Manon Chossegros, Sophia J. Wagner, Christian Matek, Daniel Stockholm, Xavier Tannier, Carsten Marr
Expert Syst. Appl.5
2024 A Benchmark Evaluation of Clinical Named Entity Recognition in French
abstract
Background: Transformer-based language models have shown strong performance on many Natural Language Processing (NLP) tasks. Masked Language Models (MLMs) attract sustained interest because they can be adapted to different languages and sub-domains through training or fine-tuning on specific corpora while remaining lighter than modern Large Language Models (MLMs). Recently, several MLMs have been released for the biomedical domain in French, and experiments suggest that they outperform standard French counterparts. However, no systematic evaluation comparing all models on the same corpora is available. Objective: This paper presents an evaluation of masked language models for biomedical French on the task of clinical named entity recognition. Material and methods: We evaluate biomedical models CamemBERT-bio and DrBERT and compare them to standard French models CamemBERT, FlauBERT and FrAlBERT as well as multilingual mBERT using three publically available corpora for clinical named entity recognition in French. The evaluation set-up relies on gold-standard corpora as released by the corpus developers. Results: Results suggest that CamemBERT-bio outperforms DrBERT consistently while FlauBERT offers competitive performance and FrAlBERT achieves the lowest carbon footprint. Conclusion: This is the first benchmark evaluation of biomedical masked language models for French clinical entity recognition that compares model performance consistently on nested entity recognition using metrics covering performance and environmental impact.
Nesrine Bannour, Christophe Servan, Aurélie Névéol, Xavier Tannier
LREC/COLING4
2024 Leveraging Information Redundancy of Real-World Data through Distant Supervision
abstract
We explore the task of event extraction and classification by harnessing the power of distant supervision. We present a novel text labeling method that leverages the redundancy of temporal information in a data lake. This method enables the creation of a large programmatically annotated corpus, allowing the training of transformer models using distant supervision. This aims to reduce expert annotation time, a scarce and expensive resource. Our approach utilizes temporal redundancy between structured sources and text, enabling the design of a replicable framework applicable to diverse real-world databases and use cases. We employ this method to create multiple silver datasets to reconstruct key events in cancer patients’ pathways, using clinical notes from a cohort of 380,000 oncological patients. By employing various noise label management techniques, we validate our end-to-end approach and compare it with a baseline classifier built on expert-annotated data. The implications of our work extend to accelerating downstream applications, such as patient recruitment for clinical trials, treatment effectiveness studies, survival analysis, and epidemiology research. While our study showcases the potential of the method, there remain avenues for further exploration, including advanced noise management techniques, semi-supervised approaches, and a deeper understanding of biases in the generated datasets and models.
Ariel Cohen 0005, Alexandrine Lanson, Emmanuelle Kempf, Xavier Tannier
LREC/COLING4
2024 Towards Semantic Interoperability Among Heterogeneous Cancer Data Models Using a Layered Modular Hyper-Ontology
abstract
Semantic interoperability is a growing and challenging subject in the healthcare domain. It aims to ensure a coherent and unambiguous exchange, use, and reuse of health information among different systems and applications. In the context of the EUCAIM (Cancer Image Europe) project, semantic interoperability among various heterogeneous cancer image data models is required to support the communication, integration, and sharing of data in a standardized and structured way. For this purpose, hyper-ontology is developed as a common semantic meta-model that bridges the disparate imaging and clinical knowledge of the various repositories in EUCAIM and supports their integration. EUCAIM’s hyper-ontology is also an application-based ontology targeted for federated semantic querying and image annotation. To facilitate the hyper-ontology building process and ensure the extensibility of the ontology model, an iterative hybrid well-founded approach that divides the ontology structure into layers and modules is established.
Mirna El Ghosh, Varvara Kalokyri, Mélanie Sambres, Morgan Vaterkowski, Catherine Duclos, Xavier Tannier, Gianna Tsakou, Manolis Tsiknakis, Christel Daniel-Le Bozec, Ferdinand Dhombres
FOIS6
2024 Collaborative and privacy-enhancing workflows on a clinical data warehouse: an example developing natural language processing pipelines to detect medical conditions
abstract
OBJECTIVE: To develop and validate a natural language processing (NLP) pipeline that detects 18 conditions in French clinical notes, including 16 comorbidities of the Charlson index, while exploring a collaborative and privacy-enhancing workflow. MATERIALS AND METHODS: The detection pipeline relied both on rule-based and machine learning algorithms, respectively, for named entity recognition and entity qualification, respectively. We used a large language model pre-trained on millions of clinical notes along with annotated clinical notes in the context of 3 cohort studies related to oncology, cardiology, and rheumatology. The overall workflow was conceived to foster collaboration between studies while respecting the privacy constraints of the data warehouse. We estimated the added values of the advanced technologies and of the collaborative setting. RESULTS: The pipeline reached macro-averaged F1-score positive predictive value, sensitivity, and specificity of 95.7 (95%CI 94.5-96.3), 95.4 (95%CI 94.0-96.3), 96.0 (95%CI 94.0-96.7), and 99.2 (95%CI 99.0-99.4), respectively. F1-scores were superior to those observed using alternative technologies or non-collaborative settings. The models were shared through a secured registry. CONCLUSIONS: We demonstrated that a community of investigators working on a common clinical data warehouse could efficiently and securely collaborate to develop, validate and use sensitive artificial intelligence models. In particular, we provided an efficient and robust NLP pipeline that detects conditions mentioned in clinical notes.
Thomas Petit-Jean, Christel Gérardin, Emmanuelle Berthelot, Gilles Chatellier, Marie Frank, Xavier Tannier, Emmanuelle Kempf, Romain Bey
J. Am. Medical Informatics Assoc.6
2022 How to improve cancer Patients ENrollment within clinical trials from rEal Life databases using the OMOP oncology Extension: the French PENELOPE initiative
Emmanuelle Kempf, Morgan Vaterkowski, Nicolas Griffon, Damien Leprovost, Stéphane Bréant, Patricia Serre, Alexandre Mouchet, Rafael Gozlan, Bastien Rance, Gilles Chatellier, Ali Bellamine, Marie Frank, Martin Hilka, Julien Guerin, Xavier Tannier, Alain Livartowski, Christel Daniel-Le Bozec
AMIA15
2022 Multilabel classification of medical concepts for patient clinical profile identification
Christel Gérardin, Perceval Wajsbürt, Pascal Vaillant, Ali Bellamine, Fabrice Carrat, Xavier Tannier
Artif. Intell. Medicine6
2022 Privacy-preserving mimic models for clinical named entity recognition in French
Nesrine Bannour, Perceval Wajsbürt, Bastien Rance, Xavier Tannier, Aurélie Névéol
J. Biomed. Informatics4
2021 Effect of Depth Order on Iterative Nested Named Entity Recognition Models
Perceval Wajsbürt, Yoann Taillé, Xavier Tannier
AIME3
2021 Medical concept normalization in French using multilingual terminologies and contextual embeddings
Perceval Wajsbürt, Arnaud Sarfati, Xavier Tannier
J. Biomed. Informatics3
2020 Terminologies augmented recurrent neural network model for clinical named entity recognition
Ivan Lerner, Nicolas Paris, Xavier Tannier
J. Biomed. Informatics3
2019 BeLink: Querying Networks of Facts, Statements and Beliefs
abstract
An important class of journalistic fact-checking scenarios involves verifying the claims and knowledge of different actors at different moments in time. Claims may be about facts, or about other claims, leading to chains of hearsay. We have recently proposed a data model for (time-anchored) facts, statements and beliefs. It builds upon the W3C's RDF standard for Linked Open Data to describe connections between agents and their statements, and to trace information propagation as agents communicate. We propose to demonstrate BeLink, a prototype capable of storing such interconnected corpora, and answer powerful queries over them relying on SPARQL 1.1. The demo will showcase the exploration of a rich real-data corpus built from Twitter and mainstream media, and interconnected through extraction of statements with their sources, time, and topics.
Tien Duc Cao, Ludivine Duroyon, François Goasdoué, Ioana Manolescu, Xavier Tannier
CIKM5
2019 Extracting Statistical Mentions from Textual Claims to Provide Trusted Content
Tien Duc Cao, Ioana Manolescu, Xavier Tannier
NLDB3
2018 Searching for Truth in a Database of Statistics
abstract
The proliferation of falsehood and misinformation, in particular through the Web, has lead to increasing energy being invested into journalistic fact-checking. Fact-checking journalists typically check the accuracy of a claim against some trusted data source. Statistic databases such as those compiled by state agencies are often used as trusted data sources, as they contain valuable, high-quality information. However, their usability is limited when they are shared in a format such as HTML or spreadsheets: this makes it hard to find the most relevant dataset for checking a specific claim, or to quickly extract from a dataset the best answer to a given query.
Tien Duc Cao, Ioana Manolescu, Xavier Tannier
WebDB3
2018 Computational fact-checking: a content management perspective
abstract
Data journalism designates journalistic work inspired by digital data sources. A particularly popular and active area of data journalism is concerned with fact-checking. The term was born in the journalist community and referred the process of verifying and ensuring the accuracy of published media content; since 2012, however, it has increasingly focused on the analysis of politics, economy, science, and news content shared in any form, but first and foremost on the Web (social and otherwise). These trends have been noticed by computer scientists working in the industry and academia. Thus, a very lively area of digital content management research has taken up these problems and works to propose foundations (models), algorithms, and implement them through concrete tools. Our tutorial: (i) Outlines the current state of affairs in the area of digital (or computational) fact-checking in newsrooms, by journalists, NGO workers, scientists and IT companies; (ii) Shows which areas of digital content management research, in particular those relying on the Web, can be leveraged to help fact-checking, and gives a comprehensive survey of efforts in this area; (iii) Highlights ongoing trends, unsolved problems, and areas where we envision future scientific and practical advances.
Sylvie Cazalens, Julien Leblay, Ioana Manolescu, Philippe Lamarre, Xavier Tannier
Proc. VLDB Endow.5
2017 Combining Word and Entity Embeddings for Entity Linking
José G. Moreno 0001, Romaric Besançon, Romain Beaumont, Eva D'hondt, Anne-Laure Ligozat, Sophie Rosset, Xavier Tannier, Brigitte Grau
ESWC (1)7
2016 Datasets for Aspect-Based Sentiment Analysis in French
Marianna Apidianaki, Xavier Tannier, Cécile Richart
LREC2
2016 A Dataset for Open Event Extraction in English
Kiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon
LREC2
2016 INEX Tweet Contextualization task: Evaluation, results and lesson learned
Patrice Bellot, Véronique Moriceau, Josiane Mothe, Eric SanJuan, Xavier Tannier
Inf. Process. Manag.5
2016 Mixed-instance querying: a lightweight integration architecture for data journalism
abstract
As the world's affairs get increasingly more digital, timely production and consumption of news require to efficiently and quickly exploit heterogeneous data sources. Discussions with journalists revealed that content management tools currently at their disposal fall very short of expectations. We demonstrate T atooine , a lightweight data integration prototype, which allows to quickly set up integration queries across (very) heterogeneous data sources, capitalizing on the many data links (joins) available in this application domain. Our demonstration is based on scenarios we study in collaboration with Le Monde, France's major newspaper.
Raphaël Bonaque, Tien Duc Cao, Bogdan Cautis, François Goasdoué, Javier Letelier, Ioana Manolescu, Oscar Mendoza, Swen Ribeiro, Xavier Tannier, Michaël Thomazo
Proc. VLDB Endow.9
2015 Generative Event Schema Induction with Entity Disambiguation
abstract
Kiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Kiem-Hieu Nguyen, Xavier Tannier, Olivier Ferret, Romaric Besançon
ACL (1)2
2015 Automatic Extraction of Time Expressions Accross Domains in French Narratives
abstract
The prevalence of temporal references across all types of natural language utterances makes temporal analysis a key issue in Natural Language Processing.This work adresses three research questions: 1/is temporal expression recognition specific to a particular domain?2/if so, can we characterize domain specificity?and 3/how can subdomain specificity be integrated in a single tool for unified temporal expression extraction?Herein, we assess temporal expression recognition from documents written in French covering three domains.We present a new corpus of clinical narratives annotated for temporal expressions, and also use existing corpora in the newswire and historical domains.We show that temporal expressions can be extracted with high performance across domains (best F-measure 0.96 obtained with a CRF model on clinical narratives).We argue that domain adaptation for the extraction of temporal expressions can be done with limited efforts and should cover pre-processing as well as temporal specific tasks.
Mike Donald Tapi-Nzali, Xavier Tannier, Aurélie Névéol
EMNLP2
2015 Supervised Machine Learning Techniques to Detect TimeML Events in French and English
Béatrice Arnulphy, Vincent Claveau, Xavier Tannier, Anne Vilnat
NLDB3
2014 How to de-identify a large clinical corpus in 10 days
Cyril Grouin, Louise Deléger, Jean-Baptiste Escudié, Gregory Groisy, Anne-Sophie Jannot, Bastien Rance, Xavier Tannier, Aurélie Névéol
AMIA7
2014 Ranking Multidocument Event Descriptions for Building Thematic Timelines
Kiem-Hieu Nguyen, Xavier Tannier, Véronique Moriceau
COLING2
2014 Evaluating Web-as-corpus Topical Document Retrieval with an Index of the OpenDirectory
Clément de Groc, Xavier Tannier
LREC2
2014 Thematic Cohesion: measuring terms discriminatory power toward themes
Clément de Groc, Xavier Tannier, Claude de Loupy
LREC2
2014 French Resources for Extraction and Normalization of Temporal Expressions with HeidelTime
Véronique Moriceau, Xavier Tannier
LREC2
2014 Extracting News Web Page Creation Time with DCTFinder
Xavier Tannier
LREC1
2013 Building Event Threads out of Multiple News Articles
abstract
We present an approach for building multidocument event threads from a large corpus of newswire articles.An event thread is basically a succession of events belonging to the same story.It helps the reader to contextualize the information contained in a single article, by navigating backward or forward in the thread from this article.A specific effort is also made on the detection of reactions to a particular event.In order to build these event threads, we use a cascade of classifiers and other modules, taking advantage of the redundancy of information in the newswire corpus.We also share interesting comments concerning our manual annotation procedure for building a training and testing set 1 .
Xavier Tannier, Véronique Moriceau
EMNLP1
2013 Eventual situations for timeline extraction from clinical reports
abstract
OBJECTIVE: To identify the temporal relations between clinical events and temporal expressions in clinical reports, as defined in the i2b2/VA 2012 challenge. DESIGN: To detect clinical events, we used rules and Conditional Random Fields. We built Random Forest models to identify event modality and polarity. To identify temporal expressions we built on the HeidelTime system. To detect temporal relations, we systematically studied their breakdown into distinct situations; we designed an oracle method to determine the most prominent situations and the most suitable associated classifiers, and combined their results. RESULTS: We achieved F-measures of 0.8307 for event identification, based on rules, and 0.8385 for temporal expression identification. In the temporal relation task, we identified nine main situations in three groups, experimentally confirming shared intuitions: within-sentence relations, section-related time, and across-sentence relations. Logistic regression and Naïve Bayes performed best on the first and third groups, and decision trees on the second. We reached a 0.6231 global F-measure, improving by 7.5 points our official submission. CONCLUSIONS: Carefully hand-crafted rules obtained good results for the detection of events and temporal expressions, while a combination of classifiers improved temporal link prediction. The characterization of the oracle recall of situations allowed us to point at directions where further work would be most useful for temporal relation detection: within-sentence relations and linking History of Present Illness events to the admission date. We suggest that the systematic situation breakdown proposed in this paper could also help improve other systems addressing this task.
Cyril Grouin, Natalia Grabar, Thierry Hamon, Sophie Rosset, Xavier Tannier, Pierre Zweigenbaum
J. Am. Medical Informatics Assoc.5
2012 Finding Salient Dates for Building Thematic Timelines
Rémy Kessler, Xavier Tannier, Caroline Hagège, Véronique Moriceau, André Bittar
ACL (1)2
2012 Automatically Generated Noun Lexicons for Event Extraction
Béatrice Arnulphy, Xavier Tannier, Anne Vilnat
CICLing (2)2
2012 Event Nominals: Annotation Guidelines and a Manually Annotated Corpus in French
Béatrice Arnulphy, Xavier Tannier, Anne Vilnat
LREC2
2012 Temporal Annotation: A Proposal for Guidelines and an Experiment with Inter-annotator Agreement
André Bittar, Caroline Hagège, Véronique Moriceau, Xavier Tannier, Charles Teissèdre
LREC4
2012 A Rough Set Formalization of Quantitative Evaluation with Ambiguity
Patrick Paroubek, Xavier Tannier
LREC2
2012 WebAnnotator, an Annotation Tool for Web Pages
Xavier Tannier
LREC1
2012 Evolution of Event Designation in Media: Preliminary Study
Xavier Tannier, Véronique Moriceau, Béatrice Arnulphy, Ruixin He
LREC1
2012 Experiments on Pseudo Relevance Feedback Using Graph Random Walks
Clément de Groc, Xavier Tannier
SPIRE2
2011 Evaluating Temporal Graphs Built from Texts via Transitive Reduction
abstract
Temporal information has been the focus of recent attention in information extraction, leading to some standardization effort, in particular for the task of relating events in a text. This task raises the problem of comparing two annotations of a given text, because relations between events in a story are intrinsically interdependent and cannot be evaluated separately. A proper evaluation measure is also crucial in the context of a machine learning approach to the problem. Finding a common comparison referent at the text level is not obvious, and we argue here in favor of a shift from event-based measures to measures on a unique textual object, a minimal underlying temporal graph, or more formally the transitive reduction of the graph of relations between event boundaries. We support it by an investigation of its properties on synthetic data and on a well-know temporal corpus.
Xavier Tannier, Philippe Muller
J. Artif. Intell. Res.1
2010 Named and Specific Entity Detection in Varied Data: The Quæro Named Entity Baseline Evaluation
Olivier Galibert, Ludovic Quintard, Sophie Rosset, Pierre Zweigenbaum, Claire Nedellec, Sophie Aubin, Laurent Gillard, Jean-Pierre Raysz, Delphine Pois, Xavier Tannier, Louise Deléger, Dominique Laurent 0003
LREC10
2010 Hybrid Citation Extraction from Patents
Olivier Galibert, Sophie Rosset, Xavier Tannier, Fanny Grandry
LREC3
2010 A Corpus for Studying Full Answer Justification
Arnaud Grappy, Brigitte Grau, Olivier Ferret, Cyril Grouin, Véronique Moriceau, Isabelle Robba, Xavier Tannier, Anne Vilnat, Vincent Barbier
LREC7
2010 Question Answering on Web Data: The QA Evaluation in Quæro
Ludovic Quintard, Olivier Galibert, Gilles Adda, Brigitte Grau, Dominique Laurent 0003, Véronique Moriceau, Sophie Rosset, Xavier Tannier, Anne Vilnat
LREC8
2010 FIDJI: Web Question-Answering at Quaero 2009
Xavier Tannier, Véronique Moriceau
LREC1
2010 FIDJI: using syntax for validating answers in multiple documents
Véronique Moriceau, Xavier Tannier
Inf. Retr.2
2008 XTM: A Robust Temporal Text Processor
Caroline Hagège, Xavier Tannier
CICLing2
2008 Evaluation Metrics for Automatic Temporal Annotation of Texts
Xavier Tannier, Philippe Muller
LREC1
2005 Classifying XML tags through "reading contexts"
abstract
Some tags used in XML documents create arbitrary breaks in the natural flow of the text. This may constitute an impediment to the application of some methods of document engineering. This article introduces the concept of ``reading contexts'', and gives clues to handle it theorically and in practice. This work should notably allow to recognize emphasis tags in a text, to define a new concept of term proximity in structured documents, to improve indexing techniques, and also to open up the way to advanced linguistic analyses of XML corpora.
Xavier Tannier, Jean-Jacques Girardot, Mihaela Juganaru-Mathieu
ACM Symposium on Document Engineering1
2005 Retrieval Status Values in Information Retrieval Evaluation
Amélie Imafouo, Xavier Tannier
SPIRE2
2005 XML Retrieval with a Natural Language Interface
Xavier Tannier, Shlomo Geva
SPIRE1
2004 Annotating and measuring temporal relations in texts
Philippe Muller, Xavier Tannier
COLING2