VLDB 2026 Research / reviewers in the wild / expert
Pierre Zweigenbaum
dblp:61/1252
· DBLP profile ↗
83ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0001-8410-4808ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Assessing the Difficulty of Inference Types in Natural Language Inference for Clinical TrialsabstractInternational audience Mathilde Aguiar, Pierre Zweigenbaum, Nona Naderi |
LREC | 2 |
| 2026 | Is Clinical Text Enough? A Multimodal Study on Mortality Prediction in Heart Failure PatientsabstractInternational audience Oumaima El Khettari, Virgile Barthet, Guillaume Hocquet, Joconde Weller, Emmanuel Morin, Pierre Zweigenbaum |
LREC | 6 |
| 2024 | Enriching a Time-Domain Astrophysics Corpus with Named Entity, Coreference and Astrophysical Relationship AnnotationsabstractInterest in Astrophysical Natural Language Processing (NLP) has increased recently, fueled by the development of specialized language models for information extraction. However, the scarcity of annotated resources for this domain is still a significant challenge. Most existing corpora are limited to Named Entity Recognition (NER) tasks, leaving a gap in resource diversity. To address this gap and facilitate a broader spectrum of NLP research in astrophysics, we introduce astroECR, an extension of our previously built Time-Domain Astrophysics Corpus (TDAC). Our contributions involve expanding it to cover named entities, coreferences, annotations related to astrophysical relationships, and normalizing celestial object names. We showcase practical utility through baseline models for four NLP tasks and provide the research community access to our corpus, code, and models. Atilla Kaan Alkan, Félix Grèzes, Cyril Grouin, Fabian Schüssler, Pierre Zweigenbaum |
LREC/COLING | 5 |
| 2024 | A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across LanguagesabstractUser-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages. Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas 0001, Tomohiro Nishiyama, Sebastian Möller 0001, Eiji Aramaki, Yuji Matsumoto 0001, Roland Roller, Pierre Zweigenbaum |
LREC/COLING | 14 |
| 2024 | Exploiting Graph Embeddings from Knowledge Bases for Neural Biomedical Relation Extraction
Anfu Tang, Louise Deléger, Robert Bossy, Pierre Zweigenbaum, Claire Nedellec |
NLDB (1) | 4 |
| 2022 | Building Comparable Corpora for Assessing Multi-Word Term AlignmentabstractRecent work has demonstrated the importance of dealing with Multi-Word Terms (MWTs) in several Natural Language Processing applications. In particular, MWTs pose serious challenges for alignment and machine translation systems because of their syntactic and semantic properties. Thus, developing algorithms that handle MWTs is becoming essential for many NLP tasks. However, the availability of bilingual and more generally multi-lingual resources is limited, especially for low-resourced languages and in specialized domains. In this paper, we propose an approach for building comparable corpora and bilingual term dictionaries that help evaluate bilingual term alignment in comparable corpora. To that aim, we exploit parallel corpora to perform automatic bilingual MWT extraction and comparable corpus construction. Parallel information helps to align bilingual MWTs and makes it easier to build comparable specialized sub-corpora. Experimental validation on an existing dataset and on manually annotated data shows the interest of the proposed methodology. Omar Adjali, Emmanuel Morin, Pierre Zweigenbaum |
LREC | 3 |
| 2022 | Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical DomainabstractBERT models used in specialized domains all seem to be the result of a simple strategy: initializing with the original BERT and then resuming pre-training on a specialized corpus. This method yields rather good performance (e.g. BioBERT (Lee et al., 2020), SciBERT (Beltagy et al., 2019), BlueBERT (Peng et al., 2019)). However, it seems reasonable to think that training directly on a specialized corpus, using a specialized vocabulary, could result in more tailored embeddings and thus help performance. To test this hypothesis, we train BERT models from scratch using many configurations involving general and medical corpora. Based on evaluations using four different tasks, we find that the initial corpus only has a weak influence on the performance of BERT models when these are further pre-trained on a medical corpus. Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Pierre Zweigenbaum |
LREC | 4 |
| 2022 | Cross-lingual Approaches for the Detection of Adverse Drug Reactions in German from a Patient's PerspectiveabstractIn this work, we present the first corpus for German Adverse Drug Reaction (ADR) detection in patient-generated content. The data consists of 4,169 binary annotated documents from a German patient forum, where users talk about health issues and get advice from medical doctors. As is common in social media data in this domain, the class labels of the corpus are very imbalanced. This and a high topic imbalance make it a very challenging dataset, since often, the same symptom can have several causes and is not always related to a medication intake. We aim to encourage further multi-lingual efforts in the domain of ADR detection and provide preliminary experiments for binary classification using different methods of zero- and few-shot learning based on a multi-lingual model. When fine-tuning XLM-RoBERTa first on English patient forum data and then on the new German data, we achieve an F1-score of 37.52 for the positive class. We make the dataset and models publicly available for the community. Lisa Raithel, Philippe Thomas 0002, Roland Roller, Oliver Sapina, Sebastian Möller 0001, Pierre Zweigenbaum |
LREC | 6 |
| 2022 | Decorate the Examples: A Simple Method of Prompt Design for Biomedical Relation ExtractionabstractRelation extraction is a core problem for natural language processing in the biomedical domain. Recent research on relation extraction showed that prompt-based learning improves the performance on both fine-tuning on full training set and few-shot training. However, less effort has been made on domain-specific tasks where good prompt design can be even harder. In this paper, we investigate prompting for biomedical relation extraction, with experiments on the ChemProt dataset. We present a simple yet effective method to systematically generate comprehensive prompts that reformulate the relation extraction task as a cloze-test task under a simple prompt formulation. In particular, we experiment with different ranking scores for prompt selection. With BioMed-RoBERTa-base, our results show that prompting-based fine-tuning obtains gains by 14.21 F1 over its regular fine-tuning baseline, and 1.14 F1 over SciFive-Large, the current state-of-the-art on ChemProt. Besides, we find prompt-based learning requires fewer training examples to make reasonable predictions. The results demonstrate the potential of our methods in such a domain-specific relation extraction task. Hui-Syuan Yeh, Thomas Lavergne, Pierre Zweigenbaum |
LREC | 3 |
| 2021 | Reproducibility in biomedical natural language processing: A FAIR approach to what we need to know
Kevin Cohen 0001, Anna Ripple, Asma Ben Abacha, Olivier Bodenreider, Orin Hargraves, Karin Verspoor, Pierre Zweigenbaum, Dina Demner-Fushman |
AMIA | 7 |
| 2020 | CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From CharactersabstractDue to the compelling improvements brought by BERT, many recent representation models adopted the Transformer architecture as their main building block, consequently inheriting the wordpiece tokenization system despite it not being intrinsically linked to the notion of Transformers.While this system is thought to achieve a good balance between the flexibility of characters and the efficiency of full words, using predefined wordpiece vocabularies from the general domain is not always suitable, especially when building models for specialized domains (e.g., the medical domain).Moreover, adopting a wordpiece tokenization shifts the focus from the word level to the subword level, making the models conceptually more complex and arguably less convenient in practice.For these reasons, we propose CharacterBERT, a new variant of BERT that drops the wordpiece system altogether and uses a Character-CNN module instead to represent entire words by consulting their characters.We show that this new model improves the performance of BERT on a variety of medical domain tasks while at the same time producing robust, word-level, and open-vocabulary representations. Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, Hiroshi Noji, Pierre Zweigenbaum, Jun'ichi Tsujii |
COLING | 5 |
| 2020 | The Multilingual Anonymisation Toolkit for Public Administrations (MAPA) ProjectabstractWe describe the MAPA project, funded under the Connecting Europe Facility programme, whose goal is the development of an open-source de-identification toolkit for all official European Union languages. It will be developed since January 2020 until December 2021. Eriks Ajausks, Victoria Arranz, Laurent Bié, Aleix Cerdà-i-Cucó, Khalid Choukri, Montse Cuadros, Hans Degroote, Amando Estela, Thierry Etchegoyhen, Mercedes García-Martínez, Aitor García-Pablos, Manuel Herranz, Alejandro Kohan, Maite Melero, Mike Rosner, Roberts Rozis, Patrick Paroubek, Arturs Vasilevskis, Pierre Zweigenbaum |
EAMT | 19 |
| 2020 | Automatic Removal of Identifying Information in Official EU Languages for Public Administrations: The MAPA ProjectabstractThe European MAPA (Multilingual Anonymisation for Public Administrations) project aims at developing an open-source solution for automatic de-identification of medical and legal documents. We introduce here the context, partners and aims of the project, and report on preliminary results. Lucie Gianola, Eriks Ajausks, Victoria Arranz, Chomicha Bendahman, Laurent Bié, Claudia Borg, Aleix Cerdà-i-Cucó, Khalid Choukri, Montse Cuadros, Ona de Gibert Bonet, Hans Degroote, Elena Edelman, Thierry Etchegoyhen, Ángela Franco Torres, Mercedes García Hernandez, Aitor García-Pablos, Albert Gatt, Cyril Grouin, Manuel Herranz, Alejandro Kohan, Thomas Lavergne, Maite Melero, Patrick Paroubek, Mickaël Rigault, Mike Rosner, Roberts Rozis, Lonneke van der Plas, Rinalds Viksna, Pierre Zweigenbaum |
JURIX | 29 |
| 2020 | Handling Entity Normalization with no Annotated Corpus: Weakly Supervised Methods Based on Distributional Representation and Ontological InformationabstractEntity normalization (or entity linking) is an important subtask of information extraction that links entity mentions in text to categories or concepts in a reference vocabulary. Machine learning based normalization methods have good adaptability as long as they have enough training data per reference with a sufficient quality. Distributional representations are commonly used because of their capacity to handle different expressions with similar meanings. However, in specific technical and scientific domains, the small amount of training data and the relatively small size of specialized corpora remain major challenges. Recently, the machine learning-based CONTES method has addressed these challenges for reference vocabularies that are ontologies, as is often the case in life sciences and biomedical domains. And yet, its performance is dependent on manually annotated corpus. Furthermore, like other machine learning based methods, parametrization remains tricky. We propose a new approach to address the scarcity of training data that extends the CONTES method by corpus selection, pre-processing and weak supervision strategies, which can yield high-performance results without any manually annotated examples. We also study which hyperparameters are most influential, with sometimes different patterns compared to previous work. The results show that our approach significantly improves accuracy and outperforms previous state-of-the-art algorithms. Arnaud Ferré, Robert Bossy, Mouhamadou Ba, Louise Deléger, Thomas Lavergne, Pierre Zweigenbaum, Claire Nedellec |
LREC | 6 |
| 2020 | C-Norm: a neural approach to few-shot entity normalizationabstractBACKGROUND: Entity normalization is an important information extraction task which has gained renewed attention in the last decade, particularly in the biomedical and life science domains. In these domains, and more generally in all specialized domains, this task is still challenging for the latest machine learning-based approaches, which have difficulty handling highly multi-class and few-shot learning problems. To address this issue, we propose C-Norm, a new neural approach which synergistically combines standard and weak supervision, ontological knowledge integration and distributional semantics. RESULTS: Our approach greatly outperforms all methods evaluated on the Bacteria Biotope datasets of BioNLP Open Shared Tasks 2019, without integrating any manually-designed domain-specific rules. CONCLUSIONS: Our results show that relatively shallow neural network methods can perform well in domains that present highly multi-class and few-shot learning problems. Arnaud Ferré, Louise Deléger, Robert Bossy, Pierre Zweigenbaum, Claire Nedellec |
BMC Bioinform. | 4 |
| 2020 | Designing a virtual patient dialogue system based on terminology-rich resources: Challenges and evaluationabstractAbstract Virtual patient software allows health professionals to practise their skills by interacting with tools simulating clinical scenarios. A natural language dialogue system can provide natural interaction for medical history-taking. However, the large number of concepts and terms in the medical domain makes the creation of such a system a demanding task. We designed a dialogue system that stands out from current research by its ability to handle a wide variety of medical specialties and clinical cases. To address the task, we designed a patient record model, a knowledge model for the task and a termino-ontological model that hosts structured thesauri with linguistic, terminological and ontological knowledge. We used a frame- and rule-based approach and terminology-rich resources to handle the medical dialogue. This work focuses on the termino-ontological model, the challenges involved and how the system manages resources for the French language. We adopted a comprehensive approach to collect terms and ontological knowledge, and dictionaries of affixes, synonyms and derivational variants. Resources include domain lists containing over 161,000 terms, and dictionaries with over 959,000 word/concept entries. We assessed our approach by having 71 participants (39 medical doctors and 32 non-medical evaluators) interact with the system and use 35 cases from 18 specialities. We conducted a quantitative evaluation of all components by analysing interaction logs (11,834 turns). Natural language understanding achieved an F-measure of 95.8%. Dialogue management provided on average 74.3 (±9.5)% of correct answers. We performed a qualitative evaluation by collecting 171 five-point Likert scale questionnaires. All evaluated aspects obtained mean scores above the Likert mid-scale point. We analysed the vocabulary coverage with regard to unseen cases: the system covered 97.8% of their terms. Evaluations showed that the system achieved high vocabulary coverage on unseen cases and was assessed as relevant for the task. Leonardo Campillos-Llanos, Catherine Thomas, Eric Bilinski, Pierre Zweigenbaum, Sophie Rosset |
Nat. Lang. Eng. | 4 |
| 2018 | Three Dimensions of Reproducibility in Natural Language Processing
Kevin Cohen 0001, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster R. Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, Lawrence Hunter |
LREC | 3 |
| 2018 | Combining rule-based and embedding-based approaches to normalize textual entities with an ontology
Arnaud Ferré, Louise Deléger, Pierre Zweigenbaum, Claire Nedellec |
LREC | 3 |
| 2018 | Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat
Christopher R. Norman, Mariska M. G. Leeflang, Pierre Zweigenbaum, Aurélie Névéol |
LREC | 3 |
| 2018 | A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora
Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp |
LREC | 1 |
| 2017 | Reproducibility in Biomedical Natural Language Processing
Kevin Cohen 0001, Aurélie Névéol, Jingbo Xia, Negacy D. Hailu, Cyril Grouin, Lawrence Hunter, Pierre Zweigenbaum |
AMIA | 7 |
| 2016 | Integrating a Dialogue System into a Virtual Patient Consultation
Leonardo Campillos-Llanos, Dhouha Bouamor, Eric Bilinski, Anne-Laure Ligozat, Sophie Rosset, Pierre Zweigenbaum |
AMIA | 6 |
| 2016 | Natural Language Processing Working Group Pre-Symposium: Graduate Student Consortium and 'Hackathon'
Stéphane M. Meystre, Sivaram Arabandi, Kavishwar B. Wagholikar, Jon D. Patrick, Guergana K. Savova, Chunhua Weng, Pierre Zweigenbaum, Dina Demner-Fushman, Özlem Uzuner, Hua Xu 0001 |
AMIA | 9 |
| 2016 | Adaptation of a Term Extractor to Arabic Specialised Texts: First Experiments and Limits
Wafa Neifar, Thierry Hamon, Pierre Zweigenbaum, Mariem Ellouze, Lamia Hadrich Belguith |
CICLing (1) | 3 |
| 2016 | Transfer-Based Learning-to-Rank Assessment of Medical Term Technicality
Dhouha Bouamor, Leonardo Campillos-Llanos, Anne-Laure Ligozat, Sophie Rosset, Pierre Zweigenbaum |
LREC | 5 |
| 2016 | Managing Linguistic and Terminological Variation in a Medical Dialogue System
Leonardo Campillos-Llanos, Dhouha Bouamor, Pierre Zweigenbaum, Sophie Rosset |
LREC | 3 |
| 2016 | Identification of Drug-Related Medical Conditions in Social Media
François Morlane-Hondère, Cyril Grouin, Pierre Zweigenbaum |
LREC | 3 |
| 2016 | PrefaceabstractAfter several decades of work on rule-based machine translation (MT) where linguists try to manually encode their knowledge about language, the time around 1990 brought a paradigm change towards automatic systems which try to learn how to translate by looking at large collections of high-quality sample translations as produced by professional translators. The first such attempts were called example- or analogy-based translation, and somewhat later the so-called statistical approach to MT was introduced. Both can be subsumed under the label data-driven approaches to MT. It took about 10 years until these self-learning systems became serious competitors of the traditional rule-based systems, and by now some of the most successful MT systems, such as Google Translate and Moses, are based on the statistical approach. Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum |
Nat. Lang. Eng. | 3 |
| 2016 | Recent advances in machine translation using comparable corporaabstractAbstract This paper highlights some of the recent developments in the field of machine translation using comparable corpora. We start by updating previous definitions of comparable corpora and then look at bilingual versions of continuous vector space models. Recently, neural networks have been used to obtain latent context representations with only few dimensions which are often called word embeddings. These promising new techniques cannot only be applied to parallel but also to comparable corpora. Subsequent sections of the paper discuss work specifically targeting at machine translation using comparable corpora, as well as work dealing with the extraction of parallel segments from comparable corpora. Finally, we give an overview on the design and the results of a recent shared task on measuring document comparability across languages. Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum |
Nat. Lang. Eng. | 3 |
| 2015 | Description of the PatientGenesys Dialogue SystemabstractLeonardo Campillos Llanos, Dhouha Bouamor, Éric Bilinski, Anne-Laure Ligozat, Pierre Zweigenbaum, Sophie Rosset. Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2015. Leonardo Campillos-Llanos, Dhouha Bouamor, Eric Bilinski, Anne-Laure Ligozat, Pierre Zweigenbaum, Sophie Rosset |
SIGDIAL Conference | 5 |
| 2015 | The contribution of co-reference resolution to supervised relation detection between bacteria and biotopes entitiesabstractBACKGROUND: The acquisition of knowledge about relations between bacteria and their locations (habitats and geographical locations) in short texts about bacteria, as defined in the BioNLP-ST 2013 Bacteria Biotope task, depends on the detection of co-reference links between mentions of entities of each of these three types. To our knowledge, no participant in this task has investigated this aspect of the situation. The present work specifically addresses issues raised by this situation: (i) how to detect these co-reference links and associated co-reference chains; (ii) how to use them to prepare positive and negative examples to train a supervised system for the detection of relations between entity mentions; (iii) what context around which entity mentions contributes to relation detection when co-reference chains are provided. RESULTS: We present experiments and results obtained both with gold entity mentions (task 2 of BioNLP-ST 2013) and with automatically detected entity mentions (end-to-end system, in task 3 of BioNLP-ST 2013). Our supervised mention detection system uses a linear chain Conditional Random Fields classifier, and our relation detection system relies on a Logistic Regression (aka Maximum Entropy) classifier. They use a set of morphological, morphosyntactic and semantic features. To minimize false inferences, co-reference resolution applies a set of heuristic rules designed to optimize precision. They take into account the types of the detected entity mentions, and take advantage of the didactic nature of the texts of the corpus, where a large proportion of bacteria naming is fairly explicit (although natural referring expressions such as "the bacteria" are common). The resulting system achieved a 0.495 F-measure on the official test set when taking as input the gold entity mentions, and a 0.351 F-measure when taking as input entity mentions predicted by our CRF system, both of which are above the best BioNLP-ST 2013 participant system. CONCLUSIONS: We show that co-reference resolution substantially improves over a baseline system which does not use co-reference information: about 3.5 F-measure points on the test corpus for the end-to-end system (5.5 points on the development corpus) and 7 F-measure points on both development and test corpora when gold mentions are used. While this outperforms the best published system on the BioNLP-ST 2013 Bacteria Biotope dataset, we consider that it provides mostly a stronger baseline from which more work can be started. We also emphasize the importance and difficulty of designing a comprehensive gold standard co-reference annotation, which we explain is a key point to further progress on the task. Thomas Lavergne, Cyril Grouin, Pierre Zweigenbaum |
BMC Bioinform. | 3 |
| 2015 | MEANS: A medical question-answering system combining NLP techniques and semantic Web technologies
Asma Ben Abacha, Pierre Zweigenbaum |
Inf. Process. Manag. | 2 |
| 2015 | Text mining for pharmacovigilance: Using machine learning for drug name recognition and drug-drug interaction extraction and classification
Asma Ben Abacha, Md. Faisal Mahbub Chowdhury, Aikaterini Karanasiou, Yassine Mrabet, Alberto Lavelli, Pierre Zweigenbaum |
J. Biomed. Informatics | 6 |
| 2015 | Combining glass box and black box evaluations in the identification of heart disease risk factors and their temporal relations from clinical recordsabstractBACKGROUND: The determination of risk factors and their temporal relations in natural language patient records is a complex task which has been addressed in the i2b2/UTHealth 2014 shared task. In this context, in most systems it was broadly decomposed into two sub-tasks implemented by two components: entity detection, and temporal relation determination. Task-level ("black box") evaluation is relevant for the final clinical application, whereas component-level evaluation ("glass box") is important for system development and progress monitoring. Unfortunately, because of the interaction between entity representation and temporal relation representation, glass box and black box evaluation cannot be managed straightforwardly at the same time in the setting of the i2b2/UTHealth 2014 task, making it difficult to assess reliably the relative performance and contribution of the individual components to the overall task. OBJECTIVE: To identify obstacles and propose methods to cope with this difficulty, and illustrate them through experiments on the i2b2/UTHealth 2014 dataset. METHODS: We outline several solutions to this problem and examine their requirements in terms of adequacy for component-level and task-level evaluation and of changes to the task framework. We select the solution which requires the least modifications to the i2b2 evaluation framework and illustrate it with our system. This system identifies risk factor mentions with a CRF system complemented by hand-designed patterns, identifies and normalizes temporal expressions through a tailored version of the Heideltime tool, and determines temporal relations of each risk factor with a One Rule classifier. RESULTS: Giving a fixed value to the temporal attribute in risk factor identification proved to be the simplest way to evaluate the risk factor detection component independently. This evaluation method enabled us to identify the risk factor detection component as most contributing to the false negatives and false positives of the global system. This led us to redirect further effort to this component, focusing on medication detection, with gains of 7 to 20 recall points and of 3 to 6 F-measure points depending on the corpus and evaluation. CONCLUSION: We proposed a method to achieve a clearer glass box evaluation of risk factor detection and temporal relation detection in clinical texts, which can provide an example to help system development in similar tasks. This glass box evaluation was instrumental in refocusing our efforts and obtaining substantial improvements in risk factor detection. Cyril Grouin, Véronique Moriceau, Pierre Zweigenbaum |
J. Biomed. Informatics | 3 |
| 2014 | Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Hercules Dalianis, Guergana K. Savova, Pierre Zweigenbaum |
AMIA | 4 |
| 2014 | Use of unsupervised word classes for entity recognition: Application to the detection of disorders in clinical reports
Maria Chatzimina, Cyril Grouin, Pierre Zweigenbaum |
LREC | 3 |
| 2014 | Annotation of specialized corpora using a comprehensive entity and relation scheme
Louise Deléger, Anne-Laure Ligozat, Cyril Grouin, Pierre Zweigenbaum, Aurélie Névéol |
LREC | 4 |
| 2014 | Language Resources for French in the Biomedical Domain
Aurélie Névéol, Julien Grosjean, Stéfan Jacques Darmoni, Pierre Zweigenbaum |
LREC | 4 |
| 2013 | Building Specialized Bilingual Lexicons Using Large Scale Background KnowledgeabstractBilingual lexicons are central components of machine translation and cross-lingual information retrieval systems.Their manual construction requires strong expertise in both languages involved and is a costly process.Several automatic methods were proposed as an alternative but they often rely on resources available in a limited number of languages and their performances are still far behind the quality of manual translations.We introduce a novel approach to the creation of specific domain bilingual lexicon that relies on Wikipedia.This massively multilingual encyclopedia makes it possible to create lexicons for a large number of language pairs.Wikipedia is used to extract domains in each language, to link domains between languages and to create generic translation dictionaries.The approach is tested on four specialized domains and is compared to three state of the art approaches using two language pairs: French-English and Romanian-English.The newly introduced method compares favorably to existing methods in all configurations tested. Dhouha Bouamor, Adrian Popescu 0001, Nasredine Semmar, Pierre Zweigenbaum |
EMNLP | 4 |
| 2013 | Building Specialized Bilingual Lexicons Using Word Sense Disambiguation
Dhouha Bouamor, Nasredine Semmar, Pierre Zweigenbaum |
IJCNLP | 3 |
| 2013 | Towards a Generic Approach for Bilingual Lexicon Extraction from Comparable Corpora
Dhouha Bouamor, Nasredine Semmar, Pierre Zweigenbaum |
MTSummit | 3 |
| 2013 | Eventual situations for timeline extraction from clinical reportsabstractOBJECTIVE: To identify the temporal relations between clinical events and temporal expressions in clinical reports, as defined in the i2b2/VA 2012 challenge. DESIGN: To detect clinical events, we used rules and Conditional Random Fields. We built Random Forest models to identify event modality and polarity. To identify temporal expressions we built on the HeidelTime system. To detect temporal relations, we systematically studied their breakdown into distinct situations; we designed an oracle method to determine the most prominent situations and the most suitable associated classifiers, and combined their results. RESULTS: We achieved F-measures of 0.8307 for event identification, based on rules, and 0.8385 for temporal expression identification. In the temporal relation task, we identified nine main situations in three groups, experimentally confirming shared intuitions: within-sentence relations, section-related time, and across-sentence relations. Logistic regression and Naïve Bayes performed best on the first and third groups, and decision trees on the second. We reached a 0.6231 global F-measure, improving by 7.5 points our official submission. CONCLUSIONS: Carefully hand-crafted rules obtained good results for the detection of events and temporal expressions, while a combination of classifiers improved temporal link prediction. The characterization of the oracle recall of situations allowed us to point at directions where further work would be most useful for temporal relation detection: within-sentence relations and linking History of Present Illness events to the admission date. We suggest that the systematic situation breakdown proposed in this paper could also help improve other systems addressing this task. Cyril Grouin, Natalia Grabar, Thierry Hamon, Sophie Rosset, Xavier Tannier, Pierre Zweigenbaum |
J. Am. Medical Informatics Assoc. | 6 |
| 2013 | A controlled greedy supervised approach for co-reference resolution on clinical text
Md. Faisal Mahbub Chowdhury, Pierre Zweigenbaum |
J. Biomed. Informatics | 2 |
| 2012 | Identifying bilingual Multi-Word Expressions for Statistical Machine Translation
Dhouha Bouamor, Nasredine Semmar, Pierre Zweigenbaum |
LREC | 3 |
| 2012 | Extended Named Entities Annotation on OCRed Documents: From Corpus Constitution to Evaluation Campaign
Olivier Galibert, Sophie Rosset, Cyril Grouin, Pierre Zweigenbaum, Ludovic Quintard |
LREC | 4 |
| 2011 | A Hybrid Approach for the Extraction of Semantic Relations from MEDLINE Abstracts
Asma Ben Abacha, Pierre Zweigenbaum |
CICLing (2) | 2 |
| 2011 | Structured and Extended Named Entity Evaluation in Automatic Speech Transcriptions
Olivier Galibert, Sophie Rosset, Cyril Grouin, Pierre Zweigenbaum, Ludovic Quintard |
IJCNLP | 4 |
| 2011 | Hybrid methods for improving information access in clinical documents: concept, assertion, and relation identificationabstractOBJECTIVE: This paper describes the approaches the authors developed while participating in the i2b2/VA 2010 challenge to automatically extract medical concepts and annotate assertions on concepts and relations between concepts. DESIGN: The authors'approaches rely on both rule-based and machine-learning methods. Natural language processing is used to extract features from the input texts; these features are then used in the authors' machine-learning approaches. The authors used Conditional Random Fields for concept extraction, and Support Vector Machines for assertion and relation annotation. Depending on the task, the authors tested various combinations of rule-based and machine-learning methods. RESULTS: The authors'assertion annotation system obtained an F-measure of 0.931, ranking fifth out of 21 participants at the i2b2/VA 2010 challenge. The authors' relation annotation system ranked third out of 16 participants with a 0.709 F-measure. The 0.773 F-measure the authors obtained on concept extraction did not make it to the top 10. CONCLUSION: On the one hand, the authors confirm that the use of only machine-learning methods is highly dependent on the annotated training data, and thus obtained better results for well-represented classes. On the other hand, the use of only a rule-based method was not sufficient to deal with new types of data. Finally, the use of hybrid approaches combining machine-learning and rule-based approaches yielded higher scores. Anne-Lyse Minard, Anne-Laure Ligozat, Asma Ben Abacha, Delphine Bernhard, Bruno Cartoni, Louise Deléger, Brigitte Grau, Sophie Rosset, Pierre Zweigenbaum, Cyril Grouin |
J. Am. Medical Informatics Assoc. | 9 |
| 2010 | Semi-Automated Extension of a Specialized Medical Lexicon for French
Bruno Cartoni, Pierre Zweigenbaum |
LREC | 2 |
| 2010 | Identifying Paraphrases between Technical and Lay Corpora
Louise Deléger, Pierre Zweigenbaum |
LREC | 2 |
| 2010 | Named and Specific Entity Detection in Varied Data: The Quæro Named Entity Baseline Evaluation
Olivier Galibert, Ludovic Quintard, Sophie Rosset, Pierre Zweigenbaum, Claire Nedellec, Sophie Aubin, Laurent Gillard, Jean-Pierre Raysz, Delphine Pois, Xavier Tannier, Louise Deléger, Dominique Laurent 0003 |
LREC | 4 |
| 2010 | Extracting medical information from narrative patient records: the case of medication-related informationabstractOBJECTIVE: While essential for patient care, information related to medication is often written as free text in clinical records and, therefore, difficult to use in computerized systems. This paper describes an approach to automatically extract medication information from clinical records, which was developed to participate in the i2b2 2009 challenge, as well as different strategies to improve the extraction. DESIGN: Our approach relies on a semantic lexicon and extraction rules as a two-phase strategy: first, drug names are recognized and, then, the context of these names is explored to extract drug-related information (mode, dosage, etc) according to rules capturing the document structure and the syntax of each kind of information. Different configurations are tested to improve this baseline system along several dimensions, particularly drug name recognition-this step being a determining factor to extract drug-related information. Changes were tested at the level of the lexicons and of the extraction rules. RESULTS: The initial system participating in i2b2 achieved good results (global F-measure of 77%). Further testing of different configurations substantially improved the system (global F-measure of 81%), performing well for all types of information (eg, 84% for drug names and 88% for modes), except for durations and reasons, which remain problematic. CONCLUSION: This study demonstrates that a simple rule-based system can achieve good performance on the medication extraction task. We also showed that controlled modifications (lexicon filtering and rule refinement) were the improvements that best raised the performance. Louise Deléger, Cyril Grouin, Pierre Zweigenbaum |
J. Am. Medical Informatics Assoc. | 3 |
| 2009 | Improvements in Analogical Learning: Application to Translating Multi-Terms of the Medical Domain
Philippe Langlais, François Yvon, Pierre Zweigenbaum |
EACL | 3 |
| 2009 | Translating medical terminologies through word alignment in parallel text corpora
Louise Deléger, Magnus Merkel, Pierre Zweigenbaum |
J. Biomed. Informatics | 3 |
| 2008 | Paraphrase Acquisition from Comparable Medical Corpora of Specialized and Lay Texts
Louise Deléger, Pierre Zweigenbaum |
AMIA | 2 |
| 2007 | MetaCoDe: A Lightweight UMLS Mapping Tool
Thierry Delbecque, Pierre Zweigenbaum |
AIME | 2 |
| 2007 | Frontiers of biomedical text mining: current progressabstractIt is now almost 15 years since the publication of the first paper on text mining in the genomics domain, and decades since the first paper on text mining in the medical domain. Enormous progress has been made in the areas of information retrieval, evaluation methodologies and resource construction. Some problems, such as abbreviation-handling, can essentially be considered solved problems, and others, such as identification of gene mentions in text, seem likely to be solved soon. However, a number of problems at the frontiers of biomedical text mining continue to present interesting challenges and opportunities for great improvements and interesting research. In this article we review the current state of the art in biomedical text mining or 'BioNLP' in general, focusing primarily on papers published within the past year. Pierre Zweigenbaum, Dina Demner-Fushman, Hong Yu 0001, Kevin Cohen 0001 |
Briefings Bioinform. | 1 |
| 2006 | Contribution to Terminology Internationalization by Word Alignment in Parallel Corpora
Louise Deléger, Magnus Merkel, Pierre Zweigenbaum |
AMIA | 3 |
| 2006 | Towards a Multilingual Medical Lexicon
Kornél G. Markó, Robert H. Baud, Pierre Zweigenbaum, Lars Borin, Magnus Merkel, Stefan Schulz 0001 |
AMIA | 3 |
| 2005 | Translating Biomedical Terms by Inferring Transducers
Vincent Claveau, Pierre Zweigenbaum |
AIME | 2 |
| 2005 | Interchanging Lexical Information for a Multilingual Dictionary
Robert H. Baud, Mikael Nyström, Lars Borin, Roger Evans, Stefan Schulz 0001, Pierre Zweigenbaum |
AMIA | 6 |
| 2003 | Learning Derived Words from Medical Corpora
Pierre Zweigenbaum, Natalia Grabar |
AIME | 1 |
| 2003 | VUMeF: Extending the French Involvement in the UMLS metathesaurus
Stéfan Jacques Darmoni, Éric Jarrousse, Pierre Zweigenbaum, Pierre Le Beux, Fiammetta Namer, Robert H. Baud, Michel Joubert, Huguette Vallée, Roger A. Côté, Antoine Buemi, Didier Bourigault, Gaëlle Recourcé, S. Jeanneau, Jean Marie Rodrigues |
AMIA | 3 |
| 2003 | UMLF: a Unified Medical Lexicon for French
Pierre Zweigenbaum, Robert H. Baud, Anita Burgun-Parenthoine, Fiammetta Namer, Éric Jarrousse, Natalia Grabar, Patrick Ruch, Franck Le Duff, Benoît Thirion, Stéfan Jacques Darmoni |
AMIA | 1 |
| 2003 | Corpus-Based Associations Provide Additional Morphological Variants to Medical Terminologies
Pierre Zweigenbaum, Natalia Grabar |
AMIA | 1 |
| 2002 | Looking for French-English translations in comparable medical corpora
Yun-Chuang Chiao, Pierre Zweigenbaum |
AMIA | 2 |
| 2002 | A Study of the Adequacy of User and Indexing Vocabularies in Natural Language Queries to a MeSH-indexed Health Gateway
Natalia Grabar, Pierre Zweigenbaum, Lina Fatima Soualmia, Stéfan Jacques Darmoni |
AMIA | 2 |
| 2002 | An assessment of the visibility of MeSH-indexed medical web catalogs through search engines
Pierre Zweigenbaum, Stéfan Jacques Darmoni, Natalia Grabar, Magaly Douyère, Jacques Benichou |
AMIA | 1 |
| 2002 | Looking for Candidate Translational Equivalents in Specialized, Comparable Corpora
Yun-Chuang Chiao, Pierre Zweigenbaum |
COLING | 2 |
| 2001 | The contribution of morphological knowledge to French MeSH mapping for information retrieval
Pierre Zweigenbaum, Stéfan Jacques Darmoni, Natalia Grabar |
AMIA | 1 |
| 2000 | A general method for sifting linguistic knowledge from structured terminologies
Natalia Grabar, Pierre Zweigenbaum |
AMIA | 2 |
| 1999 | A Lexical Method for Assisted Extraction and Coding of ICD-10 Diagnoses from Free Text Patient Discharge Summaries
Alexandre Blanquet, Pierre Zweigenbaum |
AMIA | 2 |
| 1999 | Language-independent automatic acquisition of morphological knowledge from synonym pairs
Natalia Grabar, Pierre Zweigenbaum |
AMIA | 2 |
| 1998 | Hospitexte: towards a document-based hypertextual electronic medical record
Jean Charlet, Bruno Bachimont, Vincent Brunie, Sawsan El Kassar, Pierre Zweigenbaum, Jean-François Boisvieux |
AMIA | 5 |
| 1998 | Tuning an Existing Nomenclature for Specific Domain Corpora: A Syntax-Based Similarity Method
Pierre Zweigenbaum, Benoit Habert, Adeline Nazarenko, Jacques Bouaud |
AMIA | 1 |
| 1998 | Extending an existing specialized semantic lexicon
Benoit Habert, Adeline Nazarenko, Pierre Zweigenbaum, Jacques Bouaud |
LREC | 3 |
| 1997 | Building a Normalized Conceptual Representation of Medical Language: A Semantic Composition Method Driven by Domain Knowledge Models
Jacques Bouaud, Pierre Zweigenbaum, Bruno Bachimont |
AMIA | 2 |
| 1997 | Corpus-based identification and refinement of semantic classes
Adeline Nazarenko, Pierre Zweigenbaum, Jacques Bouaud, Benoit Habert |
AMIA | 2 |
| 1997 | Evaluating a normalized conceptual representation produced from natural language patient discharge summaries
Pierre Zweigenbaum, Jacques Bouaud, Bruno Bachimont, Jean Charlet, Jean-François Boisvieux |
AMIA | 1 |
| 1996 | Processing Metonymy- a Domain-Model Heuristic Graph Traversal Approach
Jacques Bouaud, Bruno Bachimont, Pierre Zweigenbaum |
COLING | 3 |
| 1992 | First Results of a French Linguistic Development Environment
Lorne H. Bouchard, Louisette Emirkanian, Dominique Estival, Christine Fay-Varnier, Christophe Fouqueré, Gilles Prigent, Pierre Zweigenbaum |
COLING | 7 |
| 1992 | Extracting Implicit Information from Free Text Technical Reports
Marc Cavazza, Pierre Zweigenbaum |
Inf. Process. Manag. | 2 |
| 1990 | Deep Sentence Understanding in a Restricted Domain
Pierre Zweigenbaum, Marc Cavazza |
COLING | 1 |