Aurélie Névéol

dblp:37/3908 · DBLP profile ↗
← Back
59ranked-venue papers
18as first author
17since 2021 · last 2026
0000-0002-1846-9144ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Code-switching as a Bias Indicator in LLMs: "the Consequences Are Not the Same Para Nosotros"
abstract
International audience
Fanny Ducel, Aurélie Névéol, Vidit Khazanchi, Loïc Leclere, Arthur Pedrini, Léa Bouchet, Benjamin Caissial, Karën Fort
LREC2
2026 Supporting workflow reproducibility by linking bioinformatics tools across papers and executable code
abstract
MOTIVATION: The rapid growth of biological data has intensified the need for transparent, reproducible, and well-documented computational workflows. The ability to clearly connect the steps of a workflow in the code with their description in a paper would improve workflow comprehension, support reproducibility, and facilitate reuse. This task requires the linking of bioinformatics tools in workflow code with their mentions in a published workflow description. RESULTS: We present CoPaLink, an automated approach that integrates three components: named entity recognition (NER) for identifying tool mentions in scientific text, NER for tool mentions in workflow code, and entity resolution based on word embedding similarity. We propose approaches for all three steps, achieving a high individual F1-measure (77-90) and a joint accuracy of 66 when evaluated on Nextflow workflows using Sentence-BERT. CoPaLink leverages corpora of scientific articles and workflow executable code with curated tool annotations to bridge the gap between narrative descriptions and workflow implementations. AVAILABILITY AND IMPLEMENTATION: The code is available at https://gitlab.liris.cnrs.fr/sharefair/copalink-experiments and https://gitlab.liris.cnrs.fr/sharefair/copalink. The corpora are also available: CPL-Article (https://doi.org/10.5281/zenodo.20746904), CPL-Code (https://doi.org/10.5281/zenodo.20746970) and CPL-Gold-Entity-Resolution (https://doi.org/10.5281/zenodo.20746994).
Clémence Sebe, Olivier Ferret, Aurélie Névéol, Mahdi Esmailoghli, Ulf Leser, Sarah Cohen Boulakia
Bioinform.3
2025 Evaluating the Confidentiality of Synthetic Clinical Texts Generated by Language Models
Foucauld Estignard, Sahar Ghannay, Julien Girard-Satabin, Nicolas Hiebel, Aurélie Névéol
AIME (1)5
2025 Extracting Information in a Low-Resource Setting: Case Study on Bioinformatics Workflows
Clémence Sebe, Sarah Cohen Boulakia, Olivier Ferret, Aurélie Névéol
IDA4
2025 SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
abstract
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna-Adriana Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L. Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir R. Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh D. Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Jay Gala, Hamdan Al-Ali, Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
NAACL (Long Papers)51
2025 Efficient extraction of medication information from clinical notes: an evaluation in 2 languages
abstract
OBJECTIVE: To evaluate the accuracy, computational cost, and portability of a new natural language processing (NLP) method for extracting medication information from clinical narratives. MATERIALS AND METHODS: We propose an original transformer-based architecture for the extraction of entities and their relations pertaining to patients' medication regimen. First, we used this approach to train and evaluate a model on French clinical notes, using a newly annotated corpus from Hôpitaux Universitaires de Strasbourg. Second, the portability of the approach was assessed by conducting an evaluation on clinical documents in English from the 2018 n2c2 shared task. Information extraction accuracy and computational cost were assessed by comparison with an available method using transformers. RESULTS: The proposed architecture achieves on the task of relation extraction itself performance that are competitive with the state-of-the-art on both French and English (F-measures 0.82 and 0.96 vs 0.81 and 0.95), but reduces the computational cost by 10. End-to-end (Named Entity recognition and Relation Extraction) F1 performance is 0.69 and 0.82 for French and English corpus. DISCUSSION: While an existing system developed for English notes was deployed in a French hospital setting with reasonable effort, we found that an alternative architecture offered end-to-end drug information extraction with comparable extraction performance and lower computational impact for both French and English clinical text processing, respectively. CONCLUSION: The proposed architecture can be used to extract medication information from clinical text with high performance and low computational cost and consequently suits with usually limited hospital IT resources.
Thibaut Fabacher, Erik-André Sauleau, Emmanuelle Arcay, Bineta Faye, Maxime Alter, Archia Chahard, Nathan Miraillet, Adrien Coulet, Aurélie Névéol
J. Am. Medical Informatics Assoc.9
2024 Limitations of Human Identification of Automatically Generated Text
abstract
Neural text generation is receiving broad attention with the publication of new tools such as ChatGPT. The main reason for that is that the achieved quality of the generated text may be attributed to a human writer by the naked eye of a human evaluator. In this paper, we propose a new corpus in French and English for the task of recognising automatically generated texts and we conduct a study of how humans perceive the text. Our results show, as previous work before the ChatGPT era, that the generated texts by tools such as ChatGPT share some common characteristics but they are not clearly identifiable which generates different perceptions of these texts.
Nadège Alavoine, Maximin Coavoux, Emmanuelle Esperança-Rodier, Romane Gallienne, Carlos E. González-Gallardo, Jérôme Goulian, José G. Moreno 0001, Aurélie Névéol, Didier Schwab, Vincent Segonne, Johanna Simoens
LREC/COLING8
2024 A Benchmark Evaluation of Clinical Named Entity Recognition in French
abstract
Background: Transformer-based language models have shown strong performance on many Natural Language Processing (NLP) tasks. Masked Language Models (MLMs) attract sustained interest because they can be adapted to different languages and sub-domains through training or fine-tuning on specific corpora while remaining lighter than modern Large Language Models (MLMs). Recently, several MLMs have been released for the biomedical domain in French, and experiments suggest that they outperform standard French counterparts. However, no systematic evaluation comparing all models on the same corpora is available. Objective: This paper presents an evaluation of masked language models for biomedical French on the task of clinical named entity recognition. Material and methods: We evaluate biomedical models CamemBERT-bio and DrBERT and compare them to standard French models CamemBERT, FlauBERT and FrAlBERT as well as multilingual mBERT using three publically available corpora for clinical named entity recognition in French. The evaluation set-up relies on gold-standard corpora as released by the corpus developers. Results: Results suggest that CamemBERT-bio outperforms DrBERT consistently while FlauBERT offers competitive performance and FrAlBERT achieves the lowest carbon footprint. Conclusion: This is the first benchmark evaluation of biomedical masked language models for French clinical entity recognition that compares model performance consistently on nested entity recognition using metrics covering performance and environmental impact.
Nesrine Bannour, Christophe Servan, Aurélie Névéol, Xavier Tannier
LREC/COLING3
2024 Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts
abstract
Warning: This paper contains explicit statements of offensive stereotypes which may be upsetting The study of bias, fairness and social impact in Natural Language Processing (NLP) lacks resources in languages other than English. Our objective is to support the evaluation of bias in language models in a multilingual setting. We use stereotypes across nine types of biases to build a corpus containing contrasting sentence pairs, one sentence that presents a stereotype concerning an underadvantaged group and another minimally changed sentence, concerning a matching advantaged group. We build on the French CrowS-Pairs corpus and guidelines to provide translations of the existing material into seven additional languages. In total, we produce 11,139 new sentence pairs that cover stereotypes dealing with nine types of biases in seven cultural contexts. We use the final resource for the evaluation of relevant monolingual and multilingual masked language models. We find that language models in all languages favor sentences that express stereotypes in most bias categories. The process of creating a resource that covers a wide range of language types and cultural settings highlights the difficulty of bias evaluation, in particular comparability across languages and contexts.
Karën Fort, Laura Alonso Alemany, Luciana Benotti, Julien Bezançon, Claudia Borg, Marthese Borg, Yongjian Chen, Fanny Ducel, Yoann Dupont, Guido Ivetta, Margot Mieskes, Marco Naguib, Yuyan Qian, Matteo Radaelli, Wolfgang Schmeisser-Nieto, Emma Raimundo Schulz, Thiziri Saci, Sarah Saidi, Javier Torroba Marchante, Shilin Xie, Sergio E. Zanotto, Aurélie Névéol
LREC/COLING23
2024 A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages
abstract
User-generated data sources have gained significance in uncovering Adverse Drug Reactions (ADRs), with an increasing number of discussions occurring in the digital world. However, the existing clinical corpora predominantly revolve around scientific articles in English. This work presents a multilingual corpus of texts concerning ADRs gathered from diverse sources, including patient fora, social media, and clinical reports in German, French, and Japanese. Our corpus contains annotations covering 12 entity types, four attribute types, and 13 relation types. It contributes to the development of real-world multilingual language models for healthcare. We provide statistics to highlight certain challenges associated with the corpus and conduct preliminary experiments resulting in strong baselines for extracting entities and relations between these entities, both within and across languages.
Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas 0001, Tomohiro Nishiyama, Sebastian Möller 0001, Eiji Aramaki, Yuji Matsumoto 0001, Roland Roller, Pierre Zweigenbaum
LREC/COLING6
2023 The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research
abstract
Mohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aurélie Névéol, Fanny Ducel, Saif Mohammad, Karen Fort. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Mohamed Abdalla 0001, Jan Philip Wahle, Terry Ruas, Aurélie Névéol, Fanny Ducel, Saif M. Mohammad, Karën Fort
ACL (1)4
2023 Can Synthetic Text Help Clinical Named Entity Recognition? A Study of Electronic Health Records in French
abstract
In sensitive domains, the sharing of corpora is restricted due to confidentiality, copyrights, or trade secrets.Automatic text generation can help alleviate these issues by producing synthetic texts that mimic the linguistic properties of real documents while preserving confidentiality.In this study, we assess the usability of synthetic corpus as a substitute training corpus for clinical information extraction.Our goal is to automatically produce a clinical case corpus annotated with clinical entities and to evaluate it for a named entity recognition (NER) task.We use two auto-regressive neural models partially or fully trained on generic French texts and fine-tuned on clinical cases to produce a corpus of synthetic clinical cases.We study variants of the generation process: (i) fine-tuning on annotated vs. plain text (in that case, annotations are obtained a posteriori) and (ii) selection of generated texts based on models' parameters and filtering criteria.We then train NER models with the resulting synthetic text and evaluate them on a gold standard clinical corpus.Our experiments suggest that synthetic text is useful for clinical NER.
Nicolas Hiebel, Olivier Ferret, Karën Fort, Aurélie Névéol
EACL4
2022 French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English
abstract
Warning: This paper contains explicit statements of offensive stereotypes which may be upsetting Much work on biases in natural language processing has addressed biases linked to the social and cultural experience of English speaking individuals in the United States.We seek to widen the scope of bias studies by creating material to measure social bias in language models (LMs) against specific demographic groups in France.We build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages while also characterizing biases that are specific to each country and language.We introduce 1,677 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.1,467 sentence pairs are translated from CrowS-pairs and 210 are newly crowdsourced and translated back into English.The sentence pairs contrast stereotypes concerning underadvantaged groups with the same sentence concerning advantaged groups.We find that four widely used language models (three French, one multilingual) favor sentences that express stereotypes in most bias categories.We report on the translation process, which led to a characterization of stereotypes in CrowS-pairs including the identification of US-centric cultural traits.We offer guidelines to further extend the dataset to other languages and cultural environments.
Aurélie Névéol, Yoann Dupont, Julien Bezançon, Karën Fort
ACL (1)1
2022 CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives
abstract
Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluating models. Such resources are scarce, especially for specialized domains in languages other than English. In particular, there are very few resources for semantic similarity in the clinical domain in French. This can be useful for many biomedical natural language processing applications, including text generation. We introduce a definition of similarity that is guided by clinical facts and apply it to the development of a new French corpus of 1,000 sentence pairs manually annotated according to similarity scores. This new sentence similarity corpus is made freely available to the community. We further evaluate the corpus through experiments of automatic similarity measurement. We show that a model of sentence embeddings can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set (Spearman=0.8343). We also show that the corpus is complementary to DEFT STS.
Nicolas Hiebel, Olivier Ferret, Karën Fort, Aurélie Névéol
LREC4
2022 Privacy-preserving mimic models for clinical named entity recognition in French
Nesrine Bannour, Perceval Wajsbürt, Bastien Rance, Xavier Tannier, Aurélie Névéol
J. Biomed. Informatics5
2021 Can reproducibility be improved in clinical natural language processing? A study of 7 clinical NLP suites
William Digan, Aurélie Névéol, Antoine Neuraz, Maxime Wack, David Baudoin, Anita Burgun-Parenthoine, Bastien Rance
AMIA2
2021 Can reproducibility be improved in clinical natural language processing? A study of 7 clinical NLP suites
abstract
BACKGROUND: The increasing complexity of data streams and computational processes in modern clinical health information systems makes reproducibility challenging. Clinical natural language processing (NLP) pipelines are routinely leveraged for the secondary use of data. Workflow management systems (WMS) have been widely used in bioinformatics to handle the reproducibility bottleneck. OBJECTIVE: To evaluate if WMS and other bioinformatics practices could impact the reproducibility of clinical NLP frameworks. MATERIALS AND METHODS: Based on the literature across multiple researcho fields (NLP, bioinformatics and clinical informatics) we selected articles which (1) review reproducibility practices and (2) highlight a set of rules or guidelines to ensure tool or pipeline reproducibility. We aggregate insight from the literature to define reproducibility recommendations. Finally, we assess the compliance of 7 NLP frameworks to the recommendations. RESULTS: We identified 40 reproducibility features from 8 selected articles. Frameworks based on WMS match more than 50% of features (26 features for LAPPS Grid, 22 features for OpenMinted) compared to 18 features for current clinical NLP framework (cTakes, CLAMP) and 17 features for GATE, ScispaCy, and Textflows. DISCUSSION: 34 recommendations are endorsed by at least 2 articles from our selection. Overall, 15 features were adopted by every NLP Framework. Nevertheless, frameworks based on WMS had a better compliance with the features. CONCLUSION: NLP frameworks could benefit from lessons learned from the bioinformatics field (eg, public repositories of curated tools and workflows or use of containers for shareability) to enhance the reproducibility in a clinical setting.
William Digan, Aurélie Névéol, Antoine Neuraz, Maxime Wack, David Baudoin, Anita Burgun-Parenthoine, Bastien Rance
J. Am. Medical Informatics Assoc.2
2020 MEDLINE as a Parallel Corpus: a Survey to Gain Insight on French-, Spanish- and Portuguese-speaking Authors' Abstract Writing Practice
abstract
Background: Parallel corpora are used to train and evaluate machine translation systems. To alleviate the cost of producing parallel resources for evaluation campaigns, existing corpora are leveraged. However, little information may be available about the methods used for producing the corpus, including translation direction. Objective: To gain insight on MEDLINE parallel corpus used in the biomedical task at the Workshop on Machine Translation in 2019 (WMT 2019). Material and Methods: Contact information for the authors of MEDLINE articles included in the English/Spanish (EN/ES), English/French (EN/FR), and English/Portuguese (EN/PT) WMT 2019 test sets was obtained from PubMed and publisher websites. The authors were asked about their abstract writing practices in a survey. Results: The response rate was above 20%. Authors reported that they are mainly native speakers of languages other than English. Although manual translation, sometimes via professional translation services, was commonly used for abstract translation, authors of articles in the EN/ES and EN/PT sets also relied on post-edited machine translation. Discussion: This study provides a characterization of MEDLINE authors’ language skills and abstract writing practices. Conclusion: The information collected in this study will be used to inform test set design for the next WMT biomedical task.
Aurélie Névéol, Antonio Jimeno-Yepes, Mariana L. Neves
LREC1
2020 Ten simple rules to make your research more sustainable
abstract
DOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone.
Anne-Laure Ligozat, Aurélie Névéol, Bénédicte Daly, Emmanuelle Frenoux
PLoS Comput. Biol.2
2019 Measuring the impact of screening automation on meta-analyses of diagnostic test accuracy
Christopher R. Norman, Mariska M. G. Leeflang, Raphaël Porcher, Aurélie Névéol
AMIA4
2018 Methods and Tools to Enhance Rigor and Reproducibility of Biomedical Research
Halil Kilicoglu, Aurélie Névéol, Timothy Clark, Hua Xu 0001, Neil R. Smalheiser
AMIA2
2018 Data Extraction and Synthesis in Systematic Reviews of Diagnostic Test Accuracy: A Corpus for Automating and Evaluating the Process
Christopher R. Norman, Mariska M. G. Leeflang, Aurélie Névéol
AMIA3
2018 Three Dimensions of Reproducibility in Natural Language Processing
Kevin Cohen 0001, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster R. Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, Lawrence Hunter
LREC8
2018 Parallel Corpora for the Biomedical Domain
Aurélie Névéol, Antonio Jimeno-Yepes, Mariana L. Neves, Karin Verspoor
LREC1
2018 Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat
Christopher R. Norman, Mariska M. G. Leeflang, Pierre Zweigenbaum, Aurélie Névéol
LREC4
2017 Reproducibility in Biomedical Natural Language Processing
Kevin Cohen 0001, Aurélie Névéol, Jingbo Xia, Negacy D. Hailu, Cyril Grouin, Lawrence Hunter, Pierre Zweigenbaum
AMIA2
2017 Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Noémie Elhadad, Sumithra Velupillai, Hua Xu 0001, Guergana K. Savova
AMIA1
2016 The Scielo Corpus: a Parallel Corpus of Scientific Publications for Biomedicine
Mariana L. Neves, Antonio Jimeno-Yepes, Aurélie Névéol
LREC3
2016 Automatic classification of registered clinical trials towards the Global Burden of Diseases taxonomy of diseases and injuries
abstract
BACKGROUND: Clinical trial registries may allow for producing a global mapping of health research. However, health conditions are not described with standardized taxonomies in registries. Previous work analyzed clinical trial registries to improve the retrieval of relevant clinical trials for patients. However, no previous work has classified clinical trials across diseases using a standardized taxonomy allowing a comparison between global health research and global burden across diseases. We developed a knowledge-based classifier of health conditions studied in registered clinical trials towards categories of diseases and injuries from the Global Burden of Diseases (GBD) 2010 study. The classifier relies on the UMLS® knowledge source (Unified Medical Language System®) and on heuristic algorithms for parsing data. It maps trial records to a 28-class grouping of the GBD categories by automatically extracting UMLS concepts from text fields and by projecting concepts between medical terminologies. The classifier allows deriving pathways between the clinical trial record and candidate GBD categories using natural language processing and links between knowledge sources, and selects the relevant GBD classification based on rules of prioritization across the pathways found. We compared automatic and manual classifications for an external test set of 2,763 trials. We automatically classified 109,603 interventional trials registered before February 2014 at WHO ICTRP. RESULTS: In the external test set, the classifier identified the exact GBD categories for 78 % of the trials. It had very good performance for most of the 28 categories, especially "Neoplasms" (sensitivity 97.4 %, specificity 97.5 %). The sensitivity was moderate for trials not relevant to any GBD category (53 %) and low for trials of injuries (16 %). For the 109,603 trials registered at WHO ICTRP, the classifier did not assign any GBD category to 20.5 % of trials while the most common GBD categories were "Neoplasms" (22.8 %) and "Diabetes" (8.9 %). CONCLUSIONS: We developed and validated a knowledge-based classifier allowing for automatically identifying the diseases studied in registered trials by using the taxonomy from the GBD 2010 study. This tool is freely available to the research community and can be used for large-scale public health studies.
Ignacio Atal, Jean-David Zeitoun, Aurélie Névéol, Philippe Ravaud, Raphaël Porcher, Ludovic Trinquart
BMC Bioinform.3
2015 Automatic Extraction of Time Expressions Accross Domains in French Narratives
abstract
The prevalence of temporal references across all types of natural language utterances makes temporal analysis a key issue in Natural Language Processing.This work adresses three research questions: 1/is temporal expression recognition specific to a particular domain?2/if so, can we characterize domain specificity?and 3/how can subdomain specificity be integrated in a single tool for unified temporal expression extraction?Herein, we assess temporal expression recognition from documents written in French covering three domains.We present a new corpus of clinical narratives annotated for temporal expressions, and also use existing corpora in the newswire and historical domains.We show that temporal expressions can be extracted with high performance across domains (best F-measure 0.96 obtained with a CRF model on clinical narratives).We argue that domain adaptation for the extraction of temporal expressions can be done with limited efforts and should cover pre-processing as well as temporal specific tasks.
Mike Donald Tapi-Nzali, Xavier Tannier, Aurélie Névéol
EMNLP3
2014 Automatic Content Extraction for Designing a French Clinical Corpus
Louise Deléger, Cyril Grouin, Aurélie Névéol
AMIA3
2014 How to de-identify a large clinical corpus in 10 days
Cyril Grouin, Louise Deléger, Jean-Baptiste Escudié, Gregory Groisy, Anne-Sophie Jannot, Bastien Rance, Xavier Tannier, Aurélie Névéol
AMIA8
2014 Clinical Natural Language Processing in Languages Other Than English
Aurélie Névéol, Hercules Dalianis, Guergana K. Savova, Pierre Zweigenbaum
AMIA1
2014 Annotation of specialized corpora using a comprehensive entity and relation scheme
Louise Deléger, Anne-Laure Ligozat, Cyril Grouin, Pierre Zweigenbaum, Aurélie Névéol
LREC5
2014 Language Resources for French in the Biomedical Domain
Aurélie Névéol, Julien Grosjean, Stéfan Jacques Darmoni, Pierre Zweigenbaum
LREC1
2014 Natural language processing of radiology reports for the detection of thromboembolic diseases and clinically relevant incidental findings
abstract
BACKGROUND: Natural Language Processing (NLP) has been shown effective to analyze the content of radiology reports and identify diagnosis or patient characteristics. We evaluate the combination of NLP and machine learning to detect thromboembolic disease diagnosis and incidental clinically relevant findings from angiography and venography reports written in French. We model thromboembolic diagnosis and incidental findings as a set of concepts, modalities and relations between concepts that can be used as features by a supervised machine learning algorithm. A corpus of 573 radiology reports was de-identified and manually annotated with the support of NLP tools by a physician for relevant concepts, modalities and relations. A machine learning classifier was trained on the dataset interpreted by a physician for diagnosis of deep-vein thrombosis, pulmonary embolism and clinically relevant incidental findings. Decision models accounted for the imbalanced nature of the data and exploited the structure of the reports. RESULTS: The best model achieved an F measure of 0.98 for pulmonary embolism identification, 1.00 for deep vein thrombosis, and 0.80 for incidental clinically relevant findings. The use of concepts, modalities and relations improved performances in all cases. CONCLUSIONS: This study demonstrates the benefits of developing an automated method to identify medical concepts, modality and relations from radiology reports in French. An end-to-end automatic system for annotation and classification which could be applied to other radiology reports databases would be valuable for epidemiological surveillance, performance monitoring, and accreditation in French hospitals.
Anne-Dominique Pham, Aurélie Névéol, Thomas Lavergne, Daisuke Yasunaga, Olivier Clément, Guy Meyer, Rémy Morello, Anita Burgun-Parenthoine
BMC Bioinform.2
2014 De-identification of clinical notes in French: towards a protocol for reference corpus development
Cyril Grouin, Aurélie Névéol
J. Biomed. Informatics2
2013 A systematic comparison of current sources of disease knowledge
Aurélie Névéol, Bastien Rance, Zhiyong Lu
AMIA1
2013 Combining MEDLINE and publisher data to create parallel corpora for the automatic translation of biomedical text
abstract
BACKGROUND: Most of the institutional and research information in the biomedical domain is available in the form of English text. Even in countries where English is an official language, such as the United States, language can be a barrier for accessing biomedical information for non-native speakers. Recent progress in machine translation suggests that this technique could help make English texts accessible to speakers of other languages. However, the lack of adequate specialized corpora needed to train statistical models currently limits the quality of automatic translations in the biomedical domain. RESULTS: We show how a large-sized parallel corpus can automatically be obtained for the biomedical domain, using the MEDLINE database. The corpus generated in this work comprises article titles obtained from MEDLINE and abstract text automatically retrieved from journal websites, which substantially extends the corpora used in previous work. After assessing the quality of the corpus for two language pairs (English/French and English/Spanish) we use the Moses package to train a statistical machine translation model that outperforms previous models for automatic translation of biomedical text. CONCLUSIONS: We have built translation data sets in the biomedical domain that can easily be extended to other languages available in MEDLINE. These sets can successfully be applied to train statistical machine translation models. While further progress should be made by incorporating out-of-domain corpora and domain-specific lexicons, we believe that this work improves the automatic translation of biomedical texts.
Antonio Jimeno-Yepes, Élise Prieur, Aurélie Névéol
BMC Bioinform.3
2011 Extraction of data deposition statements from the literature: a method for automatically tracking research results
abstract
MOTIVATION: Research in the biomedical domain can have a major impact through open sharing of the data produced. For this reason, it is important to be able to identify instances of data production and deposition for potential re-use. Herein, we report on the automatic identification of data deposition statements in research articles. RESULTS: We apply machine learning algorithms to sentences extracted from full-text articles in PubMed Central in order to automatically determine whether a given article contains a data deposition statement, and retrieve the specific statements. With an Support Vector Machine classifier using conditional random field determined deposition features, articles containing deposition statements are correctly identified with 81% F-measure. An error analysis shows that almost half of the articles classified as containing a deposition statement by our method but not by the gold standard do indeed contain a deposition statement. In addition, our system was used to process articles in PubMed Central, predicting that a total of 52 932 articles report data deposition, many of which are not currently included in the Secondary Source Identifier [si] field for MEDLINE citations. AVAILABILITY: All annotated datasets described in this study are freely available from the NLM/NCBI website at http://www.ncbi.nlm.nih.gov/CBBresearch/Fellows/Neveol/DepositionDataSets.zip CONTACT: [email protected]; [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Aurélie Névéol, W. John Wilbur, Zhiyong Lu
Bioinform.1
2011 A context-blocks model for identifying clinical relationships in patient records
abstract
BACKGROUND: Patient records contain valuable information regarding explanation of diagnosis, progression of disease, prescription and/or effectiveness of treatment, and more. Automatic recognition of clinically important concepts and the identification of relationships between those concepts in patient records are preliminary steps for many important applications in medical informatics, ranging from quality of care to hypothesis generation. METHODS: In this work we describe an approach that facilitates the automatic recognition of eight relationships defined between medical problems, treatments and tests. Unlike the traditional bag-of-words representation, in this work, we represent a relationship with a scheme of five distinct context-blocks determined by the position of concepts in the text. As a preliminary step to relationship recognition, and in order to provide an end-to-end system, we also addressed the automatic extraction of medical problems, treatments and tests. Our approach combined the outcome of a statistical model for concept recognition and simple natural language processing features in a conditional random fields model. A set of 826 patient records from the 4th i2b2 challenge was used for training and evaluating the system. RESULTS: Results show that our concept recognition system achieved an F-measure of 0.870 for exact span concept detection. Moreover the context-block representation of relationships was more successful (F-Measure = 0.775) at identifying relationships than bag-of-words (F-Measure = 0.402). Most importantly, the performance of the end-to-end system of relationship extraction using automatically extracted concepts (F-Measure = 0.704) was comparable to that obtained using manually annotated concepts (F-Measure = 0.711), and their difference was not statistically significant. CONCLUSIONS: We extracted important clinical relationships from text in an automated manner, starting with concept recognition, and ending with relationship identification. The advantage of the context-blocks representation scheme was the correct management of word position information, which may be critical in identifying certain relationships. Our results may serve as benchmark for comparison to other systems developed on i2b2 challenge data. Finally, our system may serve as a preliminary step for other discovery tasks in medical informatics.
Rezarta Islamaj Dogan, Aurélie Névéol, Zhiyong Lu
BMC Bioinform.2
2011 Recommending MeSH terms for annotating biomedical articles
abstract
BACKGROUND: Due to the high cost of manual curation of key aspects from the scientific literature, automated methods for assisting this process are greatly desired. Here, we report a novel approach to facilitate MeSH indexing, a challenging task of assigning MeSH terms to MEDLINE citations for their archiving and retrieval. METHODS: Unlike previous methods for automatic MeSH term assignment, we reformulate the indexing task as a ranking problem such that relevant MeSH headings are ranked higher than those irrelevant ones. Specifically, for each document we retrieve 20 neighbor documents, obtain a list of MeSH main headings from neighbors, and rank the MeSH main headings using ListNet-a learning-to-rank algorithm. We trained our algorithm on 200 documents and tested on a previously used benchmark set of 200 documents and a larger dataset of 1000 documents. RESULTS: Tested on the benchmark dataset, our method achieved a precision of 0.390, recall of 0.712, and mean average precision (MAP) of 0.626. In comparison to the state of the art, we observe statistically significant improvements as large as 39% in MAP (p-value <0.001). Similar significant improvements were also obtained on the larger document set. CONCLUSION: Experimental results show that our approach makes the most accurate MeSH predictions to date, which suggests its great potential in making a practical impact on MeSH indexing. Furthermore, as discussed the proposed learning framework is robust and can be adapted to many other similar tasks beyond MeSH indexing in the biomedical domain. All data sets are available at: http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/indexing.
Minlie Huang, Aurélie Névéol, Zhiyong Lu
J. Am. Medical Informatics Assoc.2
2011 Semi-automatic semantic annotation of PubMed queries: A study on quality, efficiency, satisfaction
Aurélie Névéol, Rezarta Islamaj Dogan, Zhiyong Lu
J. Biomed. Informatics1
2010 A Textual Representation Scheme for Identifying Clinical Relationships in Patient Records
abstract
The identification of relationships between clinical concepts in patient records is a preliminary step for many important applications in medical informatics, ranging from quality of care to hypothesis generation. In this work we describe an approach that facilitates the automatic recognition of relationships defined between two different concepts in text. Unlike the traditional bag-of-words representation, in this work, a relationship is represented with a scheme of five distinct context-blocks based on the position of concepts in the text. This scheme was applied to eight different relationships, between medical problems, treatments and tests, on a set of 349 patient records from the 4th i2b2 challenge. Results show that the context-block representation was very successful (F-Measure = 0.775) compared to the bag-of-words model (F-Measure = 0.402). The advantage of this representation scheme was the correct management of word position information, which may be critical in identifying certain relationships.
Rezarta Islamaj Dogan, Aurélie Névéol, Zhiyong Lu
ICMLA2
2010 Extracting Rx information from clinical narrative
abstract
OBJECTIVE: The authors used the i2b2 Medication Extraction Challenge to evaluate their entity extraction methods, contribute to the generation of a publicly available collection of annotated clinical notes, and start developing methods for ontology-based reasoning using structured information generated from the unstructured clinical narrative. DESIGN: Extraction of salient features of medication orders from the text of de-identified hospital discharge summaries was addressed with a knowledge-based approach using simple rules and lookup lists. The entity recognition tool, MetaMap, was combined with dose, frequency, and duration modules specifically developed for the Challenge as well as a prototype module for reason identification. MEASUREMENTS: Evaluation metrics and corresponding results were provided by the Challenge organizers. RESULTS: The results indicate that robust rule-based tools achieve satisfactory results in extraction of simple elements of medication orders, but more sophisticated methods are needed for identification of reasons for the orders and durations. LIMITATIONS: Owing to the time constraints and nature of the Challenge, some obvious follow-on analysis has not been completed yet. CONCLUSIONS: The authors plan to integrate the new modules with MetaMap to enhance its accuracy. This integration effort will provide guidance in retargeting existing tools for better processing of clinical text.
James G. Mork, Olivier Bodenreider, Dina Demner-Fushman, Rezarta Islamaj Dogan, François-Michel Lang, Zhiyong Lu, Aurélie Névéol, Lee B. Peters, Sonya E. Shooshan, Alan R. Aronson
J. Am. Medical Informatics Assoc.7
2009 Multi-terminology indexing for the assignment of MeSH descriptors to medical abstracts in French
Suzanne Pereira, Saoussen Sakji, Aurélie Névéol, Ivan Kergourlay, Gaétan Kerdelhué, Elisabeth Serrot Damatte, Michel Joubert, Stéfan Jacques Darmoni
AMIA3
2009 Comment on 'MeSH-up: effective MeSH text classification for improved document retrieval'
abstract
Abstract Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Aurélie Névéol, James G. Mork, Alan R. Aronson
Bioinform.1
2009 Comparing a rule-based versus statistical system for automatic categorization of MEDLINE documents according to biomedical specialty
abstract
Automatic document categorization is an important research problem in Information Science and Natural Language Processing. Many applications, including Word Sense Disambiguation and Information Retrieval in large collections, can benefit from such categorization. This paper focuses on automatic categorization of documents from the biomedical literature into broad discipline-based categories. Two different systems are described and contrasted: CISMeF, which uses rules based on human indexing of the documents by the Medical Subject Headings(®) (MeSH(®)) controlled vocabulary in order to assign metaterms (MTs), and Journal Descriptor Indexing (JDI) based on human categorization of about 4,000 journals and statistical associations between journal descriptors (JDs) and textwords in the documents. We evaluate and compare the performance of these systems against a gold standard of humanly assigned categories for one hundred MEDLINE documents, using six measures selected from trec_eval. The results show that for five of the measures, performance is comparable, and for one measure, JDI is superior. We conclude that these results favor JDI, given the significantly greater intellectual overhead involved in human indexing and maintaining a rule base for mapping MeSH terms to MTs. We also note a JDI method that associates JDs with MeSH indexing rather than textwords, and it may be worthwhile to investigate whether this JDI method (statistical) and CISMeF (rule based) might be combined and then evaluated showing they are complementary to one another.
Susanne M. Humphrey, Aurélie Névéol, Allen C. Browne, Julien Gobeill, Patrick Ruch, Stéfan Jacques Darmoni
J. Assoc. Inf. Sci. Technol.2
2009 Natural language processing versus content-based image analysis for medical document retrieval
abstract
One of the most significant recent advances in health information systems has been the shift from paper to electronic documents. While research on automatic text and image processing has taken separate paths, there is a growing need for joint efforts, particularly for electronic health records and biomedical literature databases. This work aims at comparing text-based versus image-based access to multimodal medical documents using state-of-the-art methods of processing text and image components. A collection of 180 medical documents containing an image accompanied by a short text describing it was divided into training and test sets. Content-based image analysis and natural language processing techniques are applied individually and combined for multimodal document analysis. The evaluation consists of an indexing task and a retrieval task based on the "gold standard" codes manually assigned to corpus documents. The performance of text-based and image-based access, as well as combined document features, is compared. Image analysis proves more adequate for both the indexing and retrieval of the images. In the indexing task, multimodal analysis outperforms both independent image and text analysis. This experiment shows that text describing images can be usefully analyzed in the framework of a hybrid text/image retrieval system.
Aurélie Névéol, Thomas M. Deserno, Stéfan Jacques Darmoni, Mark Oliver Güld, Alan R. Aronson
J. Assoc. Inf. Sci. Technol.1
2009 A recent advance in the automatic indexing of the biomedical literature
Aurélie Névéol, Sonya E. Shooshan, Susanne M. Humphrey, James G. Mork, Alan R. Aronson
J. Biomed. Informatics1
2008 Methodology for Creating UMLS Content Views Appropriate for Biomedical Natural Language Processing
Alan R. Aronson, James G. Mork, Aurélie Névéol, Sonya E. Shooshan, Dina Demner-Fushman
AMIA3
2008 Using multi-terminology indexing for the assignment of MeSH descriptors to health resources in a French online catalogue
Suzanne Pereira, Aurélie Névéol, Gaétan Kerdelhué, Elisabeth Serrot Damatte, Michel Joubert, Stéfan Jacques Darmoni
AMIA2
2008 Automatic inference of indexing rules for MEDLINE
abstract
BACKGROUND: Indexing is a crucial step in any information retrieval system. In MEDLINE, a widely used database of the biomedical literature, the indexing process involves the selection of Medical Subject Headings in order to describe the subject matter of articles. The need for automatic tools to assist MEDLINE indexers in this task is growing with the increasing number of publications being added to MEDLINE. METHODS: In this paper, we describe the use and the customization of Inductive Logic Programming (ILP) to infer indexing rules that may be used to produce automatic indexing recommendations for MEDLINE indexers. RESULTS: Our results show that this original ILP-based approach outperforms manual rules when they exist. In addition, the use of ILP rules also improves the overall performance of the Medical Text Indexer (MTI), a system producing automatic indexing recommendations for MEDLINE. CONCLUSION: We expect the sets of ILP rules obtained in this experiment to be integrated into MTI.
Aurélie Névéol, Sonya E. Shooshan, Vincent Claveau
BMC Bioinform.1
2007 Fine-Grained Indexing of the Biomedical Literature: MeSH Subheading Attachment for a MEDLINE Indexing Tool
Aurélie Névéol, Sonya E. Shooshan, James G. Mork, Alan R. Aronson
AMIA1
2006 Besides Precision & Recall: Exploring Alternative Approaches to Evaluating an Automatic Indexing Tool for MEDLINE
Aurélie Névéol, Kelly Zeng, Olivier Bodenreider
AMIA1
2006 Automatic indexing of online health resources for a French quality controlled gateway
Aurélie Névéol, Alexandrina Rogozan, Stéfan Jacques Darmoni
Inf. Process. Manag.1
2005 A Benchmark Evaluation of the French MeSH Indexers
Aurélie Névéol, Vincent Mary, Arnaud Gaudinat, Célia Boyer, Alexandrina Rogozan, Stéfan Jacques Darmoni
AIME1
2005 Evaluation of French and English MeSH Indexing Systems with a Parallel Corpus
Aurélie Névéol, James G. Mork, Alan R. Aronson, Stéfan Jacques Darmoni
AMIA1
2003 Text Categorization pror to Indexing for the CISMEF Health Catalogue
Alexandrina Rogozan, Aurélie Névéol, Stéfan Jacques Darmoni
AIME2