Adrien Coulet

dblp:82/4513 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-1466-062XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Privacy-Preserving Generation of Synthetic Pathology Reports for Information Extraction
Alejandra Lorenzo, Adrien Coulet, Claire Gardent
AIME (1)2
2026 Analysing Lightweight Large Language Models for Biomedical Named Entity Recognition on Diverse Ouput Formats
Pierre Epron, Adrien Coulet, Mehwish Alam
LREC2
2025 Predicting Clinical Outcomes from Patient Care Pathways Represented with Temporal Knowledge Graphs
Jong Ho Jhee, Alberto Megina, Pacôme Constant dit Beaufils, Matilde Karakachoff, Richard Redon, Alban Gaignard, Adrien Coulet
ESWC (1)7
2025 Efficient extraction of medication information from clinical notes: an evaluation in 2 languages
abstract
OBJECTIVE: To evaluate the accuracy, computational cost, and portability of a new natural language processing (NLP) method for extracting medication information from clinical narratives. MATERIALS AND METHODS: We propose an original transformer-based architecture for the extraction of entities and their relations pertaining to patients' medication regimen. First, we used this approach to train and evaluate a model on French clinical notes, using a newly annotated corpus from Hôpitaux Universitaires de Strasbourg. Second, the portability of the approach was assessed by conducting an evaluation on clinical documents in English from the 2018 n2c2 shared task. Information extraction accuracy and computational cost were assessed by comparison with an available method using transformers. RESULTS: The proposed architecture achieves on the task of relation extraction itself performance that are competitive with the state-of-the-art on both French and English (F-measures 0.82 and 0.96 vs 0.81 and 0.95), but reduces the computational cost by 10. End-to-end (Named Entity recognition and Relation Extraction) F1 performance is 0.69 and 0.82 for French and English corpus. DISCUSSION: While an existing system developed for English notes was deployed in a French hospital setting with reasonable effort, we found that an alternative architecture offered end-to-end drug information extraction with comparable extraction performance and lower computational impact for both French and English clinical text processing, respectively. CONCLUSION: The proposed architecture can be used to extract medication information from clinical text with high performance and low computational cost and consequently suits with usually limited hospital IT resources.
Thibaut Fabacher, Erik-André Sauleau, Emmanuelle Arcay, Bineta Faye, Maxime Alter, Archia Chahard, Nathan Miraillet, Adrien Coulet, Aurélie Névéol
J. Am. Medical Informatics Assoc.8
2025 Enhancing clinical data warehousing with provenance data to support longitudinal analyses and large file management: The gitOmmix approach for genomic and image data
abstract
BACKGROUND: If hospital Clinical Data Warehouses are to address today's focus in personalized medicine, they need to be able to track patients longitudinally and manage the large data sets generated by whole genome sequencing, RNA analyses, and complex imaging studies. Current Clinical Data Warehouses address neither issue. This paper reports on methods to enrich current systems by providing provenance data allowing patient histories to be followed longitudinally and managing the linking and versioning of large data sets from whatever source. The methods are open source and applicable to any clinical data warehouse system, whether data schema it uses. METHOD: We introduce gitOmmix, an approach that overcomes these limitations, and illustrate its usefulness in the management of medical omics data. gitOmmix relies on (i) a file versioning system: git, (ii) an extension that handles large files: git-annex, (iii) a provenance knowledge graph: PROV-O, and (iv) an alignment between the git versioning information and the provenance knowledge graph. RESULTS: Capabilities inherited from git and git-annex enable retracing the history of a clinical interpretation back to the patient sample, through supporting data and analyses. In addition, the provenance knowledge graph, aligned with the git versioning information, enables querying and browsing provenance relationships between these elements. CONCLUSION: gitOmmix adds a provenance layer to CDWs, while scaling to large files and being agnostic of the CDW system. For these reasons, we think that it is a viable and generalizable solution for omics clinical studies.
Maxime Wack, Adrien Coulet, Anita Burgun-Parenthoine, Bastien Rance
J. Biomed. Informatics2
2024 Deep Reinforcement Learning for personalized diagnostic decision pathways using Electronic Health Records: A comparative study on anemia and Systemic Lupus Erythematosus
abstract
BACKGROUND: Clinical diagnoses are typically made by following a series of steps recommended by guidelines that are authored by colleges of experts. Accordingly, guidelines play a crucial role in rationalizing clinical decisions. However, they suffer from limitations, as they are designed to cover the majority of the population and often fail to account for patients with uncommon conditions. Moreover, their updates are long and expensive, making them unsuitable for emerging diseases and new medical practices. METHODS: Inspired by guidelines, we formulate the task of diagnosis as a sequential decision-making problem and study the use of Deep Reinforcement Learning (DRL) algorithms to learn the optimal sequence of actions to perform in order to obtain a correct diagnosis from Electronic Health Records (EHRs), which we name a diagnostic decision pathway. We apply DRL to synthetic yet realistic EHRs and develop two clinical use cases: Anemia diagnosis, where the decision pathways follow a decision tree schema, and Systemic Lupus Erythematosus (SLE) diagnosis, which follows a weighted criteria score. We particularly evaluate the robustness of our approaches to noise and missing data, as these frequently occur in EHRs. RESULTS: In both use cases, even with imperfect data, our best DRL algorithms exhibit competitive performance compared to traditional classifiers, with the added advantage of progressively generating a pathway to the suggested diagnosis, which can both guide and explain the decision-making process. CONCLUSION: DRL offers the opportunity to learn personalized decision pathways for diagnosis. Our two use cases illustrate the advantages of this approach: they generate step-by-step pathways that are explainable, and their performance is competitive when compared to state-of-the-art methods.
Lillian Muyama, Antoine Neuraz, Adrien Coulet
Artif. Intell. Medicine3
2024 Facilitating phenotyping from clinical texts: the medkit library
abstract
SUMMARY: Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical information of EHRs are lying in texts, phenotyping from text takes an important role in studies that rely on the secondary use of EHRs. However, the heterogeneity and highly specialized aspect of both the content and form of clinical texts makes this task particularly tedious, and is the source of time and cost constraints in observational studies. To facilitate the development, evaluation and reproducibility of phenotyping pipelines, we developed an open-source Python library named medkit. It enables composing data processing pipelines made of easy-to-reuse software bricks, named medkit operations. In addition to the core of the library, we share the operations and pipelines we already developed and invite the phenotyping community for their reuse and enrichment. AVAILABILITY AND IMPLEMENTATION: medkit is available at https://github.com/medkit-lib/medkit.
Antoine Neuraz, Ghislain Vaillant, Camila Arias, Olivier Birot, Kim Tâm Huynh, Thibaut Fabacher, Alice Rogier, Nicolas Garcelon, Ivan Lerner, Bastien Rance, Adrien Coulet
Bioinform.11
2024 Machine learning approaches for the discovery of clinical pathways from patient data: A systematic review
abstract
BACKGROUND: Clinical pathways are sequences of events followed during the clinical care of a group of patients who meet pre-defined criteria. They have many applications ranging from healthcare evaluation and optimization to clinical decision support. These pathways can be discovered from existing healthcare data, in particular with machine learning which is a family of methods used to learn patterns from data. This review provides a comprehensive overview of the literature concerning the use of machine learning methods for clinical pathway discovery from patient data. METHODS: Guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) method , we conducted a systematic review of the existing literature. We searched 6 databases, i.e., ACM Digital Library, ScienceDirect, Web of Science, PubMed, IEEE Xplore, and Scopus spanning from January 2004 to December 2023 using search terms pertinent to clinical pathways and their development. Subsequently, the retrieved papers were analyzed to assess their relevance to the scope of this study. RESULTS: In total, 131 papers that met the specified inclusion criteria were identified. These papers expressed diverse motivations behind data-driven clinical pathway discovery ranging from knowledge discovery to conformance checking with established clinical guidelines (derived from existing literature and clinical experts). Notably, the predominant methods employed (67.2%, n=88) involved unsupervised machine learning techniques, such as clustering and process mining. CONCLUSIONS: Relevant clinical pathways can be discovered from patient data using machine learning methods, with the desirable potential to aid clinical decision-making in healthcare. However, to reach this objective, the methods used to discover pathways should be reproducible, and rigorous performance evaluation by clinical experts needs to be conducted for validation.
Lillian Muyama, Antoine Neuraz, Adrien Coulet
J. Biomed. Informatics3
2019 Creating Fair Models of Atherosclerotic Cardiovascular Disease Risk
abstract
Guidelines for the management of atherosclerotic cardiovascular disease (ASCVD) recommend the use of risk stratification models to identify patients most likely to benefit from cholesterol-lowering and other therapies. These models have differential performance across race and gender groups with inconsistent behavior across studies, potentially resulting in an inequitable distribution of beneficial therapy. In this work, we leverage adversarial learning and a large observational cohort extracted from electronic health records (EHRs) to develop a "fair" ASCVD risk prediction model with reduced variability in error rates across groups. We empirically demonstrate that our approach is capable of aligning the distribution of risk predictions conditioned on the outcome across several groups simultaneously for models built from high-dimensional EHR data. We also discuss the relevance of these results in the context of the empirical trade-off between fairness and model performance.
Stephen Pfohl, Ben J. Marafino, Adrien Coulet, Fátima Rodriguez, Latha Palaniappan, Nigam H. Shah
AIES3
2019 PGxO and PGxLOD: a reconciliation of pharmacogenomic knowledge of various provenances, enabling further comparison
abstract
BACKGROUND: Pharmacogenomics (PGx) studies how genomic variations impact variations in drug response phenotypes. Knowledge in pharmacogenomics is typically composed of units that have the form of ternary relationships gene variant - drug - adverse event. Such a relationship states that an adverse event may occur for patients having the specified gene variant and being exposed to the specified drug. State-of-the-art knowledge in PGx is mainly available in reference databases such as PharmGKB and reported in scientific biomedical literature. But, PGx knowledge can also be discovered from clinical data, such as Electronic Health Records (EHRs), and in this case, may either correspond to new knowledge or confirm state-of-the-art knowledge that lacks "clinical counterpart" or validation. For this reason, there is a need for automatic comparison of knowledge units from distinct sources. RESULTS: In this article, we propose an approach, based on Semantic Web technologies, to represent and compare PGx knowledge units. To this end, we developed PGxO, a simple ontology that represents PGx knowledge units and their components. Combined with PROV-O, an ontology developed by the W3C to represent provenance information, PGxO enables encoding and associating provenance information to PGx relationships. Additionally, we introduce a set of rules to reconcile PGx knowledge, i.e. to identify when two relationships, potentially expressed using different vocabularies and levels of granularity, refer to the same, or to different knowledge units. We evaluated our ontology and rules by populating PGxO with knowledge units extracted from PharmGKB (2701), the literature (65,720) and from discoveries reported in EHR analysis studies (only 10, manually extracted); and by testing their similarity. We called PGxLOD (PGx Linked Open Data) the resulting knowledge base that represents and reconciles knowledge units of those various origins. CONCLUSIONS: The proposed ontology and reconciliation rules constitute a first step toward a more complete framework for knowledge comparison in PGx. In this direction, the experimental instantiation of PGxO, named PGxLOD, illustrates the ability and difficulties of reconciling various existing knowledge sources.
Pierre Monnin, Joël Legrand, Graziella Husson, Patrice Ringot, Andon Tchechmedjiev, Clément Jonquet, Amedeo Napoli, Adrien Coulet
BMC Bioinform.8
2017 Using Formal Concept Analysis for Checking the Structure of an Ontology in LOD: The Example of DBpedia
Pierre Monnin, Mario Lezoche, Amedeo Napoli, Adrien Coulet
ISMIS4
2013 Using Pattern Structures for Analyzing Ontology-Based Annotations of Biomedical Data
Adrien Coulet, Florent Domenach, Mehdi Kaytoue-Uberall, Amedeo Napoli
ICFCA1
2012 Using ontology-based annotation to profile disease research
abstract
BACKGROUND: Profiling the allocation and trend of research activity is of interest to funding agencies, administrators, and researchers. However, the lack of a common classification system hinders the comprehensive and systematic profiling of research activities. This study introduces ontology-based annotation as a method to overcome this difficulty. Analyzing over a decade of funding data and publication data, the trends of disease research are profiled across topics, across institutions, and over time. RESULTS: This study introduces and explores the notions of research sponsorship and allocation and shows that leaders of research activity can be identified within specific disease areas of interest, such as those with high mortality or high sponsorship. The funding profiles of disease topics readily cluster themselves in agreement with the ontology hierarchy and closely mirror the funding agency priorities. Finally, four temporal trends are identified among research topics. CONCLUSIONS: This work utilizes disease ontology (DO)-based annotation to profile effectively the landscape of biomedical research activity. By using DO in this manner a use-case driven mechanism is also proposed to evaluate the utility of classification hierarchies.
Adrien Coulet, Paea LePendu, Nigam H. Shah
J. Am. Medical Informatics Assoc.2
2012 The state of the art in text mining and natural language processing for pharmacogenomics
Adrien Coulet, Kevin Cohen 0001, Russ B. Altman
J. Biomed. Informatics1
2011 NCBO Resource Index: Ontology-based search and mining of biomedical resources
Clément Jonquet, Paea LePendu, Sean M. Falconer, Adrien Coulet, Natasha F. Noy, Mark A. Musen, Nigam H. Shah
J. Web Semant.4
2010 Using text to build semantic networks for pharmacogenomics
Adrien Coulet, Nigam H. Shah, Yael Garten, Mark A. Musen, Russ B. Altman
J. Biomed. Informatics1
2008 Ontology-guided data preparation for discovering genotype-phenotype relationships
abstract
BACKGROUND: Complexity and amount of post-genomic data constitute two major factors limiting the application of Knowledge Discovery in Databases (KDD) methods in life sciences. Bio-ontologies may nowadays play key roles in knowledge discovery in life science providing semantics to data and to extracted units, by taking advantage of the progress of Semantic Web technologies concerning the understanding and availability of tools for knowledge representation, extraction, and reasoning. RESULTS: This paper presents a method that exploits bio-ontologies for guiding data selection within the preparation step of the KDD process. We propose three scenarios in which domain knowledge and ontology elements such as subsumption, properties, class descriptions, are taken into account for data selection, before the data mining step. Each of these scenarios is illustrated within a case-study relative to the search of genotype-phenotype relationships in a familial hypercholesterolemia dataset. The guiding of data selection based on domain knowledge is analysed and shows a direct influence on the volume and significance of the data mining results. CONCLUSIONS: The method proposed in this paper is an efficient alternative to numerical methods for data selection based on domain knowledge. In turn, the results of this study may be reused in ontology modelling and data integration.
Adrien Coulet, Malika Smaïl-Tabbone, Pascale Benlian, Amedeo Napoli, Marie-Dominique Devignes
BMC Bioinform.1