EDBT 2026 Demo / reviewers in the wild / expert
Bastien Rance
dblp:86/722
· DBLP profile ↗
17ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-4417-1197ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing clinical data warehousing with provenance data to support longitudinal analyses and large file management: The gitOmmix approach for genomic and image dataabstractBACKGROUND: If hospital Clinical Data Warehouses are to address today's focus in personalized medicine, they need to be able to track patients longitudinally and manage the large data sets generated by whole genome sequencing, RNA analyses, and complex imaging studies. Current Clinical Data Warehouses address neither issue. This paper reports on methods to enrich current systems by providing provenance data allowing patient histories to be followed longitudinally and managing the linking and versioning of large data sets from whatever source. The methods are open source and applicable to any clinical data warehouse system, whether data schema it uses. METHOD: We introduce gitOmmix, an approach that overcomes these limitations, and illustrate its usefulness in the management of medical omics data. gitOmmix relies on (i) a file versioning system: git, (ii) an extension that handles large files: git-annex, (iii) a provenance knowledge graph: PROV-O, and (iv) an alignment between the git versioning information and the provenance knowledge graph. RESULTS: Capabilities inherited from git and git-annex enable retracing the history of a clinical interpretation back to the patient sample, through supporting data and analyses. In addition, the provenance knowledge graph, aligned with the git versioning information, enables querying and browsing provenance relationships between these elements. CONCLUSION: gitOmmix adds a provenance layer to CDWs, while scaling to large files and being agnostic of the CDW system. For these reasons, we think that it is a viable and generalizable solution for omics clinical studies. Maxime Wack, Adrien Coulet, Anita Burgun-Parenthoine, Bastien Rance |
J. Biomed. Informatics | 4 |
| 2024 | Facilitating phenotyping from clinical texts: the medkit libraryabstractSUMMARY: Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical information of EHRs are lying in texts, phenotyping from text takes an important role in studies that rely on the secondary use of EHRs. However, the heterogeneity and highly specialized aspect of both the content and form of clinical texts makes this task particularly tedious, and is the source of time and cost constraints in observational studies. To facilitate the development, evaluation and reproducibility of phenotyping pipelines, we developed an open-source Python library named medkit. It enables composing data processing pipelines made of easy-to-reuse software bricks, named medkit operations. In addition to the core of the library, we share the operations and pipelines we already developed and invite the phenotyping community for their reuse and enrichment. AVAILABILITY AND IMPLEMENTATION: medkit is available at https://github.com/medkit-lib/medkit. Antoine Neuraz, Ghislain Vaillant, Camila Arias, Olivier Birot, Kim Tâm Huynh, Thibaut Fabacher, Alice Rogier, Nicolas Garcelon, Ivan Lerner, Bastien Rance, Adrien Coulet |
Bioinform. | 10 |
| 2022 | How to improve cancer Patients ENrollment within clinical trials from rEal Life databases using the OMOP oncology Extension: the French PENELOPE initiative
Emmanuelle Kempf, Morgan Vaterkowski, Nicolas Griffon, Damien Leprovost, Stéphane Bréant, Patricia Serre, Alexandre Mouchet, Rafael Gozlan, Bastien Rance, Gilles Chatellier, Ali Bellamine, Marie Frank, Martin Hilka, Julien Guerin, Xavier Tannier, Alain Livartowski, Christel Daniel-Le Bozec |
AMIA | 9 |
| 2022 | Privacy-preserving mimic models for clinical named entity recognition in French
Nesrine Bannour, Perceval Wajsbürt, Bastien Rance, Xavier Tannier, Aurélie Névéol |
J. Biomed. Informatics | 3 |
| 2021 | Can reproducibility be improved in clinical natural language processing? A study of 7 clinical NLP suites
William Digan, Aurélie Névéol, Antoine Neuraz, Maxime Wack, David Baudoin, Anita Burgun-Parenthoine, Bastien Rance |
AMIA | 7 |
| 2021 | Can reproducibility be improved in clinical natural language processing? A study of 7 clinical NLP suitesabstractBACKGROUND: The increasing complexity of data streams and computational processes in modern clinical health information systems makes reproducibility challenging. Clinical natural language processing (NLP) pipelines are routinely leveraged for the secondary use of data. Workflow management systems (WMS) have been widely used in bioinformatics to handle the reproducibility bottleneck. OBJECTIVE: To evaluate if WMS and other bioinformatics practices could impact the reproducibility of clinical NLP frameworks. MATERIALS AND METHODS: Based on the literature across multiple researcho fields (NLP, bioinformatics and clinical informatics) we selected articles which (1) review reproducibility practices and (2) highlight a set of rules or guidelines to ensure tool or pipeline reproducibility. We aggregate insight from the literature to define reproducibility recommendations. Finally, we assess the compliance of 7 NLP frameworks to the recommendations. RESULTS: We identified 40 reproducibility features from 8 selected articles. Frameworks based on WMS match more than 50% of features (26 features for LAPPS Grid, 22 features for OpenMinted) compared to 18 features for current clinical NLP framework (cTakes, CLAMP) and 17 features for GATE, ScispaCy, and Textflows. DISCUSSION: 34 recommendations are endorsed by at least 2 articles from our selection. Overall, 15 features were adopted by every NLP Framework. Nevertheless, frameworks based on WMS had a better compliance with the features. CONCLUSION: NLP frameworks could benefit from lessons learned from the bioinformatics field (eg, public repositories of curated tools and workflows or use of containers for shareability) to enhance the reproducibility in a clinical setting. William Digan, Aurélie Névéol, Antoine Neuraz, Maxime Wack, David Baudoin, Anita Burgun-Parenthoine, Bastien Rance |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | Designing Scientific SPARQL Queries Using Autocompletion by SnippetsabstractSPARQL is the standard query language used to access RDF linked data sets available on the Web. However, designing a SPARQL query can be a tedious task, even for experienced users. This is often due to imperfect knowledge by the user of the ontologies involved in the query. To overcome this problem, a growing number of query editors offer autocompletetion features. Such features are nevertheless limited and mostly focused on typo checking. In this context, our contribution is four-fold. First, we analyze several autocompletion features proposed by the main editors, highlighting the needs currently not taken into account while met by a user community we work with, scientists. Second, we introduce the first (to our knowledge) autocompletion approach able to consider snippets (fragments of SPARQL query) based on queries expressed by previous users, enriching the user experience. Third, we introduce a usable, open and concrete solution able to consider a large panel of SPARQL autocompletion features that we have implemented in an editor. Last but not least, we demonstrate the interest of our approach on real biomedical queries involving services offered by the Wikidata collaborative knowledge base. Karima Rafes, Serge Abiteboul, Sarah Cohen Boulakia, Bastien Rance |
eScience | 4 |
| 2018 | A clinician friendly data warehouse oriented toward narrative reports: Dr. WarehouseabstractINTRODUCTION: Clinical data warehouses are often oriented toward integration and exploration of coded data. However narrative reports are of crucial importance for translational research. This paper describes Dr. Warehouse®, an open source data warehouse oriented toward clinical narrative reports and designed to support clinicians' day-to-day use. METHOD: Dr. Warehouse relies on an original database model to focus on documents in addition to facts. Besides classical querying functionalities, the system provides an advanced search engine and Graphical User Interfaces adapted to the exploration of text. Dr. Warehouse is dedicated to translational research with cohort recruitment capabilities, high throughput phenotyping and patient centric views (including similarity metrics among patients). These features leverage Natural Language Processing based on the extraction of UMLS® concepts, as well as negation and family history detection. RESULTS: A survey conducted after 6 months of use at the Necker Children's Hospital shows a high rate of satisfaction among the users (96.6%). During this period, 122 users performed 2837 queries, accessed 4,267 patients' records and included 36,632 patients in 131 cohorts. The source code is available at this github link https://github.com/imagine-bdd/DRWH. A demonstration based on PubMed abstracts is available at https://imagine-plateforme-bdd.fr/dwh_pubmed/. Nicolas Garcelon, Antoine Neuraz, Rémi Salomon, Hassan Faour, Vincent Benoit, Arthur Delapalme, Arnold Munnich, Anita Burgun-Parenthoine, Bastien Rance |
J. Biomed. Informatics | 9 |
| 2017 | Exploring and visualizing multidimensional data in translational research platformsabstractThe unprecedented advances in technology and scientific research over the past few years have provided the scientific community with new and more complex forms of data. Large data sets collected from single groups or cross-institution consortiums containing hundreds of omic and clinical variables corresponding to thousands of patients are becoming increasingly commonplace in the research setting. Before any core analyses are performed, visualization often plays a key role in the initial phases of research, especially for projects where no initial hypotheses are dominant. Proper visualization of data at a high level facilitates researcher's abilities to find trends, identify outliers and perform quality checks. In addition, research has uncovered the important role of visualization in data analysis and its implied benefits facilitating our understanding of disease and ultimately improving patient care. In this work, we present a review of the current landscape of existing tools designed to facilitate the visualization of multidimensional data in translational research platforms. Specifically, we reviewed the biomedical literature for translational platforms allowing the visualization and exploration of clinical and omics data, and identified 11 platforms: cBioPortal, interactive genomics patient stratification explorer, Igloo-Plot, The Georgetown Database of Cancer Plus, tranSMART, an unnamed data-cube-based model supporting heterogeneous data, Papilio, Caleydo Domino, Qlucore Omics, Oracle Health Sciences Translational Research Center and OmicsOffice® powered by TIBCO Spotfire. In a health sector continuously witnessing an increase in data from multifarious sources, visualization tools used to better grasp these data will grow in their importance, and we believe our work will be useful in guiding investigators in similar situations. William Dunn Jr, Anita Burgun-Parenthoine, Marie-Odile Krebs, Bastien Rance |
Briefings Bioinform. | 4 |
| 2015 | Reviewing 741 patients records in two hours with FASTVISU
Jean-Baptiste Escudié, Anne-Sophie Jannot, Eric Zapletal, Sarah Cohen Boulakia, Georgia Malamut, Anita Burgun-Parenthoine, Bastien Rance |
AMIA | 7 |
| 2015 | Translational research platforms integrating clinical and omics data: a review of publicly available solutionsabstractThe rise of personalized medicine and the availability of high-throughput molecular analyses in the context of clinical care have increased the need for adequate tools for translational researchers to manage and explore these data. We reviewed the biomedical literature for translational platforms allowing the management and exploration of clinical and omics data, and identified seven platforms: BRISK, caTRIP, cBio Cancer Portal, G-DOC, iCOD, iDASH and tranSMART. We analyzed these platforms along seven major axes. (1) The community axis regrouped information regarding initiators and funders of the project, as well as availability status and references. (2) We regrouped under the information content axis the nature of the clinical and omics data handled by each system. (3) The privacy management environment axis encompassed functionalities allowing control over data privacy. (4) In the analysis support axis, we detailed the analytical and statistical tools provided by the platforms. We also explored (5) interoperability support and (6) system requirements. The final axis (7) platform support listed the availability of documentation and installation procedures. A large heterogeneity was observed in regard to the capability to manage phenotype information in addition to omics data, their security and interoperability features. The analytical and visualization features strongly depend on the considered platform. Similarly, the availability of the systems is variable. This review aims at providing the reader with the background to choose the platform best suited to their needs. To conclude, we discuss the desiderata for optimal translational research platforms, in terms of privacy, interoperability and technical features. Vincent Canuel, Bastien Rance, Paul Avillach, Patrice Degoulet, Anita Burgun-Parenthoine |
Briefings Bioinform. | 2 |
| 2014 | How to de-identify a large clinical corpus in 10 days
Cyril Grouin, Louise Deléger, Jean-Baptiste Escudié, Gregory Groisy, Anne-Sophie Jannot, Bastien Rance, Xavier Tannier, Aurélie Névéol |
AMIA | 6 |
| 2013 | Network Visualization of UMLS Source Vocabularies using Semantic Groups
Thai Le, Bastien Rance, Olivier Bodenreider |
AMIA | 2 |
| 2013 | A systematic comparison of current sources of disease knowledge
Aurélie Névéol, Bastien Rance, Zhiyong Lu |
AMIA | 2 |
| 2012 | A mutation-centric approach to identifying pharmacogenomic relations in text
Bastien Rance, Emily Doughty, Dina Demner-Fushman, Maricel G. Kann, Olivier Bodenreider |
J. Biomed. Informatics | 1 |
| 2011 | Integrating clinical research with the Healthcare Enterprise: From the RE-USE project to the EHR4CR platformabstractBACKGROUND: There are different approaches for repurposing clinical data collected in the Electronic Healthcare Record (EHR) for use in clinical research. Semantic integration of "siloed" applications across domain boundaries is the raison d'être of the standards-based profiles developed by the Integrating the Healthcare Enterprise (IHE) initiative - an initiative by healthcare professionals and industry promoting the coordinated use of established standards such as DICOM and HL7 to address specific clinical needs in support of optimal patient care. In particular, the combination of two IHE profiles - the integration profile "Retrieve Form for Data Capture" (RFD), and the IHE content profile "Clinical Research Document" (CRD) - offers a straightforward approach to repurposing EHR data by enabling the pre-population of the case report forms (eCRF) used for clinical research data capture by Clinical Data Management Systems (CDMS) with previously collected EHR data. OBJECTIVE: Implement an alternative solution of the RFD-CRD integration profile centered around two approaches: (i) Use of the EHR as the single-source data-entry and persistence point in order to ensure that all the clinical data for a given patient could be found in a single source irrespective of the data collection context, i.e. patient care or clinical research; and (ii) Maximize the automatic pre-population process through the use of a semantic interoperability services that identify duplicate or semantically-equivalent eCRF/EHR data elements as they were collected in the EHR context. METHODS: The RE-USE architecture and associated profiles are focused on defining a set of scalable, standards-based, IHE-compliant profiles that can enable single-source data collection/entry and cross-system data reuse through semantic integration. Specifically, data reuse is realized through the semantic mapping of data collection fields in electronic Case Report Forms (eCRFs) to data elements previously defined as part of patient care-centric templates in the EHR context. The approach was evaluated in the context of a multi-center clinical trial conducted in a large, multi-disciplinary hospital with an installed EHR. RESULTS: Data elements of seven eCRFs used in a multi-center clinical trial were mapped to data elements of patient care-centric templates in use in the EHR at the George Pompidou hospital. 13.4% of the data elements of the eCRFs were found to be represented in EHR templates and were therefore candidate for pre-population. During the execution phase of the clinical study, the semantic mapping architecture enabled data persisted in the EHR context as part of clinical care to be used to pre-populate eCRFS for use without secondary data entry. To ensure that the pre-populated data is viable for use in the clinical research context, all pre-populated eCRF data needs to be first approved by a trial investigator prior to being persisted in a research data store within a CDMS. CONCLUSION: Single-source data entry in the clinical care context for use in the clinical research context - a process enabled through the use of the EHR as single point of data entry, can - if demonstrated to be a viable strategy - not only significantly reduce data collection efforts while simultaneously increasing data collection accuracy secondary to elimination of transcription or double-entry errors between the two contexts but also ensure that all the clinical data for a given patient, irrespective of the data collection context, are available in the EHR for decision support and treatment planning. The RE-USE approach used mapping algorithms to identify semantic coherence between clinical care and clinical research data elements and pre-populate eCRFs. The RE-USE project utilized SNOMED International v.3.5 as its "pivot reference terminology" to support EHR-to-eCRF mapping, a decision that likely enhanced the "recall" of the mapping algorithms. The RE-USE results demonstrate the difficult challenges involved in semantic integration between the clinical care and clinical research contexts. Abdennaji El Fadly, Bastien Rance, Noël Lucas, Charles N. Mead, Gilles Chatellier, Pierre-Yves Lastic, Marie-Christine Jaulent, Christel Daniel-Le Bozec |
J. Biomed. Informatics | 2 |
| 2007 | Extracting Sequential Nuggets of Knowledge
Christine Froidevaux, Frédérique Lisacek, Bastien Rance |
DEXA | 3 |