Eric Prud'hommeaux

dblp:10/5324 · also Eric G. Prud'hommeaux · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-1775-9921ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Computer networks · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Shape Expressions with Inheritance
Iovka Boneva, José Emilio Labra Gayo, Eric Prud'hommeaux, Katherine Thornton, Andra Waagmeester
ESWC (1)3
2023 Shape Expressions (ShEx) schemas for the FHIR R5 specification
Deepak K. Sharma, Eric Prud'hommeaux, David Booth, Claude J. Nanjo, Guoqian Jiang
J. Biomed. Informatics2
2022 Maximizing Interoperability, Enriching EHR Data: Transforming HL7 FHIR Data to RDF Using the FHIR RDF Playground
James Champion, Eric Prud'hommeaux, David Booth, Gaurav Vaidya, James P. Balhoff, Deepak K. Sharma, Guoqian Jiang, Emily R. Pfaff
AMIA2
2022 Modeling a Cancer Symptom Control Domain Using HL7 FHIR: Applicability of the Minimal Common Oncology Data Elements (mCODE)
Nan Huo, Yue Yu 0012, Nansu Zong, Andrea Cheville, Claude J. Nanjo, Eric Prud'hommeaux, Deirdre Pachman, Guohui Xiao 0001, Emily R. Pfaff, Christopher G. Chute, Guoqian Jiang, Kathryn J. Ruddy
AMIA6
2022 FHIR-Ontop-OMOP: Building clinical knowledge graphs in FHIR RDF with the OMOP Common data Model
abstract
BACKGROUND: Knowledge graphs (KGs) play a key role to enable explainable artificial intelligence (AI) applications in healthcare. Constructing clinical knowledge graphs (CKGs) against heterogeneous electronic health records (EHRs) has been desired by the research and healthcare AI communities. From the standardization perspective, community-based standards such as the Fast Healthcare Interoperability Resources (FHIR) and the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) are increasingly used to represent and standardize EHR data for clinical data analytics, however, the potential of such a standard on building CKG has not been well investigated. OBJECTIVE: To develop and evaluate methods and tools that expose the OMOP CDM-based clinical data repositories into virtual clinical KGs that are compliant with FHIR Resource Description Framework (RDF) specification. METHODS: We developed a system called FHIR-Ontop-OMOP to generate virtual clinical KGs from the OMOP relational databases. We leveraged an OMOP CDM-based Medical Information Mart for Intensive Care (MIMIC-III) data repository to evaluate the FHIR-Ontop-OMOP system in terms of the faithfulness of data transformation and the conformance of the generated CKGs to the FHIR RDF specification. RESULTS: A beta version of the system has been released. A total of more than 100 data element mappings from 11 OMOP CDM clinical data, health system and vocabulary tables were implemented in the system, covering 11 FHIR resources. The generated virtual CKG from MIMIC-III contains 46,520 instances of FHIR Patient, 716,595 instances of Condition, 1,063,525 instances of Procedure, 24,934,751 instances of MedicationStatement, 365,181,104 instances of Observations, and 4,779,672 instances of CodeableConcept. Patient counts identified by five pairs of SQL (over the MIMIC database) and SPARQL (over the virtual CKG) queries were identical, ensuring the faithfulness of the data transformation. Generated CKG in RDF triples for 100 patients were fully conformant with the FHIR RDF specification. CONCLUSION: The FHIR-Ontop-OMOP system can expose OMOP database as a FHIR-compliant RDF graph. It provides a meaningful use case demonstrating the potentials that can be enabled by the interoperability between FHIR and OMOP CDM. Generated clinical KGs in FHIR RDF provide a semantic foundation to enable explainable AI applications in healthcare.
Guohui Xiao 0001, Emily R. Pfaff, Eric Prud'hommeaux, David Booth, Deepak K. Sharma, Nan Huo, Yue Yu 0012, Nansu Zong, Kathryn J. Ruddy, Christopher G. Chute, Guoqian Jiang
J. Biomed. Informatics3
2021 Development of a FHIR RDF data transformation and validation framework and its evaluation
Eric Prud'hommeaux, Josh Collins, David Booth, Kevin J. Peterson, Harold R. Solbrig, Guoqian Jiang
J. Biomed. Informatics1
2020 Exploring JSON-LD as an Executable Definition of FHIR RDF to Enable Semantics of FHIR Data
Harold R. Solbrig, Dazhi Jiao, Eric Prud'hommeaux, David Booth, Cory M. Endle, Daniel J. Stone, Guoqian Jiang
AMIA3
2019 Using Shape Expressions (ShEx) to Share RDF Data Models and to Guide Curation with Rigorous Validation
abstract
Abstract We discuss Shape Expressions (ShEx), a concise, formal, modeling and validation language for RDF structures. For instance, a Shape Expression could prescribe that subjects in a given RDF graph that fall into the shape “Paper” are expected to have a section called “Abstract”, and any ShEx implementation can confirm whether that is indeed the case for all such subjects within a given graph or subgraph. There are currently five actively maintained ShEx implementations. We discuss how we use the JavaScript, Scala and Python implementations in RDF data validation workflows in distinct, applied contexts. We present examples of how ShEx can be used to model and validate data from two different sources, the domain-specific Fast Healthcare Interoperability Resources (FHIR) and the domain-generic Wikidata knowledge base, which is the linked database built and maintained by the Wikimedia Foundation as a sister project to Wikipedia. Example projects that are using Wikidata as a data curation platform are presented as well, along with ways in which they are using ShEx for modeling and validation. When reusing RDF graphs created by others, it is important to know how the data is represented. Current practices of using human-readable descriptions or ontologies to communicate data structures often lack sufficient precision for data consumers to quickly and easily understand data representation details. We provide concrete examples of how we use ShEx as a constraint and validation language that allows humans and machines to communicate unambiguously about data assets. We use ShEx to exchange and understand data models of different origins, and to express a shared model of a resource’s footprint in a Linked Data source. We also use ShEx to agilely develop data models, test them against sample data, and revise or refine them. The expressivity of ShEx allows us to catch disagreement, inconsistencies, or errors efficiently, both at the time of input, and through batch inspections. ShEx addresses the need of the Semantic Web community to ensure data quality for RDF graphs. It is currently being used in the development of FHIR/RDF. The language is sufficiently expressive to capture constraints in FHIR, and the intuitive syntax helps people to quickly grasp the range of conformant documents. The publication workflow for FHIR tests all of these examples against the ShEx schemas, catching non-conformant data before they reach the public. ShEx is also currently used in Wikidata projects such as Gene Wiki and WikiCite to develop quality-control pipelines to maintain data integrity and incorporate or harmonize differences in data across different parts of the pipelines.
Katherine Thornton, Harold R. Solbrig, Gregory S. Stupp, José Emilio Labra Gayo, Daniel Mietchen, Eric Prud'hommeaux, Andra Waagmeester
ESWC6
2017 Building an FHIR Ontology based Data Access Framework with the OHDSI Data Repositories
Guoqian Jiang, Guohui Xiao 0001, Richard C. Kiefer, Eric Prud'hommeaux, Harold R. Solbrig
AMIA4
2017 Semantics and Validation of Shapes Schemas for RDF
Iovka Boneva, José Emilio Labra Gayo, Eric Prud'hommeaux
ISWC (1)3
2017 Modeling and validating HL7 FHIR profiles using semantic web Shape Expressions (ShEx)
Harold R. Solbrig, Eric Prud'hommeaux, Grahame Grieve, Lloyd McKenzie, Joshua C. Mandel, Deepak K. Sharma, Guoqian Jiang
J. Biomed. Informatics2
2016 Standardized Representation of Clinical Study Data Dictionaries with CIMI Archetypes
Deepak K. Sharma, Harold R. Solbrig, Eric Prud'hommeaux, Jyotishman Pathak, Guoqian Jiang
AMIA3
2016 Knowledge Representation on the Web Revisited: The Case for Prototypes
Michael Cochez, Stefan Decker, Eric Prud'hommeaux
ISWC (1)3
2015 Quality Assurance of Cancer Study Common Data Elements Using A Post-Coordination Approach
Guoqian Jiang, Harold R. Solbrig, Eric Prud'hommeaux, Cui Tao, Chunhua Weng, Christopher G. Chute
AMIA3
2015 Representing and Validating Cancer Study Metadata Standard Using RDF Shapes Expression Language
Harold R. Solbrig, Eric Prud'hommeaux, Christopher G. Chute, Guoqian Jiang
AMIA2
2015 Complexity and Expressiveness of ShEx for RDF
abstract
Graph data abstractions are often assumed to be intuitive, but experience shows that they are not equally understandable or usable in practice. In this vision and challenges paper, we examine the human-centricity of contemporary graph data abstractions through four lenses: researchability, usability, teachability, and societal impact. Drawing on diverse real-world use cases, ranging from clinical data and collaborative knowledge bases to biological and pangenomic graphs, we distill insights from database research, human-computer interaction, and education. Based on this analysis, we identify open research challenges that must be addressed to make graph abstractions easier to study, use, learn, and reason about.
Slawomir Staworko, Iovka Boneva, José Emilio Labra Gayo, Samuel Hym, Eric Prud'hommeaux, Harold R. Solbrig
ICDT5
2012 Translating standards into practice - One Semantic Web API for Gene Expression
Helena F. Deus, Eric Prud'hommeaux, Michael Miller 0001, Jun Zhao 0003, James Malone, Tomasz Adamusiak, Jamie P. McCusker, Sudeshna Das 0001, Philippe Rocca-Serra, Ronan Fox, M. Scott Marshall
J. Biomed. Informatics2
2012 Emerging practices for mapping and linking life sciences data using RDF - A case series
abstract
Members of the W3C Health Care and Life Sciences Interest Group (HCLS IG) have published a variety of genomic and drug-related data sets as Resource Description Framework (RDF) triples. This experience has helped the interest group define a general data workflow for mapping health care and life science (HCLS) data to RDF and linking it with other Linked Data sources. This paper presents the workflow along with four case studies that demonstrate the workflow and addresses many of the challenges that may be faced when creating new Linked Data resources. The first case study describes the creation of linked RDF data from microarray data sets while the second discusses a linked RDF data set created from a knowledge base of drug therapies and drug targets. The third case study describes the creation of an RDF index of biomedical concepts present in unstructured clinical reports and how this index was linked to a drug side-effect knowledge base. The final case study describes the initial development of a linked data set from a knowledge base of small molecules. This paper also provides a detailed set of recommended practices for creating and publishing Linked Data sources in the HCLS domain in such a way that they are discoverable and usable by people, software agents, and applications. These practices are based on the cumulative experience of the Linked Open Drug Data (LODD) task force of the HCLS IG. While no single set of recommendations can address all of the heterogeneous information needs that exist within the HCLS domains, practitioners wishing to create Linked Data should find the recommendations useful for identifying the tools, techniques, and practices employed by earlier developers. In addition to clarifying available methods for producing Linked Data, the recommendations for metadata should also make the discovery and consumption of Linked Data easier.
M. Scott Marshall, Richard D. Boyce, Helena F. Deus, Jun Zhao 0003, Egon L. Willighagen, Matthias Samwald, Elgar Pichler, Janos G. Hajagos, Eric Prud'hommeaux, Susie Stephens
J. Web Semant.9
2011 Interpreting relational databases in the RDF domain
abstract
The W3C's "Direct Mapping of Relational Data to RDF" defines a simple, practical and intuitive interpretation of SQL database tables as RDF graphs. This document specifies the formal data models for RDB (Relational DataBase) and RDF and defines a denotational semantics of RDB in the RDF domain. We show how this mapping treats all of the important features of SQL tables, like cardinality and NULLs, and yields an RDF graph which preserves the relational information.
Alexandre Bertails, Eric Prud'hommeaux
K-CAP2
2009 Semantic Web for Health Care and Life Sciences: a review of the state of the art
abstract
Biomedical researchers need to be able to ask questions that span many heterogeneous data sources in order to make well-informed decisions that may lead to important scientific breakthroughs. For this to be achieved, diverse types of data about drugs, patients, diseases, proteins, cells, pathways and so on must be effectively integrated. Yet, linking disparate biomedical data continues to be a challenge due to inconsistency in naming and heterogeneity in data models and formats. Many organizations are now exploring the use of Semantic Web technologies in the hope of easing the cost of data integration [1]. The benefits promised by the Semantic Web include integration of heterogeneous data using explicit semantics, simplified annotation and sharing of findings, rich explicit models for data representation, aggregation and search, easier re-use of data in unanticipated ways, and the application of logic to infer additional information [2]. The World Wide Web Consortium (W3C) (http://www.w3.org/) has established the Semantic Web for Health Care and Life Sciences Interest Group (HCLS IG) (http://www.w3.org/2001/sw/hcls/) to help organizations in their adoption of the Semantic Web. The HCLS IG is chartered to develop and support the use of Semantic Web technologies to improve collaboration, research and development, innovation, and adoption in the domains of Health Care and Life Sciences. As a part of realizing this vision, a workshop on the Semantic Web for Health Care and Life Sciences was organized in conjunction with WWW2008 (http://esw.w3.org/topic/HCLS/WWW2008) [3]. The workshop provided a review of the latest positions and research in this domain. Five of the seven papers within this issue originated from the HCLS/WWW2008 workshop and review a range of Semantic Web technologies/approaches employed in different biomedical domains. Vandervalk et al. describe ‘The State of the Union’ for the adoption of Semantic Web standards by key institutes in bioinformatics. The paper explores the nature and connectivity of several community-driven semantic warehousing projects. It reports on the progress with the CardioSHARE/Moby-2 project, which aims to make the resources of the ‘Deep Web’ transparently accessible through SPARQL queries. It points out that the warehouse approach is limited, in that queries are confined to the resources that have been selected for inclusion. It also discusses a related problem that the majority of bioinformatics data exist in the ‘Deep Web’, that is, the data does not exist until an application or analytical tool is invoked, and therefore does not have a predictable Web address. It also highlights that the inability to utilize Uniform Resource Identifiers (URIs) to address bioinformatics data is a barrier to its accessibility in the Semantic Web. Das et al. discuss the use of ontologies to bridge diverse Web-based communities. The paper introduces the Science Collaboration Framework (SCF) as a reusable platform for advanced online collaboration in biomedical research. SCF supports structured Web 2.0 community discourse amongst researchers, makes heterogeneous data resources available to collaborating scientists, captures the semantics of the relationships between resources, and structures discourse around the resources. The first instance of the SCF framework is being used to create an open-access online community for stem cell research—StemBook (http://www.stembook.org). The SCF framework has been applied to interdisciplinary areas such as neurodegenerative disease and neuro-repair research, but has broad utility across the natural sciences. Zhao et al. describe various design patterns for representing and querying provenance information relating to mapping links between heterogeneous data from sources in the domain of functional genomics. The paper illustrates the use of named RDF graphs at different levels of granularity to make provenance assertions about linked data. It also demonstrates that these assertions are sufficient to support requirements including data currency, integrity, evidential support and historical queries. Dumontier et al. discuss a number of approaches for capturing pharmacogenomic data and other related information to facilitate data sharing and knowledge discovery. The paper describes how recent advances in Semantic Web technologies have presented exciting new opportunities for knowledge discovery related to pharmacogenomics by representing information with machine-understandable semantics. It illustrates progress in this area with respect to a personalized medicine project which aims to facilitate pharmacogenomics knowledge discovery through intuitive knowledge capture and sophisticated question answering using automated reasoning over expressive ontologies. Manning et al. review several data integration approaches that involve extracting data from a wide variety of public and private data repositories, each of which is associated with a unique vocabulary and schema. The paper presents an implemented data architecture that leverages semantic mapping of experimental metadata to support the rapid development of scientific discovery applications. This achieves the twin goals of reducing architectural complexity while leveraging Semantic Web technologies to provide flexibility, efficiency and more fully characterized data relationships. The architecture consists of a metadata ontology, a metadata repository and an interface that allows access to the repository. The paper describes how this approach allows scientists to discover and link relevant data across diverse data sources. It provides a platform for development of integrative informatics applications. Chen et al. survey the feasibility and state of the art for using Semantic Web technology to represent, integrate and analyze knowledge in a range of biomedical networks. The paper introduces a conceptual framework to enable researchers to integrate graph mining with ontology reasoning in network data analysis. Four case studies are used to demonstrate how semantic graph mining can be applied to the analysis of disease-causal genes, Gene Ontology (GO) category cross-talks, drug efficacy analysis and herb–drug interaction analysis. Ruttenberg et al. review the use of Semantic Web technologies for assembling and querying biomedical knowledge from multiple sources and disciplines. The paper presents the Neurocommons prototype knowledge base, a demonstration intended to show the feasibility and benefits of using Semantic Web technologies. The prototype allows one to explore the scalability of current Semantic Web tools and methods for creating such a resource, and to reveal issues that will need to be addressed in order to further expand its scope and use. The paper demonstrates the utility of the knowledge base by reviewing a few example queries that provide answers to precise questions relevant to the understanding of the disease. There has been a considerable increase in the adoption of Semantic Web technologies in the life sciences and health care over the last 5 years. Much of the adoption has resulted from a strong need to be able to integrate and analyze data across databases, applications and communities. It has been fascinating to witness the breadth of solutions that have been implemented to meet these needs. Some applications have focused on demonstrating the scalability of extremely large triple stores, running on platforms as diverse as clusters, PCs and iPhones. Another group of users is primarily focused on using the latest capabilities in OWL to further knowledge discovery through inference. Others have focused on Linked Data, which allows people to use data browsers to surf across silos of data that have been converted into linked data using approaches like relational to RDF mapping. Recently, there has been much interest in Web 2.0, developments in social networking and mashups (Web applications that combine data from more than one source into a single integrated tool). However, many researchers are now exploring the additional capabilities that come with the Semantic Web (or Web 3.0) as it provides a machine-readable framework as to how people can say things about data. Using these technologies in concert transforms social networking sites from being fun pastimes for teenagers, to being serious research tools for sharing knowledge across communities. The incorporation of the Semantic Web into Wikis greatly enhances their usability through improved categorization and search of data. Scientific publishing has the potential to be transformed through support for community annotations, interconnected citations and services for semantically tagged key concepts and statements within papers. Going forwards, we are expecting to see increased adoption of Semantic Web technologies within both industry and academic settings. It is looking likely that many implementations will focus on light-weight approaches that enable the linking of data across silos, while a few large-scale efforts will rely on the use of heavy-weight ontologies as a top-down approach to data integration. We are also expecting a continued convergence of the Semantic Web and social networking, thereby enabling a more collaborative approach to science. National Institutes of Health Grants (P01 DC04732 and U24 NS051869 to K.-H.C., in part). We thank the many reviewers who contributed their time and expertise to evaluation and revision of papers in this issue.
Kei-Hoi Cheung, Eric Prud'hommeaux, Susie Stephens
Briefings Bioinform.2
2009 A journey to Semantic Web query federation in the life sciences
abstract
BACKGROUND: As interest in adopting the Semantic Web in the biomedical domain continues to grow, Semantic Web technology has been evolving and maturing. A variety of technological approaches including triplestore technologies, SPARQL endpoints, Linked Data, and Vocabulary of Interlinked Datasets have emerged in recent years. In addition to the data warehouse construction, these technological approaches can be used to support dynamic query federation. As a community effort, the BioRDF task force, within the Semantic Web for Health Care and Life Sciences Interest Group, is exploring how these emerging approaches can be utilized to execute distributed queries across different neuroscience data sources. METHODS AND RESULTS: We have created two health care and life science knowledge bases. We have explored a variety of Semantic Web approaches to describe, map, and dynamically query multiple datasets. We have demonstrated several federation approaches that integrate diverse types of information about neurons and receptors that play an important role in basic, clinical, and translational neuroscience research. Particularly, we have created a prototype receptor explorer which uses OWL mappings to provide an integrated list of receptors and executes individual queries against different SPARQL endpoints. We have also employed the AIDA Toolkit, which is directed at groups of knowledge workers who cooperatively search, annotate, interpret, and enrich large collections of heterogeneous documents from diverse locations. We have explored a tool called "FeDeRate", which enables a global SPARQL query to be decomposed into subqueries against the remote databases offering either SPARQL or SQL query interfaces. Finally, we have explored how to use the vocabulary of interlinked Datasets (voiD) to create metadata for describing datasets exposed as Linked Data URIs or SPARQL endpoints. CONCLUSION: We have demonstrated the use of a set of novel and state-of-the-art Semantic Web technologies in support of a neuroscience query federation scenario. We have identified both the strengths and weaknesses of these technologies. While Semantic Web offers a global data model including the use of Uniform Resource Identifiers (URI's), the proliferation of semantically-equivalent URI's hinders large scale data integration. Our work helps direct research and tool development, which will be of benefit to this community.
Kei-Hoi Cheung, H. Robert Frost, M. Scott Marshall, Eric Prud'hommeaux, Matthias Samwald, Jun Zhao 0003, Adrian Paschke
BMC Bioinform.4
2008 Report on semantic web for health care and life sciences workshop
abstract
The Semantic Web for Health Care and Life Sciences Workshop will be held in Beijing, China, on April 22, 2008. The goal of the workshop is to foster the development and advancement in the use of Semantic Web technologies to facilitate collaboration, research and development, and innovation adoption in the domains of Health Care and Life Sciences, We also encourage the participation of all research communities in this event, with enhanced participation from Asia due to the location of the event. The workshop consists of two invited keynote talks, eight peer-reviewed presentations, and one panel discussion.
Huajun Chen, Kei-Hoi Cheung, Michel Dumontier, Eric Prud'hommeaux, Alan Ruttenberg, Susie Stephens
WWW4
2002 Annotea: an open RDF infrastructure for shared Web annotations
José Kahan, Marja-Riitta Koivunen, Eric Prud'hommeaux, Ralph R. Swick
Comput. Networks3
1997 Network Performance Effects of HTTP/1.1, CSS1, and PNG
abstract
We describe our investigation of the effect of persistent connections, pipelining and link level document compression on our client and server HTTP implementations. A simple test setup is used to verify HTTP/1.1's design and understand HTTP/1.1 implementation strategies. We present TCP and real time performance data between the libwww robot [27] and both the W3C's Jigsaw [28] and Apache [29] HTTP servers using HTTP/1.0, HTTP/1.1 with persistent connections, HTTP/1.1 with pipelined requests, and HTTP/1.1 with pipelined requests and deflate data compression [22]. We also investigate whether the TCP Nagle algorithm has an effect on HTTP/1.1 performance. While somewhat artificial and possibly overstating the benefits of HTTP/1.1, we believe the tests and results approximate some common behavior seen in browsers. The results confirm that HTTP/1.1 is meeting its major design goals. Our experience has been that implementation details are very important to achieve all of the benefits of HTTP/1.1.For all our tests, a pipelined HTTP/1.1 implementation outperformed HTTP/1.0, even when the HTTP/1.0 implementation used multiple connections in parallel, under all network environments tested. The savings were at least a factor of two, and sometimes as much as a factor of ten, in terms of packets transmitted. Elapsed time improvement is less dramatic, and strongly depends on your network connection.Some data is presented showing further savings possible by changes in Web content, specifically by the use of CSS style sheets [10], and the more compact PNG [20] image representation, both recent recommendations of W3C. Time did not allow full end to end data collection on these cases. The results show that HTTP/1.1 and changes in Web content will have dramatic results in Internet and Web performance as HTTP/1.1 and related technologies deploy over the near future. Universal use of style sheets, even without deployment of HTTP/1.1, would cause a very significant reduction in network traffic.This paper does not investigate further performance and network savings enabled by the improved caching facilities provided by the HTTP/1.1 protocol, or by sophisticated use of range requests.
Henrik Frystyk Nielsen, James Gettys, Anselm Baird-Smith, Eric Prud'hommeaux, Håkon Wium Lie, Chris Lilley
SIGCOMM4