VLDB 2026 Research / reviewers in the wild / expert
Guoqian Jiang
dblp:61/1149
· DBLP profile ↗
90ranked-venue papers
21as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 82 · 18 first-author · 21 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identifying autism spectrum disorder diagnosis in children via temporal-frequency hyperdimensional computing with resting-state EEG
Guoqian Jiang, Junxia Han, Xiaoli Li 0002 |
Expert Syst. Appl. | 2 |
| 2026 | A graph-based twin-stream network for EEG motion intention recognition
Nanqing Zhang, Kai Wang 0105, Guoqian Jiang, Lejun Wang |
Expert Syst. Appl. | 3 |
| 2025 | Causal Factorization Graph Network for Interpretable and Scalable Fault Diagnosis of Multiple Wind TurbinesabstractRecent advances in data-driven methods have significantly improved wind turbine (WT) fault diagnosis using SCADA data. However, existing methods predominantly suffer from two critical limitations: correlation-based models lack causal interpretability, and isolated turbine modeling hinders industrial scalability. We propose a Causal Factorization Graph Network (CFGN) that jointly models multiple turbines through adaptive causal graph learning. CFGN dynamically infers sensor causality via attention-guided edge pruning (e.g., blade angle→generator power), disentangles fault-related causal features from turbine-specific biases via dual GNN branches, and enforces domain-invariant fault detection through a hybrid loss function. Evaluated on four operational turbines, CFGN achieves 96.46% accuracy with 18% fewer false alarms, while maintaining 95.27% average F1 score across heterogeneous units. This framework bridges the gap between black-box AI and industrial needs for interpretable, cross-turbine diagnostics, offering a scalable solution for sustainable wind farm operations. Guoqian Jiang, Qun He |
INDIN | 2 |
| 2025 | Dual-Path Contrastive Learning For Wind Turbine Icing DetectionabstractWind energy, characterized by its clean and replenishable nature, is increasingly used worldwide due to its environmental friendliness and wide distribution of resources. However, ice accretion on turbine blades in cold regions, often resulting from cold weather conditions, significantly impacts both the operational performance and security of wind energy production, which significantly increases maintenance costs, resulting in a significant reduction in the energy output performance of wind turbines. Specifically, blade icing alters the aerodynamic characteristics of the blade surface, increases wind resistance, and reduces wind energy conversion efficiency. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity. The key challenges are complex sensor parameter variations, high labeling costs, and data imbalance, making accurate icing prediction difficult. In order to tackle these difficulties, this study introduces a technique based on dual-path contrastive learning. This method balances the dataset using a sliding window technique and utilizes both icing loss features and expert features for dual-path processing to fully exploit the feature sets. Furthermore, ice accretion may result in a non-uniform mass allocation across the blade surfaces. This imbalance can induce vibrations within the turbine system, thereby compromising its operational stability and structural integrity, particularly exhibiting excellent performance in handling data imbalance. Aili Xu, Jiamei Zhou, Xu Cheng 0003, Fan Shi 0001, Yongming Han, Guoqian Jiang |
INDIN | 6 |
| 2025 | Hierarchical Spatiotemporal Graph Network for Fault Diagnosis of Industrial ProcessesabstractIntelligent fault diagnosis of industrial processes has received enormous attention in recent years, and deep learning-based methods have excellent performance in accurate health state detection. However, existing methods cannot fully exploit the complex relationship between different subsystems of industrial processes. To address this limit, we convert industrial data into graph structure data with sensors (as nodes) and topological connections between sensors (as edges) to represent complex interactive information. Specifically, we propose a spatiotemporal graph convolutional network with a hierarchical structure (HiSTGCN) for fault diagnosis of industrial processes. First, a local-global graph framework is constructed to explore the correlation between subsystems fully. Particularly, we propose a hierarchical graph structure with a global graph representing the correlation of sensors between subsystems and several local graphs capturing the correlation of sensors within subsystems to enrich the feature extraction. Then, we design a hierarchical spatiotemporal graph neural network to perform a local-global graph framework in both temporal and spatial dimensions. Finally, a synthesized residual health monitor module based on the principal component analysis (PCA) is designed for fault detection and location. Experiment results on an industrial simulation process dataset and a real wind farm dataset show that HiSTGCN has reliable and superior fault diagnosis performance compared to existing methods. Guoqian Jiang, Kaili Shen, Xiufeng Liu 0001, Xu Cheng 0003 |
IEEE Internet Things J. | 1 |
| 2024 | A Federated Learning Framework for Cloud-Edge Collaborative Fault Diagnosis of Wind TurbinesabstractIn modern Internet of Things-enhanced wind power systems, most existing data-driven fault diagnosis approaches for wind turbines (WTs) are performed under a centralized paradigm that ignores data privacy. Recently, federated learning (FL) presented a solution to enable edge WTs located at isolated sites to collaboratively learn a shared diagnosis model without accessing local privacy-sensitive data. However, the practical issues of fault label heterogeneity among edge clients and scarcity of labeled data still severely impede the generation of a satisfactory diagnosis model. To address these issues, we propose a diagnostic knowledge-based FL framework (DKFLWT) for collaborative fault diagnosis of distributed edge WTs. In our DKFLWT framework, independently learned diagnostic knowledge from each edge client, rather than model parameters in conventional FL, is uploaded to the cloud server to enrich the client-specific information visible to the server and mitigate the adverse effects on model performance caused by label heterogeneity. To enhance the overall efficiency of the framework, we develop a two-stage, single-round training mechanism, in which the cloud server serves as a universal platform that can accommodate the customized requirements of users, implying the convenient integration of semi-supervised learning to enhance the diagnosis performance in scenarios with limited labeled data. Furthermore, a spatio-temporal memory-enhanced autoencoder is designed to sufficiently exploit essential diagnostic knowledge of different fault patterns from each client. Experimental results demonstrate superior diagnosis performance of our DKFLWT framework with an improvement of more than 22.1% in accuracy and 37.2% in training efficiency against several compared methods in all seriously heterogeneous scenarios. Guoqian Jiang, Xiufeng Liu 0001, Xu Cheng 0003 |
IEEE Internet Things J. | 1 |
| 2024 | Class-Imbalanced Spatial-Temporal Feature Learning for Blade Icing Recognition of Wind TurbineabstractBlade icing detection is vital for wind turbines in cold climates, as it can prevent revenue loss and power degradation. Many machine learning models have been proposed to improve the detection of blade icing; however, earlier studies do not adequately address these issues due to the dynamics of sensor correlations and the imbalance of blade icing data, resulting in low precision and a high false alarm rate. In this study, we aim to address both of these challenges in order to identify blade icing more accurately. On this premise, we develop a spatial–temporal graph convolutional network (SGCN) that leverages the graph convolutional network for adaptively analyzing the dynamics of sensor correlations and a distance-based classifier to improve imbalanced learning. Experiments on the public UEA time series classification datasets and the real-world wind turbine datasets indicate that SGCN is capable of state-of-the-art accuracy, especially in the case of extremely imbalanced data. Renfang Wang, Hong Qiu, Guoqian Jiang, Xiufeng Liu 0001, Xu Cheng 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Unsupervised fault diagnosis of wind turbine bearing via a deep residual deformable convolution network based on subdomain adaptation under time-varying speeds
Pengfei Liang 0005, Guoqian Jiang, Lijie Zhang 0002 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Imbalanced learning for wind turbine blade icing detection via spatio-temporal attention model with a self-adaptive weight loss function
Guoqian Jiang, Ruxu Yue, Qun He, Xiaoli Li 0002 |
Expert Syst. Appl. | 1 |
| 2023 | AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
J. Biomed. Informatics | 10 |
| 2023 | Shape Expressions (ShEx) schemas for the FHIR R5 specification
Deepak K. Sharma, Eric Prud'hommeaux, David Booth, Claude J. Nanjo, Guoqian Jiang |
J. Biomed. Informatics | 5 |
| 2022 | Maximizing Interoperability, Enriching EHR Data: Transforming HL7 FHIR Data to RDF Using the FHIR RDF Playground
James Champion, Eric Prud'hommeaux, David Booth, Gaurav Vaidya, James P. Balhoff, Deepak K. Sharma, Guoqian Jiang, Emily R. Pfaff |
AMIA | 7 |
| 2022 | Modeling a Cancer Symptom Control Domain Using HL7 FHIR: Applicability of the Minimal Common Oncology Data Elements (mCODE)
Nan Huo, Yue Yu 0012, Nansu Zong, Andrea Cheville, Claude J. Nanjo, Eric Prud'hommeaux, Deirdre Pachman, Guohui Xiao 0001, Emily R. Pfaff, Christopher G. Chute, Guoqian Jiang, Kathryn J. Ruddy |
AMIA | 11 |
| 2022 | A Comparative Study on the Capability of Real-World Antineoplastic Drug Data Collection by CanMED, ATC and HemOnc
Yue Yu 0012, Kathryn J. Ruddy, Nan Huo, Nansu Zong, Deirdre Pachman, Christopher G. Chute, Emily R. Pfaff, Andrea Cheville, Guoqian Jiang |
AMIA | 9 |
| 2022 | BETA: a comprehensive benchmark for computational drug-target predictionabstractInternal validation is the most popular evaluation strategy used for drug-target predictive models. The simple random shuffling in the cross-validation, however, is not always ideal to handle large, diverse and copious datasets as it could potentially introduce bias. Hence, these predictive models cannot be comprehensively evaluated to provide insight into their general performance on a variety of use-cases (e.g. permutations of different levels of connectiveness and categories in drug and target space, as well as validations based on different data sources). In this work, we introduce a benchmark, BETA, that aims to address this gap by (i) providing an extensive multipartite network consisting of 0.97 million biomedical concepts and 8.5 million associations, in addition to 62 million drug-drug and protein-protein similarities and (ii) presenting evaluation strategies that reflect seven cases (i.e. general, screening with different connectivity, target and drug screening based on categories, searching for specific drugs and targets and drug repurposing for specific diseases), a total of seven Tests (consisting of 344 Tasks in total) across multiple sampling and validation strategies. Six state-of-the-art methods covering two broad input data types (chemical structure- and gene sequence-based and network-based) were tested across all the developed Tasks. The best-worst performing cases have been analyzed to demonstrate the ability of the proposed benchmark to identify limitations of the tested methods for running over the benchmark tasks. The results highlight BETA as a benchmark in the selection of computational strategies for drug repurposing and target discovery. Nansu Zong, Ning Li 0045, Andrew Wen, Victoria Ngo, Yue Yu 0012, Ming Huang 0006, Shaika Chowdhury, Chao Jiang 0002, Sunyang Fu, Richard Weinshilboum, Guoqian Jiang, Lawrence Hunter |
Briefings Bioinform. | 11 |
| 2022 | Design and validation of a FHIR-based EHR-driven phenotyping toolboxabstractOBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error. Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 14 |
| 2022 | Correction to: Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbenchabstractJournal of the American Medical Informatics Association, Volume 28, Issue 11, November 2021, Pages 2313–2324, https://doi.org/10.1093/jamia/ocab174 The revised manuscript entitled “Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbench,” by Jiang, et al. has been submitted by the authors as a replacement for the originally published version of the manuscript. Following publication, the authors were alerted to the article’s noncompliance with the All of Us Research Program’s Data Access and Use Policies. The policies violated are that authorized data users will not: After an extensive effort to obscure values that correspond to fewer than 20 participants to align with the All of Us Research Program’s policies, the authors have submitted a corrected manuscript. This version of the manuscript has been reviewed by the editorial team and is being republished after retraction of the original version of the... Jie Na, Nansu Zong, David E. Midthun, Guoqian Jiang |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | FHIR-Ontop-OMOP: Building clinical knowledge graphs in FHIR RDF with the OMOP Common data ModelabstractBACKGROUND: Knowledge graphs (KGs) play a key role to enable explainable artificial intelligence (AI) applications in healthcare. Constructing clinical knowledge graphs (CKGs) against heterogeneous electronic health records (EHRs) has been desired by the research and healthcare AI communities. From the standardization perspective, community-based standards such as the Fast Healthcare Interoperability Resources (FHIR) and the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) are increasingly used to represent and standardize EHR data for clinical data analytics, however, the potential of such a standard on building CKG has not been well investigated. OBJECTIVE: To develop and evaluate methods and tools that expose the OMOP CDM-based clinical data repositories into virtual clinical KGs that are compliant with FHIR Resource Description Framework (RDF) specification. METHODS: We developed a system called FHIR-Ontop-OMOP to generate virtual clinical KGs from the OMOP relational databases. We leveraged an OMOP CDM-based Medical Information Mart for Intensive Care (MIMIC-III) data repository to evaluate the FHIR-Ontop-OMOP system in terms of the faithfulness of data transformation and the conformance of the generated CKGs to the FHIR RDF specification. RESULTS: A beta version of the system has been released. A total of more than 100 data element mappings from 11 OMOP CDM clinical data, health system and vocabulary tables were implemented in the system, covering 11 FHIR resources. The generated virtual CKG from MIMIC-III contains 46,520 instances of FHIR Patient, 716,595 instances of Condition, 1,063,525 instances of Procedure, 24,934,751 instances of MedicationStatement, 365,181,104 instances of Observations, and 4,779,672 instances of CodeableConcept. Patient counts identified by five pairs of SQL (over the MIMIC database) and SPARQL (over the virtual CKG) queries were identical, ensuring the faithfulness of the data transformation. Generated CKG in RDF triples for 100 patients were fully conformant with the FHIR RDF specification. CONCLUSION: The FHIR-Ontop-OMOP system can expose OMOP database as a FHIR-compliant RDF graph. It provides a meaningful use case demonstrating the potentials that can be enabled by the interoperability between FHIR and OMOP CDM. Generated clinical KGs in FHIR RDF provide a semantic foundation to enable explainable AI applications in healthcare. Guohui Xiao 0001, Emily R. Pfaff, Eric Prud'hommeaux, David Booth, Deepak K. Sharma, Nan Huo, Yue Yu 0012, Nansu Zong, Kathryn J. Ruddy, Christopher G. Chute, Guoqian Jiang |
J. Biomed. Informatics | 11 |
| 2022 | Developing an ETL tool for converting the PCORnet CDM into the OMOP CDM to facilitate the COVID-19 data integration
Yue Yu 0012, Nansu Zong, Andrew Wen, Sijia Liu 0002, Daniel J. Stone, David Knaack, Alanna M. Chamberlain, Emily R. Pfaff, Davera Gabriel, Christopher G. Chute, Nilay Shah, Guoqian Jiang |
J. Biomed. Informatics | 12 |
| 2021 | Automatic Construction of Biomedical Knowledge Graph from Covid-19 Literature
Chao Jiang 0002, Victoria Ngo, Richard Chapman 0001, Guoqian Jiang, Nansu Zong |
AMIA | 5 |
| 2021 | Multi-site Evaluation of Longitudinal Changes in Ejection Fraction in Heart Failure Patients Through Data-driven Phenotyping
Prakash Adekkanattu, Jennifer A. Pacheco, Joseph Kabariti, Daniel J. Stone, Yue Yu 0012, Parag Goyal, Faraz S. Ahmad, Guoqian Jiang, Yuan Luo 0001, Luke V. Rasmussen, Pascal S. Brandt, Jie Xu 0012, Fei Wang 0001, Natalie C. Benda, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 8 |
| 2021 | Supporting EHR-based Cohort Discovery Through User-centered Design: Results of an Early Formative Usability Study
Natalie C. Benda, Pascal S. Brandt, Jessica S. Ancker, Jennifer A. Pacheco, Prakash Adekkanattu, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 6 |
| 2021 | A Deep Learning Framework Using a Pre-trained BERT Model to Predict the Risk of Progression from Mild Cognitive Impairment to Alzheimer's Disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Fei Wang 0001, Richard Isaacson, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 5 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 14 |
| 2021 | On Constraints and Considerations for Extending Support for Natural Language Processing-Based FHIR Resource Generation
Andrew Wen, Luke V. Rasmussen, Daniel J. Stone, Sijia Liu 0002, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Yuan Luo 0001, Fei Wang 0001, Jyotishman Pathak, Guoqian Jiang |
AMIA | 12 |
| 2021 | Feasibility of capturing real-world data from health information technology systems at multiple centers to assess cardiac ablation device outcomes: A fit-for-purpose informatics analysis reportabstractOBJECTIVE: The study sought to conduct an informatics analysis on the National Evaluation System for Health Technology Coordinating Center test case of cardiac ablation catheters and to demonstrate the role of informatics approaches in the feasibility assessment of capturing real-world data using unique device identifiers (UDIs) that are fit for purpose for label extensions for 2 cardiac ablation catheters from the electronic health records and other health information technology systems in a multicenter evaluation. MATERIALS AND METHODS: We focused on data capture and transformation and data quality maturity model specified in the National Evaluation System for Health Technology Coordinating Center data quality framework. The informatics analysis included 4 elements: the use of UDIs for identifying device exposure data, the use of standardized codes for defining computable phenotypes, the use of natural language processing for capturing unstructured data elements from clinical data systems, and the use of common data models for standardizing data collection and analyses. RESULTS: We found that, with the UDI implementation at 3 health systems, the target device exposure data could be effectively identified, particularly for brand-specific devices. Computable phenotypes for study outcomes could be defined using codes; however, ablation registries, natural language processing tools, and chart reviews were required for validating data quality of the phenotypes. The common data model implementation status varied across sites. The maturity level of the key informatics technologies was highly aligned with the data quality maturity model. CONCLUSIONS: We demonstrated that the informatics approaches can be feasibly used to capture safety and effectiveness outcomes in real-world data for use in medical device studies supporting label extensions. Guoqian Jiang, Sanket S. Dhruva, Jiajing Chen, Wade L. Schulz, Amit A. Doshi, Peter A. Noseworthy, Yue Yu 0012, Hobart Patrick Young, Eric Brandt, Keondae R. Ervin, Nilay D. Shah, Joseph S. Ross, Paul Coplan, Joseph P. Drozda |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Characterizing phenotypic abnormalities associated with high-risk individuals developing lung cancer using electronic health records from the All of Us researcher workbenchabstractOBJECTIVE: The study sought to test the feasibility of conducting a phenome-wide association study to characterize phenotypic abnormalities associated with individuals at high risk for lung cancer using electronic health records. MATERIALS AND METHODS: We used the beta release of the All of Us Researcher Workbench with clinical and survey data from a population of 225 000 subjects. We identified 3 cohorts of individuals at high risk to develop lung cancer based on (1) the 2013 U.S. Preventive Services Task Force criteria, (2) the long-term quitters of cigarette smoking criteria, and (3) the younger age of onset criteria. We applied the logistic regression analysis to identify the significant associations between individuals' phenotypes and their risk categories. We validated our findings against a lung cancer cohort from the same population and conducted an expert review to understand whether these associations are known or potentially novel. RESULTS: We found a total of 214 statistically significant associations (P < .05 with a Bonferroni correction and odds ratio > 1.5) enriched in the high-risk individuals from 3 cohorts, and 15 enriched in the low-risk individuals. Forty significant associations enriched in the high-risk individuals and 13 enriched in the low-risk individuals were validated in the cancer cohort. Expert review identified 15 potentially new associations enriched in the high-risk individuals. CONCLUSIONS: It is feasible to conduct a phenome-wide association study to characterize phenotypic abnormalities associated in high-risk individuals developing lung cancer using electronic health records. The All of Us Research Workbench is a promising resource for the research studies to evaluate and optimize lung cancer screening criteria. Jie Na, Nansu Zong, David E. Midthun, Yuan Luo 0001, Guoqian Jiang |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Development of a FHIR RDF data transformation and validation framework and its evaluation
Eric Prud'hommeaux, Josh Collins, David Booth, Kevin J. Peterson, Harold R. Solbrig, Guoqian Jiang |
J. Biomed. Informatics | 6 |
| 2021 | A Spatio-Temporal Multiscale Neural Network Approach for Wind Turbine Fault Diagnosis With Imbalanced SCADA DataabstractThe supervisory control and data acquisition (SCADA) systems are widely installed on wind turbines (WTs) in major wind farms, which produce a large amount of sensory data that can be used for fault diagnosis of WTs. However, these SCADA data are naturally multivariate time series, which represent complex temporal correlations within each sensor variable and spatial correlations between different sensor variables. To effectively capture spatio-temporal correlations in SCADA data, we propose a new spatio-temporal multiscale neural network (STMNN). The proposed STMNN model contains two parallel feature extraction modules: first, a multiscale deep echo state network module to extract temporal multiscale features; second, a multiscale residual network module to extract spatial multiscale features. Additionally, as the WTs are in a normal working state most of the time, there are a large amount of normal SCADA data and few failure SCADA data. To address the data imbalance problem of SCADA data and enhance the fault diagnosis performance, instead of cross-entropy loss, the STMNN model adopts focal loss as loss function. Our proposed STMNN method can provide an end-to-end fault diagnosis solution with imbalanced SCADA data, and is evaluated through experiments on an SCADA dataset from a real wind farm. The experimental results and comparative analysis have proved the effectiveness of our proposed STMNN model in practical applications. Qun He, Yanhua Pang, Guoqian Jiang |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Feasibility of Cross-Platform EHR-Driven Phenotyping Using Clinical Quality Language
Pascal S. Brandt, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Faraz S. Ahmad, Jie Xu 0012, Jessica S. Ancker, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 12 |
| 2020 | Exploring JSON-LD as an Executable Definition of FHIR RDF to Enable Semantics of FHIR Data
Harold R. Solbrig, Dazhi Jiao, Eric Prud'hommeaux, David Booth, Cory M. Endle, Daniel J. Stone, Guoqian Jiang |
AMIA | 7 |
| 2020 | Identification of Alzheimer's Disease Subtypes from Electronic Health Records Using a Data-Driven Approach
Jie Xu 0012, Fei Wang 0001, Prakash Adekkanattu, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Yuan Luo 0001, Chengsheng Mao, Jennifer A. Pacheco, Luke V. Rasmussen, Yiye Zhang, Richard Isaacson, Jyotishman Pathak |
AMIA | 6 |
| 2020 | Constructing co-occurrence network embeddings to assist association extraction for COVID-19 and other coronavirus infectious diseasesabstractOBJECTIVE: As coronavirus disease 2019 (COVID-19) started its rapid emergence and gradually transformed into an unprecedented pandemic, the need for having a knowledge repository for the disease became crucial. To address this issue, a new COVID-19 machine-readable dataset known as the COVID-19 Open Research Dataset (CORD-19) has been released. Based on this, our objective was to build a computable co-occurrence network embeddings to assist association detection among COVID-19-related biomedical entities. MATERIALS AND METHODS: Leveraging a Linked Data version of CORD-19 (ie, CORD-19-on-FHIR), we first utilized SPARQL to extract co-occurrences among chemicals, diseases, genes, and mutations and build a co-occurrence network. We then trained the representation of the derived co-occurrence network using node2vec with 4 edge embeddings operations (L1, L2, Average, and Hadamard). Six algorithms (decision tree, logistic regression, support vector machine, random forest, naïve Bayes, and multilayer perceptron) were applied to evaluate performance on link prediction. An unsupervised learning strategy was also developed incorporating the t-SNE (t-distributed stochastic neighbor embedding) and DBSCAN (density-based spatial clustering of applications with noise) algorithms for case studies. RESULTS: The random forest classifier showed the best performance on link prediction across different network embeddings. For edge embeddings generated using the Average operation, random forest achieved the optimal average precision of 0.97 along with a F1 score of 0.90. For unsupervised learning, 63 clusters were formed with silhouette score of 0.128. Significant associations were detected for 5 coronavirus infectious diseases in their corresponding subgroups. CONCLUSIONS: In this study, we constructed COVID-19-centered co-occurrence network embeddings. Results indicated that the generated embeddings were able to extract significant associations for COVID-19 and coronavirus infectious diseases. David Oniani, Guoqian Jiang, Feichen Shen |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | A corpus-driven standardization framework for encoding clinical problems with HL7 FHIR
Kevin J. Peterson, Guoqian Jiang |
J. Biomed. Informatics | 2 |
| 2020 | Identifying sub-phenotypes of acute kidney injury using structured and unstructured electronic health record data with memory networks
Jingyuan Chou, Xi Sheryl Zhang, Yuan Luo 0001, Tamara Isakova, Prakash Adekkanattu, Jessica S. Ancker, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Luke V. Rasmussen, Jyotishman Pathak, Fei Wang 0001 |
J. Biomed. Informatics | 8 |
| 2019 | Evaluating the Portability of an NLP System for Processing Echocardiograms: A Retrospective, Multi-site Observational Study
Prakash Adekkanattu, Guoqian Jiang, Yuan Luo 0001, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Richard C. Kiefer, Daniel J. Stone, Pascal S. Brandt, Yizhen Zhong, Fei Wang 0001, Jessica S. Ancker, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 2 |
| 2019 | Considerations for Improving the Portability of Electronic Health Record-Based Phenotype Algorithms
Luke V. Rasmussen, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Jessica S. Ancker, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 3 |
| 2019 | Developing a FHIR-based EHR phenotyping framework: A case study for identification of patients with obesity and multiple comorbidities from discharge summaries
Na Hong, Andrew Wen, Daniel J. Stone, Shintaro Tsuji, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Prakash Adekkanattu, Fei Wang 0001, Yuan Luo 0001, Jyotishman Pathak, Guoqian Jiang |
J. Biomed. Informatics | 13 |
| 2019 | ADEpedia-on-OHDSI: A next generation pharmacovigilance signal detection platform using the OHDSI common data model
Yue Yu 0012, Kathryn J. Ruddy, Na Hong, Shintaro Tsuji, Andrew Wen, Nilay D. Shah, Guoqian Jiang |
J. Biomed. Informatics | 7 |
| 2019 | DeepLab-Based Spatial Feature Extraction for Hyperspectral Image ClassificationabstractRecently, deep learning has been used for hyperspectral image classification (HSIC) due to its powerful feature learning and classification ability. In this letter, a novel deep learning-based framework based on DeepLab is proposed for HSIC. Inspired by the excellent performance of DeepLab in semantic segmentation, the proposed framework applies DeepLab to excavate spatial features of the hyperspectral image (HSI) pixel to pixel. It breaks through the limitation of patch-wise feature learning in the most of existing deep learning methods used in HSIC. More importantly, it can extract features at multiple scales and effectively avoid the reduction of spatial resolution. Furthermore, to improve the HSIC performance, the spatial features extracted by DeepLab and the spectral features are fused by a weighted fusion method, then the fused features are input into support vector machine for final classification. Experimental results on two public HSI data sets demonstrate that the proposed framework outperformed the traditional methods and the existing deep learning-based methods, especially for small-scale classes. Zijia Niu, Guoqian Jiang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | Clinical Concept Value Sets and Interoperability in Health Data Analytics
Sigfried Gold, Andrea Batch, Robert C. McClure, Guoqian Jiang, Hadi Kharrazi, Rishi Saripalle, Vojtech Huser, Chunhua Weng, Nancy K. Roderer, Ana Szarfman, Niklas Elmqvist, David Gotz |
AMIA | 4 |
| 2018 | Standardizing Heterogeneous Annotation Corpora Using HL7 FHIR for Facilitating their Reuse and Integration in Clinical NLP
Na Hong, Andrew Wen, Majid Rastegar-Mojarad, Sunghwan Sohn, Guoqian Jiang |
AMIA | 6 |
| 2018 | Using HL7 FHIR to Improve Standardization and Interoperability of Common Data Models for Clinical and Translational Research
Guoqian Jiang, Jon D. Duke, Daniella Meeker, Mitra Rocca, Harold R. Solbrig, Shawn N. Murphy |
AMIA | 1 |
| 2018 | Representing Cancer Case Report Forms Using HL7 FHIR
Robinette Renner, Harold R. Solbrig, Guoqian Jiang |
AMIA | 3 |
| 2018 | A Multi-Institutional Review and Validation of Federated Query Results in Multiple Common Data Models
Wade L. Schulz, Hobart Patrick Young, Kathryn J. Ruddy, Nilay D. Shah, Joseph S. Ross, Mitra Rocca, Guoqian Jiang |
AMIA | 8 |
| 2018 | Automated Population of an i2b2 Clinical Data Warehouse using FHIR
Harold R. Solbrig, Na Hong, Shawn N. Murphy, Guoqian Jiang |
AMIA | 4 |
| 2018 | Developing a RadLex-based Name Entity Recognition Tool for Mining Free-Text Radiology Reports
Shintaro Tsuji, Andrew Wen, Guoqian Jiang |
AMIA | 3 |
| 2018 | Developing A Standards-based Signal Detection and Validation Framework of Immune-related Adverse Events Using the OHDSI Common Data Model
Yue Yu 0012, Kathryn J. Ruddy, Shintaro Tsuji, Na Hong, Nilay D. Shah, Guoqian Jiang |
AMIA | 6 |
| 2018 | Mapping Common Data Elements to a Domain Model Using an Artificial Neural Network
Robinette Renner, Shaobo Tan, Dongqi Li, Ada Chaeli van der Zijp-Tan, Ryan G. Benton, Glen M. Borchert, Jingshan Huang, Guoqian Jiang |
BIBM | 10 |
| 2018 | A case study evaluating the portability of an executable computable phenotype algorithm across multiple institutions and electronic health record environmentsabstractElectronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges. Jennifer A. Pacheco, Luke V. Rasmussen, Richard C. Kiefer, Thomas R. Campion Jr., Peter Speltz, Robert J. Carroll, Sarah C. Stallings, Huan Mo, Monika Ahuja, Guoqian Jiang, Eric LaRose, Peggy L. Peissig, Ning Shang 0004, Barbara Benoit, Vivian S. Gainer, Kenneth Borthwick, Kathryn L. Jackson, Ambrish Sharma, Andy Yizhou Wu, Abel N. Kho, Dan M. Roden, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
J. Am. Medical Informatics Assoc. | 10 |
| 2018 | Toward a normalized clinical drug knowledge base in China - applying the RxNorm model to Chinese clinical drugsabstractObjective: In recent years, electronic health record systems have been widely implemented in China, making clinical data available electronically. However, little effort has been devoted to making drug information exchangeable among these systems. This study aimed to build a Normalized Chinese Clinical Drug (NCCD) knowledge base, by applying and extending the information model of RxNorm to Chinese clinical drugs. Methods: Chinese drugs were collected from 4 major resources-China Food and Drug Administration, China Health Insurance Systems, Hospital Pharmacy Systems, and China Pharmacopoeia-for integration and normalization in NCCD. Chemical drugs were normalized using the information model in RxNorm without much change. Chinese patent drugs (i.e., Chinese herbal extracts), however, were represented using an expanded RxNorm model to incorporate the unique characteristics of these drugs. A hybrid approach combining automated natural language processing technologies and manual review by domain experts was then applied to drug attribute extraction, normalization, and further generation of drug names at different specification levels. Lastly, we reported the statistics of NCCD, as well as the evaluation results using several sets of randomly selected Chinese drugs. Results: The current version of NCCD contains 16 976 chemical drugs and 2663 Chinese patent medicines, resulting in 19 639 clinical drugs, 250 267 unique concepts, and 2 602 760 relations. By manual review of 1700 chemical drugs and 250 Chinese patent drugs randomly selected from NCCD (about 10%), we showed that the hybrid approach could achieve an accuracy of 98.60% for drug name extraction and normalization. Using a collection of 500 chemical drugs and 500 Chinese patent drugs from other resources, we showed that NCCD achieved coverages of 97.0% and 90.0% for chemical drugs and Chinese patent drugs, respectively. Conclusion: Evaluation results demonstrated the potential to improve interoperability across various electronic drug systems in China. Li Wang 0077, Yaoyun Zhang, Min Jiang 0007, Jiancheng Dong, Yun Liu 0020, Cui Tao, Guoqian Jiang, Yi Zhou 0005, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2017 | Big Data to Knowledge (BD2K) and the Application of Metadata
Guoqian Jiang, Walter S. Campbell, Timothy Clark, Cui Tao, Mark A. Musen |
AMIA | 1 |
| 2017 | Building an FHIR Ontology based Data Access Framework with the OHDSI Data Repositories
Guoqian Jiang, Guohui Xiao 0001, Richard C. Kiefer, Eric Prud'hommeaux, Harold R. Solbrig |
AMIA | 1 |
| 2017 | Leveraging Value Sets from the Value Set Authority Center (VSAC) in a Standards-Based Clinical Data Repository
Richard C. Kiefer, Luke V. Rasmussen, Jennifer A. Pacheco, Peter Speltz, Joshua C. Denny, William K. Thompson, Jyotishman Pathak, Guoqian Jiang |
AMIA | 8 |
| 2017 | Mining Hierarchies and Similarity Clusters from Value Set Repositories
Kevin J. Peterson, Guoqian Jiang, Scott M. Brue, Feichen Shen |
AMIA | 2 |
| 2017 | Assessing Usability of the D2Refine Platform for Harmonization and Standardization of Clinical Study Data Dictionaries
Deepak K. Sharma, Kevin J. Peterson, Guoqian Jiang |
AMIA | 3 |
| 2017 | The Phenotype Execution and Modeling Architecture: A Roadmap Towards Next-generation Phenotyping Using EHRs
Peter Speltz, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, William K. Thompson, Guoqian Jiang, Jyotishman Pathak, Joshua C. Denny |
AMIA | 6 |
| 2017 | Modeling and validating HL7 FHIR profiles using semantic web Shape Expressions (ShEx)
Harold R. Solbrig, Eric Prud'hommeaux, Grahame Grieve, Lloyd McKenzie, Joshua C. Mandel, Deepak K. Sharma, Guoqian Jiang |
J. Biomed. Informatics | 7 |
| 2016 | An NLP Extension to the Quality Data Model for EHR-Driven Phenotype Algorithm Authoring and Execution
Guoqian Jiang, William K. Thompson, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, Huan Mo, Peter Speltz, Joshua C. Denny, Jyotishman Pathak |
AMIA | 1 |
| 2016 | Leveraging Terminology Services for Extract-Transform-Load Processes: A User-Centered Approach
Kevin J. Peterson, Guoqian Jiang, Scott M. Brue |
AMIA | 2 |
| 2016 | Standardized Representation of Clinical Study Data Dictionaries with CIMI Archetypes
Deepak K. Sharma, Harold R. Solbrig, Eric Prud'hommeaux, Jyotishman Pathak, Guoqian Jiang |
AMIA | 5 |
| 2016 | A computational framework for converting textual clinical diagnostic criteria into the quality data model
Na Hong, Dingcheng Li, Yue Yu 0012, Qiongying Xiu, Guoqian Jiang |
J. Biomed. Informatics | 6 |
| 2016 | Automated learning of domain taxonomies from text using background knowledge
Julia Hoxha, Guoqian Jiang, Chunhua Weng |
J. Biomed. Informatics | 2 |
| 2016 | Developing a data element repository to support EHR-driven phenotype algorithm authoring and execution
Guoqian Jiang, Richard C. Kiefer, Luke V. Rasmussen, Harold R. Solbrig, Huan Mo, Jennifer A. Pacheco, Jie Xu 0011, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
J. Biomed. Informatics | 1 |
| 2015 | Harmonization of Quality Data Model with HL7 FHIR to Support EHR-driven Phenotype Authoring and Execution: A Pilot Study
Guoqian Jiang, Harold R. Solbrig, Richard C. Kiefer, Luke V. Rasmussen, Huan Mo, Jennifer A. Pacheco, Enid N. H. Montague, Jie Xu 0011, Peter Speltz, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
AMIA | 1 |
| 2015 | Quality Assurance of Cancer Study Common Data Elements Using A Post-Coordination Approach
Guoqian Jiang, Harold R. Solbrig, Eric Prud'hommeaux, Cui Tao, Chunhua Weng, Christopher G. Chute |
AMIA | 1 |
| 2015 | Usability of a phenotype builder prototype and lessons learned for the design of phenotyping tools
Enid N. H. Montague, Jie Xu 0011, Luke V. Rasmussen, Joshua C. Denny, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Peter Speltz, William K. Thompson, Jyotishman Pathak |
AMIA | 5 |
| 2015 | Representing and Validating Cancer Study Metadata Standard Using RDF Shapes Expression Language
Harold R. Solbrig, Eric Prud'hommeaux, Christopher G. Chute, Guoqian Jiang |
AMIA | 4 |
| 2015 | A domain ontology for the Non-Coding RNA fieldabstractIdentification of non-coding RNAs (ncRNAs) has been significantly enhanced due to the rapid advancement in sequencing technologies. On the other hand, semantic annotation of ncRNA data lag behind their identification, and there is a great need to effectively integrate discovery from relevant communities. To this end, the Non-Coding RNA Ontology (NCRO) is being developed to provide a precisely defined ncRNA controlled vocabulary, which can fill a specific and highly needed niche in unification of ncRNA biology. Jingshan Huang, Karen Eilbeck, Judith A. Blake, Dejing Dou, Darren A. Natale, Alan Ruttenberg, Barry Smith 0001, Michael T. Zimmermann, Guoqian Jiang, Bin Wu 0008, Yongqun He, Shaojie Zhang 0001, Xiaowei Wang 0006, Zixing Liu |
BIBM | 9 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational researchabstractOBJECTIVE: To review and evaluate available software tools for electronic health record-driven phenotype authoring in order to identify gaps and needs for future development. MATERIALS AND METHODS: Candidate phenotype authoring tools were identified through (1) literature search in four publication databases (PubMed, Embase, Web of Science, and Scopus) and (2) a web search. A collection of tools was compiled and reviewed after the searches. A survey was designed and distributed to the developers of the reviewed tools to discover their functionalities and features. RESULTS: Twenty-four different phenotype authoring tools were identified and reviewed. Developers of 16 of these identified tools completed the evaluation survey (67% response rate). The surveyed tools showed commonalities but also varied in their capabilities in algorithm representation, logic functions, data support and software extensibility, search functions, user interface, and data outputs. DISCUSSION: Positive trends identified in the evaluation included: algorithms can be represented in both computable and human readable formats; and most tools offer a web interface for easy access. However, issues were also identified: many tools were lacking advanced logic functions for authoring complex algorithms; the ability to construct queries that leveraged un-structured data was not widely implemented; and many tools had limited support for plug-ins or external analytic software. CONCLUSIONS: Existing phenotype authoring tools could enable clinical researchers to work with electronic health record data more efficiently, but gaps still exist in terms of the functionalities of such tools. The present work can serve as a reference point for the future development of similar tools. Jie Xu 0011, Luke V. Rasmussen, Pamela L. Shaw, Guoqian Jiang, Richard C. Kiefer, Huan Mo, Jennifer A. Pacheco, Peter Speltz, Qian Zhu 0003, Joshua C. Denny, Jyotishman Pathak, William K. Thompson, Enid N. H. Montague |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Lexical Term Standardization of ICD-11 Using Semantic Web Technologies
Guoqian Jiang, Harold R. Solbrig, Bedirhan Üstün, Christopher G. Chute |
AMIA | 1 |
| 2014 | Adverse Drug Event-based Stratification of Tumor Mutations: A Case Study of Breast Cancer Patients Receiving Aromatase Inhibitors
Michael T. Zimmermann, Naresh Prodduturi, Christopher G. Chute, Guoqian Jiang |
AMIA | 5 |
| 2013 | Building A Platform for Supporting Clinical Study Meta-Data Standards Authoring Using Scalable Semantic Web Technologies
Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
AMIA | 1 |
| 2013 | A semantic-web oriented representation of the clinical element model for secondary use of electronic health records dataabstractThe clinical element model (CEM) is an information model designed for representing clinical information in electronic health records (EHR) systems across organizations. The current representation of CEMs does not support formal semantic definitions and therefore it is not possible to perform reasoning and consistency checking on derived models. This paper introduces our efforts to represent the CEM specification using the Web Ontology Language (OWL). The CEM-OWL representation connects the CEM content with the Semantic Web environment, which provides authoring, reasoning, and querying tools. This work may also facilitate the harmonization of the CEMs with domain knowledge represented in terminology models as well as other clinical information models such as the openEHR archetype model. We have created the CEM-OWL meta ontology based on the CEM specification. A convertor has been implemented in Java to automatically translate detailed CEMs from XML to OWL. A panel evaluation has been conducted, and the results show that the OWL modeling can faithfully represent the CEM specification and represent patient data. Cui Tao, Guoqian Jiang, Thomas A. Oniki, Robert R. Freimuth, Qian Zhu 0003, Deepak K. Sharma, Jyotishman Pathak, Stanley M. Huff, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | A Proposal Provenance Model for ICD-11 Revision Beta Phase
Guoqian Jiang, Cory M. Endle, Harold R. Solbrig, Christopher G. Chute |
AMIA | 1 |
| 2012 | A Preliminary Study on Normalizing AERS Drug Names Using RxNorm
Liwei Wang 0010, Guoqian Jiang |
AMIA | 2 |
| 2012 | A Standardized Drug and Drug Class Universal Network
Qian Zhu 0003, Guoqian Jiang, Christopher G. Chute |
AMIA | 2 |
| 2012 | Quality evaluation of value sets from cancer study common data elements using the UMLS semantic groupsabstractOBJECTIVE: The objective of this study is to develop an approach to evaluate the quality of terminological annotations on the value set (ie, enumerated value domain) components of the common data elements (CDEs) in the context of clinical research using both unified medical language system (UMLS) semantic types and groups. MATERIALS AND METHODS: The CDEs of the National Cancer Institute (NCI) Cancer Data Standards Repository, the NCI Thesaurus (NCIt) concepts and the UMLS semantic network were integrated using a semantic web-based framework for a SPARQL-enabled evaluation. First, the set of CDE-permissible values with corresponding meanings in external controlled terminologies were isolated. The corresponding value meanings were then evaluated against their NCI- or UMLS-generated semantic network mapping to determine whether all of the meanings fell within the same semantic group. RESULTS: Of the enumerated CDEs in the Cancer Data Standards Repository, 3093 (26.2%) had elements drawn from more than one UMLS semantic group. A random sample (n=100) of this set of elements indicated that 17% of them were likely to have been misclassified. DISCUSSION: The use of existing semantic web tools can support a high-throughput mechanism for evaluating the quality of large CDE collections. This study demonstrates that the involvement of multiple semantic groups in an enumerated value domain of a CDE is an effective anchor to trigger an auditing point for quality evaluation activities. CONCLUSION: This approach produces a useful quality assurance mechanism for a clinical study CDE repository. Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Quality evaluation of cancer study Common Data Elements using the UMLS Semantic NetworkabstractThe binding of controlled terminology has been regarded as important for standardization of Common Data Elements (CDEs) in cancer research. However, the potential of such binding has not yet been fully explored, especially its quality assurance aspect. The objective of this study is to explore whether there is a relationship between terminological annotations and the UMLS Semantic Network (SN) that can be exploited to improve those annotations. We profiled the terminological concepts associated with the standard structure of the CDEs of the NCI Cancer Data Standards Repository (caDSR) using the UMLS SN. We processed 17798 data elements and extracted 17526 primary object class/property concept pairs. We identified dominant semantic types for the categories "object class" and "property" and determined that the preponderance of the instances were disjoint (i.e. the intersection of semantic types between the two categories is empty). We then performed a preliminary evaluation on the data elements whose asserted primary object class/property concept pairs conflict with this observation - where the semantic type of the object class fell into a SN category typically used by property or visa-versa. In conclusion, the UMLS SN based profiling approach is feasible for the quality assurance and accessibility of the cancer study CDEs. This approach could provide useful insight about how to build mechanisms of quality assurance in a meta-data repository. Guoqian Jiang, Harold R. Solbrig, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2009 | Viewpoint Paper: Auditing the Semantic Completeness of SNOMED CT Using Formal Concept AnalysisabstractOBJECTIVE: This study sought to develop and evaluate an approach for auditing the semantic completeness of the SNOMED CT contents using a formal concept analysis (FCA)-based model. DESIGN: We developed a model for formalizing the normal forms of SNOMED CT expressions using FCA. Anonymous nodes, identified through the analyses, were retrieved from the model for evaluation. Two quasi-Poisson regression models were developed to test whether anonymous nodes can evaluate the semantic completeness of SNOMED CT contents (Model 1), and for testing whether such completeness differs between 2 clinical domains (Model 2). The data were randomly sampled from all the contexts that could be formed in the 2 largest domains: Procedure and Clinical Finding. Case studies (n = 4) were performed on randomly selected anonymous node samples for validation. MEASUREMENTS: In Model 1, the outcome variable is the number of fully defined concepts within a context, while the explanatory variables are the number of lattice nodes and the number of anonymous nodes. In Model 2, the outcome variable is the number of anonymous nodes and the explanatory variables are the number of lattice nodes and a binary category for domain (Procedure/Clinical Finding). RESULTS: A total of 5,450 contexts from the 2 domains were collected for analyses. Our findings revealed that the number of anonymous nodes had a significant negative correlation with the number of fully defined concepts within a context (p < 0.001). Further, the Clinical Finding domain had fewer anonymous nodes than the Procedure domain (p < 0.001). Case studies demonstrated that the anonymous nodes are an effective index for auditing SNOMED CT. CONCLUSION: The anonymous nodes retrieved from FCA-based analyses are a candidate proxy for the semantic completeness of the SNOMED CT contents. Our novel FCA-based approach can be useful for auditing the semantic completeness of SNOMED CT contents, or any large ontology, within or across domains. Guoqian Jiang, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Formalizing ICD coding rules using Formal Concept Analysis
Guoqian Jiang, Jyotishman Pathak, Christopher G. Chute |
J. Biomed. Informatics | 1 |
| 2008 | LexValueSets: An Approach for Context-Driven Value Sets Extraction
Jyotishman Pathak, Guoqian Jiang, Sridhar O. Dwarkanath, James D. Buntrock, Christopher G. Chute |
AMIA | 2 |
| 2006 | Semantic distribution study of noun-noun compounds in the Japanese CT clinical reports
Naoki Nishimoto, Satoshi Terae, Guoqian Jiang, Masahito Uesugi, Takayoshi Terashita, Takumi Tanikawa, Akira Endou, Katsuhiko Ogasawara, Tsunetaro Sakurai |
AMIA | 3 |
| 2006 | Model Analysis for Optimal Allocation of Pediatric Emergency Center
Takumi Tanikawa, Hisateru Ohba, Takayoshi Terashita, Masahito Uesugi, Guoqian Jiang, Katsuhiko Ogasawara, Tsunetaro Sakurai |
AMIA | 5 |
| 2005 | Extraction of Specific Nursing Terms Using Corpora Comparison
Guoqian Jiang, Hitomi Sato, Akira Endoh, Katsuhiko Ogasawara, Tsunetaro Sakurai |
AMIA | 1 |
| 2003 | Clinicians' perceptions and the relevant computer-based information needs towards the practice of evidence based medicine
Guoqian Jiang, Katsuhiko Ogasawara, Akira Endoh, Tsunetaro Sakurai |
AMIA | 1 |
| 2003 | Interpretive Structural Modeling for introducing of image information system at the middle-scale hospitals
Katsuhiko Ogasawara, Hiroko Yamashina, Tomoko Kamiya, Guoqian Jiang, Tsunetaro Sakurai |
AMIA | 4 |
| 2000 | How Well Do Clinicians Use Computer-based Information for Clinical Practice? Postal Survey of Clinicians' Views in Japan
Guoqian Jiang, Katsuhiko Ogasawara, Akira Endoh, Tsunetaro Sakurai |
AMIA | 1 |
| 2000 | Influence of the Size of Color Images Taken by Digital Camera on Image Reading: ROC Curve on Red Simulative Wound
Katsuhiko Ogasawara, Kuniko Ito, Guoqian Jiang, Yukiko Uematsu, Akira Endoh, Tsunetaro Sakurai, Michio Kimura, Yasuyuki Fukuhara |
AMIA | 3 |