VLDB 2026 Research / reviewers in the wild / expert
Guo-Qiang Zhang 0001
dblp:z/GQZhang1 · also Guoqiang Zhang 0001
· DBLP profile ↗
94ranked-venue papers
33as first author
10since 2021 · last 2025
0000-0002-3663-1109ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 53 · 8 first-author · 9 since 2021Theory of computation · 33 · 22 first-authorDatabases, data management, data science and information retrieval · 8 · 5 first-authorArtificial intelligence and machine learning · 6 · 3 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Temporal Ensemble Logic for Integrative Representation of the Entirety of Clinical Trials
Yan Huang 0034, Rashmie Abeysinghe, Zenan Sun, Pengze Li, Xing He 0003, Shiqiang Tao, Cui Tao, Jiang Bian 0001, Licong Cui, Guo-Qiang Zhang 0001 |
TIME | 12 |
| 2025 | Quantitatively assessing the impact of the quality of SNOMED CT subtype hierarchy on cohort queriesabstractOBJECTIVE: SNOMED CT provides a standardized terminology for clinical concepts, allowing cohort queries over heterogeneous clinical data including Electronic Health Records (EHRs). While it is intuitive that missing and inaccurate subtype (or is-a) relations in SNOMED CT reduce the recall and precision of cohort queries, the extent of these impacts has not been formally assessed. This study fills this gap by developing quantitative metrics to measure these impacts and performing statistical analysis on their significance. MATERIAL AND METHODS: We used the Optum de-identified COVID-19 Electronic Health Record dataset. We defined micro-averaged and macro-averaged recall and precision metrics to assess the impact of missing and inaccurate is-a relations on cohort queries. Both practical and simulated analyses were performed. Practical analyses involved 407 missing and 48 inaccurate is-a relations confirmed by domain experts, with statistical testing using Wilcoxon signed-rank tests. Simulated analyses used two random sets of 400 is-a relations to simulate missing and inaccurate is-a relations. RESULTS: Wilcoxon signed-rank tests from both practical and simulated analyses (P-values < .001) showed that missing is-a relations significantly reduced the micro- and macro-averaged recall, and inaccurate is-a relations significantly reduced the micro- and macro-averaged precision. DISCUSSION: The introduced impact metrics can assist SNOMED CT maintainers in prioritizing critical hierarchical defects for quality enhancement. These metrics are generally applicable for assessing the quality impact of a terminology's subtype hierarchy on its cohort query applications. CONCLUSION: Our results indicate a significant impact of missing and inaccurate is-a relations in SNOMED CT on the recall and precision of cohort queries. Our work highlights the importance of high-quality terminology hierarchy for cohort queries over EHR data and provides valuable insights for prioritizing quality improvements of SNOMED CT's hierarchy. Xubing Hao, Yan Huang 0034, Jay Shi, Rashmie Abeysinghe, Cui Tao, Kirk Roberts, Guo-Qiang Zhang 0001, Licong Cui |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Identifying Factors Associated with COVID-19 All-Cause 90-Day Readmission: Machine Learning Approaches
Shiwei Lin, Shiqiang Tao, Yan Huang 0034, Guo-Qiang Zhang 0001 |
AIME (1) | 5 |
| 2023 | A hierarchical strategy to minimize privacy risk when linking "De-identified" data in biomedical research consortia
Lucila Ohno-Machado, Xiaoqian Jiang, Tsung-Ting Kuo, Shiqiang Tao, Pritham Ram, Guo-Qiang Zhang 0001, Hua Xu 0001 |
J. Biomed. Informatics | 7 |
| 2022 | Temporal Cohort Logic
Guo-Qiang Zhang 0001, Yan Huang 0034, Licong Cui |
AMIA | 1 |
| 2022 | UDA-CT: A General Framework for CT Image StandardizationabstractLarge-scale CT image studies often suffer from a lack of homogeneity regarding radiomic characteristics due to the images acquired with scanners from different vendors or with different reconstruction algorithms. We propose a deep learning-based framework called UDA-CT to tackle the homogeneity issue by leveraging both paired and unpaired images. Using UDA-CT, the CT images can be standardized both from different acquisition protocols of the same scanner and CT images acquired using a similar protocol but scanners from different vendors. UDA-CT incorporates recent advances in deep learning including domain adaptation and adversarial augmentation. It includes a unique design for model training batch which integrates nonstandard images and their adversarial variations to enhance model generalizability. The experimental results show that UDA-CT significantly improves the performance of the cross-scanner image standardization by utilizing both paired and unpaired data. Md. Selim, Jie Zhang 0092, Baowei Fei, Matthew A. Lewis, Guo-Qiang Zhang 0001, Jin Chen 0004 |
BIBM | 5 |
| 2021 | Identifying Sleep-Related Factors Associated with Cognitive Function in a Hispanics/Latinos Cohort: A Dual Random Forest Approach
Licong Cui, Paul E. Schulz, Guo-Qiang Zhang 0001 |
AMIA | 5 |
| 2021 | Cross-Vendor CT Image Data Harmonization Using CVH-CT
Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Gary Yeeming Ge, Jin Chen 0004 |
AMIA | 4 |
| 2021 | CT Image Harmonization for Enhancing Radiomics StudiesabstractWhile remarkable advances have been made in Computed Tomography (CT), most of the existing efforts focus on imaging enhancement while reducing radiation dose. How to normalize CT images acquired using non-standard protocols is vital for decision-making in cross-center large-scale radiomics studies but remains the boundary to explore. We develop a novel GAN-based image standardization algorithm called RadiomicGAN to mitigate the discrepancy caused by using non-standard acquisition protocols. In RadiomicGAN, a pre-trained U-Net has been adopted as part of the generator to learn radiomic feature distributions efficiently, and a novel training approach, called Window Training, has been developed to smoothly transform the pre-trained model to the medical imaging domain. In the experiments, we compared RadiomicGAN with four state-of-the-art CT image standardization approaches on both patient and phantom CT images acquired using three different reconstruction kernels. We objectively evaluated model performance based on more than 1,000 radiomic features. The results show that RadiomicGAN clearly outperforms the compared models. The source code, manual, and sample data are available at https://github.con selim-iitdu/radiomicGAN. Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Jin Chen 0004 |
BIBM | 4 |
| 2021 | ELII: A novel inverted index for fast temporal query, with application to a large Covid-19 EHR datasetabstractFast temporal query on large EHR-derived data sources presents an emerging big data challenge, as this query modality is intractable using conventional strategies that have not focused on addressing Covid-19-related research needs at scale. We introduce a novel approach called Event-level Inverted Index (ELII) to optimize time trade-offs between one-time batch preprocessing and subsequent open-ended, user-specified temporal queries. An experimental temporal query engine has been implemented in a NoSQL database using our new ELII strategy. Near-real-time performance was achieved on a large Covid-19 EHR dataset, with 1.3 million unique patients and 3.76 billion records. We evaluated the performance of ELII on several types of queries: classical (non-temporal), absolute temporal, and relative temporal. Our experimental results indicate that ELII accomplished these queries in seconds, achieving average speed accelerations of 26.8 times on relative temporal query, 88.6 times on absolute temporal query, and 1037.6 times on classical query compared to a baseline approach without using ELII. Our study suggests that ELII is a promising approach supporting fast temporal query, an important mode of cohort development for Covid-19 studies. Yan Huang 0034, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 3 |
| 2020 | STAN-CT: Standardizing CT Image using Generative Adversarial Networks
Md. Selim, Jie Zhang 0092, Baowei Fei, Guo-Qiang Zhang 0001, Jin Chen 0004 |
AMIA | 4 |
| 2020 | Temporal phenotyping for transitional disease progress: An application to epilepsy and Alzheimer's disease
Yejin Kim 0001, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Xiaoqian Jiang |
J. Biomed. Informatics | 3 |
| 2019 | SeizureBank: A Repository of Analysis-ready Seizure Signal Data
Yan Huang 0034, Shiqiang Tao, Licong Cui, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
AMIA | 6 |
| 2019 | A representation of continuous domains via relationally approximable concepts in a generalized framework of formal concept analysis
Lankun Guo, Qingguo Li, Guo-Qiang Zhang 0001 |
Int. J. Approx. Reason. | 3 |
| 2018 | A Cross-Cohort Query System for the National Sleep Research Resource (NSRR)
Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 2 |
| 2018 | Retrospective analysis of health claims to evaluate pharmacotherapies with potential for repurposing: Association of bupropion and stimulant use disorder remission
Emily R. Hankosky, Heather Bush, Linda P. Dwoskin, Daniel R. Harris, Darren W. Henderson, Guo-Qiang Zhang 0001, Patricia R. Freeman, Jeffery C. Talbert |
AMIA | 6 |
| 2018 | System Demonstration: Integration of Patient Reported Outcomes with Electronic Health Records - the EASI-PRO Project
Justin Starren, Daniella Meeker, Kenneth D. Mandl, Guo-Qiang Zhang 0001, Alyssa White, Raheel Sayeed, Daniel Gottlieb 0001, Alex Wormuth, Welmoed Van Deen, Shiqiang Tao |
AMIA | 4 |
| 2018 | The National Sleep Research Resource: towards a sleep data commonsabstractObjective: The gold standard for diagnosing sleep disorders is polysomnography, which generates extensive data about biophysical changes occurring during sleep. We developed the National Sleep Research Resource (NSRR), a comprehensive system for sharing sleep data. The NSRR embodies elements of a data commons aimed at accelerating research to address critical questions about the impact of sleep disorders on important health outcomes. Approach: We used a metadata-guided approach, with a set of common sleep-specific terms enforcing uniform semantic interpretation of data elements across three main components: (1) annotated datasets; (2) user interfaces for accessing data; and (3) computational tools for the analysis of polysomnography recordings. We incorporated the process for managing dataset-specific data use agreements, evidence of Institutional Review Board review, and the corresponding access control in the NSRR web portal. The metadata-guided approach facilitates structural and semantic interoperability, ultimately leading to enhanced data reusability and scientific rigor. Results: The authors curated and deposited retrospective data from 10 large, NIH-funded sleep cohort studies, including several from the Trans-Omics for Precision Medicine (TOPMed) program, into the NSRR. The NSRR currently contains data on 26 808 subjects and 31 166 signal files in European Data Format. Launched in April 2014, over 3000 registered users have downloaded over 130 terabytes of data. Conclusions: The NSRR offers a use case and an example for creating a full-fledged data commons. It provides a single point of access to analysis-ready physiological signals from polysomnography obtained from multiple sources, and a wide variety of clinical data to facilitate sleep research. Guo-Qiang Zhang 0001, Licong Cui, Remo Mueller, Shiqiang Tao, Matthew Kim, Michael Rueschman, Sara Mariani, Daniel R. Mobley, Susan Redline |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | An efficient, large-scale, non-lattice-detection algorithm for exhaustive structural auditing of biomedical ontologies
Guo-Qiang Zhang 0001, Guangming Xing, Licong Cui |
J. Biomed. Informatics | 1 |
| 2018 | Auditing SNOMED CT hierarchical relations based on lexical features of concepts in non-lattice subgraphs
Licong Cui, Olivier Bodenreider, Jay Shi, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 4 |
| 2018 | Quality assurance of biomedical terminologies and ontologies
James Geller, Yehoshua Perl, Licong Cui, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 4 |
| 2018 | HyCLASSS: A Hybrid Classifier for Automatic Sleep Stage ScoringabstractAutomatic identification of sleep stage is an important step in a sleep study. In this paper, we propose a hybrid automatic sleep stage scoring approach, named HyCLASSS, based on single channel electroencephalogram (EEG). HyCLASSS, for the first time, leverages both signal and stage transition features of human sleep for automatic identification of sleep stages. HyCLASSS consists of two parts: A random forest classifier and correction rules. Random forest classifier is trained using 30 EEG signal features, including temporal, frequency, and nonlinear features. The correction rules are constructed based on stage transition feature, importing the continuity property of sleep, and characteristic of sleep stage transition. Compared with the gold standard of manual scoring using Rechtschaffen and Kales criterion, the overall accuracy and kappa coefficient applied on 198 subjects has reached 85.95% and 0.8046 in our experiment, respectively. The performance of HyCLASS compared favorably to previous work, and it could be integrated with sleep evaluation or sleep diagnosis system in the future. Licong Cui, Shiqiang Tao, Jing Chen 0002, Xiang Zhang 0001, Guo-Qiang Zhang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2017 | Can SNOMED CT Changes Be Used as a Surrogate Standard for Evaluating the Performance of Its Auditing Methods?
Guo-Qiang Zhang 0001, Yan Huang 0034, Licong Cui |
AMIA | 1 |
| 2017 | SpindleSphere: A Web-based Platform for Large-scale Sleep Spindle Analysis and Visualization
Licong Cui, Shiqiang Tao, Ningzhou Zeng, Guo-Qiang Zhang 0001 |
AMIA | 5 |
| 2017 | Facilitating Cohort Discovery by Enhancing Ontology Exploration, Query Management and Query Sharing for Large Clinical Data Repositories
Shiqiang Tao, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 4 |
| 2017 | A Practical Scientific Workflow of Sharing Large Selected Dataset from Clinical Research Data Repository
Shiqiang Tao, Steven K. Roggenkamp, Guo-Qiang Zhang 0001 |
AMIA | 3 |
| 2017 | Spark-MCA: Large-scale, Exhaustive Formal Concept Analysis for Evaluating the Semantic Completeness of SNOMED CT
Wei Zhu 0010, Guo-Qiang Zhang 0001, Licong Cui |
AMIA | 2 |
| 2017 | ImageSfERe: Image sharing for epilepsy researchabstractSudden Unexpected Death in Epilepsy (SUDEP) is the leading mode of epilepsy-related death and is most common in patients with intractable, frequent, and continuing seizures. Magnetic Resonance Imaging (MRI) and other neuroimaging techniques create an important data source for investigators to collaborate and share a larger cohort of potential SUDEP patient data. To address the challenges of sharing neuroimaging data across institutions with good semantic precision, we developed ImageSfERe, a web-based image sharing module for Epilepsy Research. Since various aspects of patient information need to be harmonized and integrated, a dedicated questionnaire consisting of 248 questions about patient background information is incorporated into ImageSfERe. Eight repositories of patient de-identified images in Digital Image and Communications in Medicine (DICOM) format, over 23 GB of data, have been successfully processed and made available via ImageSfERe. To enable interoperability, we also adapted epilepsy neuroimaging common data elements from the National Institute of Health (NIH). Steven K. Roggenkamp, Shiqiang Tao, Guo-Qiang Zhang 0001 |
BIBM | 4 |
| 2017 | Evaluation of relational and NoSQL approaches for patient cohort identification from heterogeneous data sourcesabstractPatient cohort discovery across heterogeneous data sources is a challenging task, which may involve a complicated process of data loading, harmonization, and querying. Most existing cohort identification tools use a relational database model implemented in SQL for storing patient data. However, SQL databases have restrictions on the maximum number of columns in a table, which necessitates the breaking down of high-dimensional data into multiple tables and affects query performance as a result. In this paper, we proposed two NoSQL-based patient cohort query systems based on an existing SQL-based system for a cross-cohort query interface for the National Sleep Resource Research (NSRR). We used eight NSRR datasets in our experiment to evaluate the performance of NoSQL-based and SQL-based systems in data loading, harmonization, and query. Our experiment showed that NoSQL-based approaches outperformed the SQL-based, and NoSQL-based systems are rather promising for developing patient cohort query systems across heterogeneous data sources. Ningzhou Zeng, Guo-Qiang Zhang 0001, Licong Cui |
BIBM | 2 |
| 2017 | Mining non-lattice subgraphs for detecting missing hierarchical relations and concepts in SNOMED CTabstractOBJECTIVE: Quality assurance of large ontological systems such as SNOMED CT is an indispensable part of the terminology management lifecycle. We introduce a hybrid structural-lexical method for scalable and systematic discovery of missing hierarchical relations and concepts in SNOMED CT. MATERIAL AND METHODS: All non-lattice subgraphs (the structural part) in SNOMED CT are exhaustively extracted using a scalable MapReduce algorithm. Four lexical patterns (the lexical part) are identified among the extracted non-lattice subgraphs. Non-lattice subgraphs exhibiting such lexical patterns are often indicative of missing hierarchical relations or concepts. Each lexical pattern is associated with a potential specific type of error. RESULTS: Applying the structural-lexical method to SNOMED CT (September 2015 US edition), we found 6801 non-lattice subgraphs that matched these lexical patterns, of which 2046 were amenable to visual inspection. We evaluated a random sample of 100 small subgraphs, of which 59 were reviewed in detail by domain experts. All the subgraphs reviewed contained errors confirmed by the experts. The most frequent type of error was missing is-a relations due to incomplete or inconsistent modeling of the concepts. CONCLUSIONS: Our hybrid structural-lexical method is innovative and proved effective not only in detecting errors in SNOMED CT, but also in suggesting remediation for these errors. Licong Cui, Wei Zhu 0010, Shiqiang Tao, James T. Case, Olivier Bodenreider, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 6 |
| 2017 | Preface
Achim Jung, Guo-Qiang Zhang 0001 |
Math. Struct. Comput. Sci. | 2 |
| 2016 | ODaCCI: Ontology-guided Data Curation for Multisite Clinical Research Data Integration in the NINDS Center for SUDEP Research
Licong Cui, Yan Huang 0034, Shiqiang Tao, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
AMIA | 5 |
| 2016 | DCDS: A Real-time Data Capture and Personalized Decision Support System for Heart Failure Patients in Skilled Nursing Facilities
Wei Zhu 0010, Lingyun Luo, Tarun Jain, Rebecca S. Boxer, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 6 |
| 2016 | Biomedical Ontology Quality Assurance Using a Big Data ApproachabstractThis article presents recent progresses made in using scalable cloud computing environment, Hadoop and MapReduce, to perform ontology quality assurance (OQA), and points to areas of future opportunity. The standard sequential approach used for implementing OQA methods can take weeks if not months for exhaustive analyses for large biomedical ontological systems. With OQA methods newly implemented using massively parallel algorithms in the MapReduce framework, several orders of magnitude in speed-up can be achieved (e.g., from three months to three hours). Such dramatically reduced time makes it feasible not only to perform exhaustive structural analysis of large ontological hierarchies, but also to systematically track structural changes between versions for evolutional analysis. As an exemplar, progress is reported in using MapReduce to perform evolutional analysis and visualization on the Systemized Nomenclature of Medicine—Clinical Terms (SNOMED CT), a prominent clinical terminology system. Future opportunities in three areas are described: one is to extend the scope of MapReduce-based approach to existing OQA methods, especially for automated exhaustive structural analysis. The second is to apply our proposed MapReduce Pipeline for Lattice-based Evaluation (MaPLE) approach, demonstrated as an exemplar method for SNOMED CT, to other biomedical ontologies. The third area is to develop interfaces for reviewing results obtained by OQA methods and for visualizing ontological alignment and evolution, which can also take advantage of cloud computing technology to systematically pre-compute computationally intensive jobs in order to increase performance during user interactions with the visualization interface. Advances in these directions are expected to better support the ontological engineering lifecycle. Licong Cui, Shiqiang Tao, Guo-Qiang Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2015 | COBE: A Conjunctive Ontology Browser and Explorer for Visualizing SNOMED CT Fragments
Wei Zhu 0010, Shiqiang Tao, Licong Cui, Guo-Qiang Zhang 0001 |
AMIA | 5 |
| 2015 | RREV: Reconfigurable Rendering Engine for visualization of clinically annotated polysomnogramsabstractIn sleep medicine, clinical studies often use their own data dictionaries for capturing clinical sleep events using proprietary signal analysis software [1][2]. Visualization of polysomnograms and their associated events from multiple distinct studies, such as for the National Sleep Research Resource (NSRR)[3], is an unresolved issue. Currently, there is no known visualization software for the European Data Format (EDF) that can be dynamically configured to support rendering of sleep events for multiple vendor formats. To address this challenge, domain ontology has been developed as a part of NSRR to model all sleep medicine terms and concepts to provide a common schema for addressing the structural and semantic heterogeneity of multiple vendor formats [4]. A Reconfigurable Rendering Engine using Abstract Factory pattern [5] and domain ontology provides a standard interface for accessing ontology-enabled clinical events for the visualization of electrophysiological signals. About 11,078 polysomnograms (8,444 SHHS, 860 CHAT, 591 HeartBEAT, 730 CFS, 453 SOF) [12] in EDF have been processed resulting in 1.1TB of web-accessible and reusable PSGs with NSRR standardized event annotations. Catherine P. Jayapandian, Michael G. Morrical, Dennis A. Dean, Shiqiang Tao, Daniel R. Mobley, Matthew Kim, Michael Rueschman, Kenneth A. Loparo, Susan Redline, Guo-Qiang Zhang 0001 |
BIBM | 11 |
| 2015 | Phenome-driven disease genetics prediction toward drug discoveryabstractMOTIVATION: Discerning genetic contributions to diseases not only enhances our understanding of disease mechanisms, but also leads to translational opportunities for drug discovery. Recent computational approaches incorporate disease phenotypic similarities to improve the prediction power of disease gene discovery. However, most current studies used only one data source of human disease phenotype. We present an innovative and generic strategy for combining multiple different data sources of human disease phenotype and predicting disease-associated genes from integrated phenotypic and genomic data. RESULTS: To demonstrate our approach, we explored a new phenotype database from biomedical ontologies and constructed Disease Manifestation Network (DMN). We combined DMN with mimMiner, which was a widely used phenotype database in disease gene prediction studies. Our approach achieved significantly improved performance over a baseline method, which used only one phenotype data source. In the leave-one-out cross-validation and de novo gene prediction analysis, our approach achieved the area under the curves of 90.7% and 90.3%, which are significantly higher than 84.2% (P < e(-4)) and 81.3% (P < e(-12)) for the baseline approach. We further demonstrated that our predicted genes have the translational potential in drug discovery. We used Crohn's disease as an example and ranked the candidate drugs based on the rank of drug targets. Our gene prediction approach prioritized druggable genes that are likely to be associated with Crohn's disease pathogenesis, and our rank of candidate drugs successfully prioritized the Food and Drug Administration-approved drugs for Crohn's disease. We also found literature evidence to support a number of drugs among the top 200 candidates. In summary, we demonstrated that a novel strategy combining unique disease phenotype data with system approaches can lead to rapid drug discovery. AVAILABILITY AND IMPLEMENTATION: nlp. CASE: edu/public/data/DMN Yang Chen 0022, Guo-Qiang Zhang 0001 |
Bioinform. | 3 |
| 2015 | Comparative analysis of a novel disease phenotype network based on clinical manifestations
Yang Chen 0022, Xiang Zhang 0001, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 3 |
| 2014 | A Semantic-based Approach for Exploring Consumer Health Questions Using UMLS
Licong Cui, Shiqiang Tao, Guo-Qiang Zhang 0001 |
AMIA | 3 |
| 2014 | MEDCIS: Multi-Modality Epilepsy Data Capture and Integration System
Guo-Qiang Zhang 0001, Licong Cui, Samden D. Lhatoo, Satya Sanket Sahoo |
AMIA | 1 |
| 2014 | MaPLE: A MapReduce Pipeline for Lattice-based Evaluation and its application to SNOMED CTabstractNon-lattice fragments are often indicative of structural anomalies in ontological systems and, as such, represent possible areas of focus for subsequent quality assurance work. However, extracting the non-lattice fragments in large ontological systems is computationally expensive if not prohibitive, using a traditional sequential approach. In this paper we present a general MapReduce pipeline, called MaPLE (MapReduce Pipeline for Lattice-based Evaluation), for extracting non-lattice fragments in large partially ordered sets and demonstrate its applicability in ontology quality assurance. Using MaPLE in a 30-node Hadoop local cloud, we systematically extracted non-lattice fragments in 8 SNOMED CT versions from 2009 to 2014 (each containing over 300k concepts), with an average total computing time of less than 3 hours per version. With dramatically reduced time, MaPLE makes it feasible not only to perform exhaustive structural analysis of large ontological hierarchies, but also to systematically track structural changes between versions. Our change analysis showed that the average change rates on the non-lattice pairs are up to 38.6 times higher than the change rates of the background structure (concept nodes). This demonstrates that fragments around non-lattice pairs exhibit significantly higher rates of change in the process of ontological evolution. Guo-Qiang Zhang 0001, Wei Zhu 0010, Shiqiang Tao, Olivier Bodenreider, Licong Cui |
IEEE BigData | 1 |
| 2014 | Domain Ontology As Conceptual Model for Big Data Management: Application in Biomedical Informatics
Catherine P. Jayapandian, Aman Dabir, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo |
ER | 5 |
| 2014 | Heart beats in the cloud: distributed analysis of electrophysiological 'Big Data' using cloud computing for epilepsy clinical researchabstractOBJECTIVE: The rapidly growing volume of multimodal electrophysiological signal data is playing a critical role in patient care and clinical research across multiple disease domains, such as epilepsy and sleep medicine. To facilitate secondary use of these data, there is an urgent need to develop novel algorithms and informatics approaches using new cloud computing technologies as well as ontologies for collaborative multicenter studies. MATERIALS AND METHODS: We present the Cloudwave platform, which (a) defines parallelized algorithms for computing cardiac measures using the MapReduce parallel programming framework, (b) supports real-time interaction with large volumes of electrophysiological signals, and (c) features signal visualization and querying functionalities using an ontology-driven web-based interface. Cloudwave is currently used in the multicenter National Institute of Neurological Diseases and Stroke (NINDS)-funded Prevention and Risk Identification of SUDEP (sudden unexplained death in epilepsy) Mortality (PRISM) project to identify risk factors for sudden death in epilepsy. RESULTS: Comparative evaluations of Cloudwave with traditional desktop approaches to compute cardiac measures (eg, QRS complexes, RR intervals, and instantaneous heart rate) on epilepsy patient data show one order of magnitude improvement for single-channel ECG data and 20 times improvement for four-channel ECG data. This enables Cloudwave to support real-time user interaction with signal data, which is semantically annotated with a novel epilepsy and seizure ontology. DISCUSSION: Data privacy is a critical issue in using cloud infrastructure, and cloud platforms, such as Amazon Web Services, offer features to support Health Insurance Portability and Accountability Act standards. CONCLUSION: The Cloudwave platform is a new approach to leverage of large-scale electrophysiological data for advancing multicenter clinical research. Satya Sanket Sahoo, Catherine P. Jayapandian, Farhad Kaffashi, Stephanie Chung, Alireza Bozorgi, Kenneth A. Loparo, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 10 |
| 2014 | Epilepsy and seizure ontology: towards an epilepsy informatics infrastructure for clinical research and patient careabstractOBJECTIVE: Epilepsy encompasses an extensive array of clinical and research subdomains, many of which emphasize multi-modal physiological measurements such as electroencephalography and neuroimaging. The integration of structured, unstructured, and signal data into a coherent structure for patient care as well as clinical research requires an effective informatics infrastructure that is underpinned by a formal domain ontology. METHODS: We have developed an epilepsy and seizure ontology (EpSO) using a four-dimensional epilepsy classification system that integrates the latest International League Against Epilepsy terminology recommendations and National Institute of Neurological Disorders and Stroke (NINDS) common data elements. It imports concepts from existing ontologies, including the Neural ElectroMagnetic Ontologies, and uses formal concept analysis to create a taxonomy of epilepsy syndromes based on their seizure semiology and anatomical location. RESULTS: EpSO is used in a suite of informatics tools for (a) patient data entry, (b) epilepsy focused clinical free text processing, and (c) patient cohort identification as part of the multi-center NINDS-funded study on sudden unexpected death in epilepsy. EpSO is available for download at http://prism.case.edu/prism/index.php/EpilepsyOntology. DISCUSSION: An epilepsy ontology consortium is being created for community-driven extension, review, and adoption of EpSO. We are in the process of submitting EpSO to the BioPortal repository. CONCLUSIONS: EpSO plays a critical role in informatics tools for epilepsy patient care and multi-center clinical research. Satya Sanket Sahoo, Samden D. Lhatoo, Licong Cui, Catherine P. Jayapandian, Alireza Bozorgi, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2014 | Complex epilepsy phenotype extraction from narrative clinical discharge summaries
Licong Cui, Satya Sanket Sahoo, Samden D. Lhatoo, Prashant Rai, Alireza Bozorgi, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 7 |
| 2013 | Creation and Comparative Analysis of a Novel Disease Phenotype Network Based on Clinical Manifestation
Yang Chen 0022, Xiang Zhang 0001, Guo-Qiang Zhang 0001 |
AMIA | 3 |
| 2013 | Cloudwave: Distributed Processing of "Big Data" from Electrophysiological Recordings for Epilepsy Clinical Research Using Hadoop
Catherine P. Jayapandian, Alireza Bozorgi, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo |
AMIA | 5 |
| 2013 | Research and applications: Ontology-guided organ detection to retrieve web images of disease manifestation: towards the construction of a consumer-based health image libraryabstractBACKGROUND: Visual information is a crucial aspect of medical knowledge. Building a comprehensive medical image base, in the spirit of the Unified Medical Language System (UMLS), would greatly benefit patient education and self-care. However, collection and annotation of such a large-scale image base is challenging. OBJECTIVE: To combine visual object detection techniques with medical ontology to automatically mine web photos and retrieve a large number of disease manifestation images with minimal manual labeling effort. METHODS: As a proof of concept, we first learnt five organ detectors on three detection scales for eyes, ears, lips, hands, and feet. Given a disease, we used information from the UMLS to select affected body parts, ran the pretrained organ detectors on web images, and combined the detection outputs to retrieve disease images. RESULTS: Compared with a supervised image retrieval approach that requires training images for every disease, our ontology-guided approach exploits shared visual information of body parts across diseases. In retrieving 2220 web images of 32 diseases, we reduced manual labeling effort to 15.6% while improving the average precision by 3.9% from 77.7% to 81.6%. For 40.6% of the diseases, we improved the precision by 10%. CONCLUSIONS: The results confirm the concept that the web is a feasible source for automatic disease image retrieval for health image database construction. Our approach requires a small amount of manual effort to collect complex disease images, and to annotate them by standard medical ontology terms. Yang Chen 0022, Xiaofeng Ren, Guo-Qiang Zhang 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2013 | An analysis of FMA using structural self-bisimilarity
Lingyun Luo, José L. V. Mejino Jr., Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 3 |
| 2012 | EpiDEA: Extracting Structured Epilepsy and Seizure Information from Patient Discharge Summaries for Cohort Identification
Licong Cui, Samden D. Lhatoo, Guo-Qiang Zhang 0001, Satya Sanket Sahoo, Alireza Bozorgi |
AMIA | 3 |
| 2012 | OPIC: Ontology-driven Patient Information Capturing System for Epilepsy
Satya Sanket Sahoo, Lingyun Luo, Alireza Bozorgi, Samden D. Lhatoo, Guo-Qiang Zhang 0001 |
AMIA | 7 |
| 2012 | An Analysis of Multi-type Relational Interactions in FMA Using Graph Motifs with Disjointness Constraints
Guo-Qiang Zhang 0001, Lingyun Luo, Chimezie Ogbuji, Cliff A. Joslyn, José L. V. Mejino Jr., Satya Sanket Sahoo |
AMIA | 1 |
| 2011 | OnWARD: Ontology-driven web-based framework for multi-center clinical studiesabstractWith a large percentage of clinical trials still using paper forms as the primary data collection tool, there is much potential for increasing efficiency through web-based data collection systems, especially for large-scale multi-center trials. This paper presents OnWARD, an ontology-driven, secure, rapidly-deployed, web-based framework supporting data capture for large-scale multi-center clinical research. Our approach is developed using the agile methodology to provide a flexible, user-centered dynamic form generator, which can be quickly deployed and customized for any clinical study without the need of deep technical expertise. Because of the flexible framework, the data management system can be extended to accommodate a large variety of data types, including genetic, genomic and proteomic data. In this paper, we demonstrate the initial deployment of OnWARD for a Phase II multi-center clinical trial after a development period of merely three months. The study utilizes 23 clinical report forms containing more than 1500 data points. Preliminary evaluation results show that OnWARD exceeded expectations of the clinical investigators in efficiency, flexibility and ease in setting up. Van Anh Tran, Nathan Johnson, Susan Redline, Guo-Qiang Zhang 0001 |
J. Biomed. Informatics | 4 |
| 2011 | A Fast Iterated Conditional Modes Algorithm for Water-Fat Decomposition in MRIabstractDecomposition of water and fat in magnetic resonance imaging (MRI) is important for biomedical research and clinical applications. In this paper, we propose a two-phased approach for the three-point water-fat decomposition problem. Our contribution consists of two components: 1) a background-masked Markov random field (MRF) energy model to formulate the local smoothness of field inhomogeneity; 2) a new iterated conditional modes (ICM) algorithm accounting for high-performance optimization of the MRF energy model. The MRF energy model is integrated with background masking to prevent error propagation of background estimates as well as improve efficiency. The central component of our new ICM algorithm is the stability tracking (ST) mechanism intended to dynamically track iterative stability on pixels so that computation per iteration is performed only on instable pixels. The ST mechanism significantly improves the efficiency of ICM. We also develop a median-based initialization algorithm to provide good initial guesses for ICM iterations, and an adaptive gradient-based scheme for parametric configuration of the MRF model. We evaluate the robust of our approach with high-resolution mouse datasets acquired from 7T MRI. Fangping Huang, Sreenath Narayan, David L. Wilson, David H. Johnson 0001, Guo-Qiang Zhang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2010 | Segmenting and Merging Domain-specific Ontology Modules for Clinical InformaticsabstractA significant set of challenges to the use of large, source ontologies in the medical domain include: automated translation, customization of source ontologies, and performance issues associated with the use of logical reasoning systems to interpret the meaning of a domain captured in a formal knowledge representation. SNOMED-CT and FMA are two reference ontologies that cover much of the domain of clinical informatics and motivate a better means for re-use. In this paper, we present a method for segmenting and merging modules from these ontologies for a specific domain that preserve the meaning of the anatomy terms they have in common. Chimezie Ogbuji, Sivaram Arabandi, Songmao Zhang, Guo-Qiang Zhang 0001 |
FOIS | 4 |
| 2010 | Large-Scale, Exhaustive Lattice-Based Structural Auditing of SNOMED CT
Guo-Qiang Zhang 0001 |
KSEM | 1 |
| 2010 | Using SPARQL to Test for Lattices: Application to Quality Assurance in Biomedical Ontologies
Guo-Qiang Zhang 0001, Olivier Bodenreider |
ISWC (2) | 1 |
| 2010 | A set coverage problem
Guo-Qiang Zhang 0001, Licong Cui |
Inf. Process. Lett. | 1 |
| 2007 | Bifinite Chu Spaces
Manfred Droste, Guo-Qiang Zhang 0001 |
CALCO | 2 |
| 2007 | Enhancing relevance scoring with chronological term rankabstractWe introduce a new relevance scoring technique that enhances existing relevance scoring schemes with term position information. This technique uses chronological term rank (CTR) which captures the positions of terms as they occur in the sequence of words in a document. CTR is both conceptually and computationally simple when compared to other approaches that use document structure information, such as term proximity, term order and document features. CTR works well when paired with Okapi BM25. We evaluate the performance of various combinations of CTR with Okapi BM25 in order to identify the most effective formula. We then compare the performance of the selected approach against the performance of existing methods such as Okapi BM25, pivoted length normalization and language models. Significant improvements are seen consistently across a variety of TREC data and topic sets, measured by the major retrieval performance metrics. This seems to be the first use of this statistic for relevance scoring. There is likely to be greater retrieval improvements possible using chronological term rank enhanced methods in future work. Adam D. Troy, Guo-Qiang Zhang 0001 |
SIGIR | 2 |
| 2007 | Weakly distributive domains (II)
Guo-Qiang Zhang 0001 |
Frontiers Comput. Sci. China | 2 |
| 2007 | Mediating secure information flow policies
Guo-Qiang Zhang 0001 |
Inf. Comput. | 1 |
| 2006 | Visualization of Remote Hyperspectral Image Data Using Google EarthabstractOne of the greatest obstacles posed in the field of airborne remote sensing is the lack of an expeditious, cross- platform method of viewing imaging data through a standardized web interface. After an ariel photography session, clients or agencies typically wait days or weeks to obtain their images. We have been developing a solution for this issue by creating a wireless network system that allows aircraft to transfer data inflight to a grounded base station using modified, but commercially available, MIMO-based wireless networking hardware. The final goal of this project is to devise a combination of hardware and software systems to deliver images to clients in near-realtime. This paper addresses the web visualization and telemetry monitoring component of the system, the process behind its development, and its path to fruition. The system developed shows not only the feasibility of but also the realization of a software system for transferring and viewing aerial multi- and hyperspectral images during in-flight imaging sessions. Matthew David Crowley, Eric J. Sukalac, Xiuhong Sun, Patrick L. Coronado, Guo-Qiang Zhang 0001 |
IGARSS | 6 |
| 2006 | ARCS: an aggregated related column scoring scheme for aligned sequencesabstractMOTIVATION: Biologists frequently align multiple biological sequences to determine consensus sequences and/or search for predominant residues and conserved regions. Particularly, determining conserved regions in an alignment is one of the most important activities. Since protein sequences are often several-hundred residues or longer, it is difficult to distinguish biologically important conserved regions (motifs or domains) from others. The widely used tools, Logos, Al2co, Confind, and the entropy-based method, often fail to highlight such regions. Thus a computational tool that can highlight biologically important regions accurately will be highly desired. RESULTS: This paper presents a new scoring scheme ARCS (Aggregated Related Column Score) for aligned biological sequences. ARCS method considers not only the traditional character similarity measure but also column correlation. In an extensive experimental evaluation using 533 PROSITE patterns, ARCS is able to highlight the motif regions with up to 77.7% accuracy corresponding to the top three peaks. AVAILABILITY: The source code is available on http://bio.informatics.indiana.edu/projects/arcs and http://goldengate.case.edu/projects/arcs Jeong-Hyeon Choi, Guangyu Chen, Jacek Szymanski, Guo-Qiang Zhang 0001, Anthony K. H. Tung, Jaewoo Kang, Sun Kim, Jiong Yang 0001 |
Bioinform. | 5 |
| 2006 | A Categorical View on Algebraic Lattices in Formal Concept Analysis
Pascal Hitzler, Markus Krötzsch, Guo-Qiang Zhang 0001 |
Fundam. Informaticae | 3 |
| 2005 | On an open problem of Amadio and Curien: The finite antichain condition
Guo-Qiang Zhang 0001 |
Inf. Comput. | 1 |
| 2004 | Reasoning with power defaults
Guo-Qiang Zhang 0001, William C. Rounds |
Theor. Comput. Sci. | 1 |
| 2003 | Compact Coverages Generate Spectral FramesabstractThis note proves a useful characterization of spectral frames: a frame is spectral if and only if it can be generated from a compact coverage relation, where compactness is defined in the usual topological sense. Guo-Qiang Zhang 0001 |
MFPS | 1 |
| 2003 | Chu Spaces, Concept Lattices, and DomainsabstractThis paper serves to bring three independent but important areas of computer science to a common meeting point: Formal Concept Analysis (FCA), Chu Spaces, and Domain Theory (DT). Each area is given a perspective or reformulation that is conducive to the flow of ideas and to the exploration of cross-disciplinary connections. Among other results, we show that the notion of states in Scott’s information system corresponds precisely to that of formal concepts in FCA with respect to all finite Chu spaces, and the entailment relation corresponds to “association rules”. Guo-Qiang Zhang 0001 |
MFPS | 1 |
| 2003 | On transformations of formal power series
Manfred Droste, Guo-Qiang Zhang 0001 |
Inf. Comput. | 2 |
| 2003 | A representation of stably compact spaces, and patch topology
Thierry Coquand, Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 2 |
| 2001 | Rational Transformations of Formal Power Series
Manfred Droste, Guo-Qiang Zhang 0001 |
ICALP | 2 |
| 2001 | Clausal Logic and Logic Programming in Algebraic Domains
William C. Rounds, Guo-Qiang Zhang 0001 |
Inf. Comput. | 2 |
| 2001 | Domains via Graphs
Guo-Qiang Zhang 0001, Yixiang Chen 0001 |
J. Comput. Sci. Technol. | 1 |
| 2000 | Sequents, Frames, and Completeness
Thierry Coquand, Guo-Qiang Zhang 0001 |
CSL | 2 |
| 1999 | Automata, Boolean Matrices, and Ultimate Periodicity
Guo-Qiang Zhang 0001 |
Inf. Comput. | 1 |
| 1997 | Complexity of Power Default ReasoningabstractThis paper derives a new and surprisingly low complexity result for inference in a new form of Reiter's propositional default logic (1980). The problem studied here is the default inference problem whose fundamental importance was pointed out by Kraus, Lehmann, and Magidor (1980). We prove that "normal" default inference, in propositional logic, is a problem complete for co-NP(3), the third level of the Boolean hierarchy. Our result (by changing the underlying semantics) contrasts favorably with a similar result of Gottlob (1992), who proves that standard default inference is II/sub 2//sup P/-complete. Our inference relation also obeys all of the laws for preferential consequence relations set forth by Kraus, Lehmann, and Magidor (1990). In particular we get the property of being able to reason by cases and the law of cautious monotony. Both of these laws fail for standard propositional default logic. The key technique for our results is the use of Scott's domain theory to integrate defaults into partial model theory of the logic, instead of keeping defaults as quasiproof rules in the syntax. In particular, reasoning disjunctively entails using the Smyth powerdomain. Guo-Qiang Zhang 0001, William C. Rounds |
LICS | 1 |
| 1997 | Power Defaults
Guo-Qiang Zhang 0001, William C. Rounds |
LPNMR | 1 |
| 1997 | Parallel Product of Event Structures
Ilaria Castellani, Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 2 |
| 1997 | The End of Pumping?
Guo-Qiang Zhang 0001, E. Rodney Canfield |
Theor. Comput. Sci. | 1 |
| 1997 | Defaults in Domain Theory
Guo-Qiang Zhang 0001, William C. Rounds |
Theor. Comput. Sci. | 1 |
| 1996 | Quasi-Prime Algebraic Domains
Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 1 |
| 1996 | The Largest Cartesian Closed Category of Stable Domains
Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 1 |
| 1995 | Domain Theory Meets Default LogicabstractAbstract We present a development of the theory of default information structures, combining ideas from domain theory with ideas from non-monotonic logic. Conceptually, our treatment is distinguished from standard default logic in that we view default structures as generating models rather than theories. Reiter's default rules are viewed as non-deterministic algorithms for generating preferred partial models. Using domain-theoretical notions, we suggest a robust alternative to the standard definition of extensions in default logic, by introducing the notion of dilation. We prove the existence of such dilations for a new class of default information structures, called the class of rational structures, which properly includes the class of semi-normal structures. William C. Rounds, Guo-Qiang Zhang 0001 |
J. Log. Comput. | 2 |
| 1995 | On Maximal Stable Functions
Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 1 |
| 1994 | A Representation of SFP
Guo-Qiang Zhang 0001 |
Inf. Comput. | 1 |
| 1993 | Universal Quasi-Prime Algebraic Domains
Guo-Qiang Zhang 0001 |
MFPS | 1 |
| 1993 | Some Monoidal Closed Categories of Stable Domains and Event StructuresabstractThis paper introduces the following new constructions on stable domains and event structures: the tensor product; the linear function space; and the exponential. These give rise to a monoidal closed category of dI-domains and to stable event structures, which can be used to interpret intuitionistic linear logic. Finally, the usefulness of the category of stable event structures for modeling concurrency and its relation to other models are discussed. Guo-Qiang Zhang 0001 |
Math. Struct. Comput. Sci. | 1 |
| 1992 | Disjunctive Systems and L-Domains
Guo-Qiang Zhang 0001 |
ICALP | 1 |
| 1992 | dI-Domains as Prime Information Systems
Guo-Qiang Zhang 0001 |
Inf. Comput. | 1 |
| 1992 | Stable Neighbourboods
Guo-Qiang Zhang 0001 |
Theor. Comput. Sci. | 1 |
| 1991 | A Monoidal Closed Category of Event Structures
Guo-Qiang Zhang 0001 |
MFPS | 1 |
| 1989 | DI-Domains as Information Systems (Extended Abstract)
Guo-Qiang Zhang 0001 |
ICALP | 1 |
| 1984 | "NP = P?" and restricted partitions
Guo-Qiang Zhang 0001 |
Inf. Sci. | 1 |