Hua Min

dblp:76/3457 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0003-2422-0043ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Big Data Decision-Making and Racial Disparities: A Case Study Among COVID-19 Inpatient Visits
abstract
The COVID-19 pandemic has had a disproportionate impact on certain racial and ethnic groups, resulting in significant health outcome disparities. The National COVID Cohort Collaborative (N3C) provides a valuable resource for exploring these disparities through big data analytics. This study belongs to a broader work that examines decisions made during data processing and their impact on the analyses performed. Central to our analysis is the introduction of the Continuous Inpatient Encounter (CIE) concept—a novel method we propose for aggregating inpatient visits. By utilizing big data analytics, we aim to identify potential disparities in CIE rates among different racial groups. The results of this study are critical for enhancing the equity of data-driven decision-making in healthcare and for addressing the racial disparities observed in COVID-19 outcomes.
Atefehsadat Haghighathoseini, Janusz Wojtusiak, Nirup M. Menon, Hua Min, Cara Frankenfeld, Timothy Leslie
IEEE Big Data4
2024 PCAO2: an ontology for integration of prostate cancer associated genotypic, phenotypic and lifestyle data
abstract
Disease ontologies facilitate the semantic organization and representation of domain-specific knowledge. In the case of prostate cancer (PCa), large volumes of research results and clinical data have been accumulated and needed to be standardized for sharing and translational researches. A formal representation of PCa-associated knowledge will be essential to the diverse data standardization, data sharing and the future knowledge graph extraction, deep phenotyping and explainable artificial intelligence developing. In this study, we constructed an updated PCa ontology (PCAO2) based on the ontology development life cycle. An online information retrieval system was designed to ensure the usability of the ontology. The PCAO2 with a subclass-based taxonomic hierarchy covers the major biomedical concepts for PCa-associated genotypic, phenotypic and lifestyle data. The current version of the PCAO2 contains 633 concepts organized under three biomedical viewpoints, namely, epidemiology, diagnosis and treatment. These concepts are enriched by the addition of definition, synonym, relationship and reference. For the precision diagnosis and treatment, the PCa-associated genes and lifestyles are integrated in the viewpoint of epidemiological aspects of PCa. PCAO2 provides a standardized and systematized semantic framework for studying large amounts of heterogeneous PCa data and knowledge, which can be further, edited and enriched by the scientific community. The PCAO2 is freely available at https://bioportal.bioontology.org/ontologies/PCAO, http://pcaontology.net/ and http://pcaontology.net/mobile/.
Chunjiang Yu, Hui Zong, Yalan Chen, Yibin Zhou, Xingyun Liu, Xiaonan Zheng, Hua Min, Bairong Shen
Briefings Bioinform.9
2022 Keyphrase Identification with a Limited Labeled Dataset Using Deep Active Learning and Domain Adaptation
Rohan Goli, Nina C. Hubig, Hua Min, Yang Gong, Dean F. Sittig, Paul G. Biondich, Adam Wright, Christian Nøhr, Timothy Law 0001, Arild Faxvaag, Ronald W. Gimbel, Lior Rennert, Xia Jing
AMIA3
2021 A clinical decision support system (CDSS) ontology to facilitate portable vaccination CDSS rules: preliminary results
Xia Jing, Hua Min, Yang Gong, James J. Cimino, Dean F. Sittig, Paul G. Biondich, Adam Wright, Christian Nøhr, Timothy Law 0001, Arild Faxvaag, Akash Indani, Nina C. Hubig, Ronald W. Gimbel, Lior Rennert
AMIA2
2021 Fear of Missing Out (FOMO) Toward ICT Use During Public Health Emergencies: An Investigation on Predictors and Outcomes
abstract
COVID-19 has brought a great impact on people's lives around the world. This paper aims to study the influencing factors of people's fear of missing out (FOMO) toward personal ICT use and its further impact on life satisfaction during the pandemic. A sample consisting of 318 participants was obtained by an online survey in China. Partial least squares structural equation modeling (PLS-SEM) was used for data analysis. The results suggested that people's anxiety and boredom brought by the pandemic are positively correlated with their FOMO. People with higher FOMO used personal ICTs more frequently for both social and process purposes. Furthermore, the social use of ICTs promoted people's life satisfaction, while the process use of ICTs had no significant effect on life satisfaction. Several theoretical and practical implications were discussed based on the results.
Xiaokang Song, Shijie Song 0002, Yuxiang Zhao 0001, Hua Min, Qinghua Zhu 0002
J. Database Manag.4
2019 Synthetic Data for Teaching Data Integration in Informatics Graduate Program
Hedyeh Mobahi, Hua Min, Janusz Wojtusiak
AMIA2
2017 Discovering additional complex NCIt gene concepts with high error rate
abstract
The Gene hierarchy of the National Cancer Institute (NCI) Thesaurus (NCIt) is of high priority for NCI. It is important to have quality assurance (QA) techniques to improve its content quality. We present a two-step methodology concentrating on auditing the modeling of complex concepts, which are shown to have a higher error rate compared to control concepts. In the first step, we test whether concepts that appear complex in a so called “partial-area taxonomy” have a higher error rate than control concepts. In the second step, we introduce an innovative technique based on a “partial-area sub-taxonomy” (constructed with a subset of roles) to discover additional complex concepts. The results of the QA study show that these concepts are indeed statistically significantly more likely to have more errors than control concepts. This makes it easier for NCI staff to improve the modeling quality of gene concepts in NCIt.
Hua Min, Yehoshua Perl, James Geller
BIBM2
2016 Applying Machine Learning Methods to Predict Activities of Daily Living for Cancer Patients
Hua Min, Talha Oz, Sava Vukomanovic, Hedyeh Mobahi, Katherine Irvin, Ilirjeta Krasniqi, Janusz Wojtusiak
AMIA1
2016 Visualizing the Effects of Cancers on Relationships Between Comorbidities and Activities of Daily Living
Hua Min, Talha Oz, Sava Vukomanovic, Hedyeh Mobahi, Katherine Irvin, Ilirjeta Krasniqi, Janusz Wojtusiak
AMIA1
2015 Scalable quality assurance for large SNOMED CT hierarchies using subject-based subtaxonomies
abstract
OBJECTIVE: Standards terminologies may be large and complex, making their quality assurance challenging. Some terminology quality assurance (TQA) methodologies are based on abstraction networks (AbNs), compact terminology summaries. We have tested AbNs and the performance of related TQA methodologies on small terminology hierarchies. However, some standards terminologies, for example, SNOMED, are composed of very large hierarchies. Scaling AbN TQA techniques to such hierarchies poses a significant challenge. We present a scalable subject-based approach for AbN TQA. METHODS: An innovative technique is presented for scaling TQA by creating a new kind of subject-based AbN called a subtaxonomy for large hierarchies. New hypotheses about concentrations of erroneous concepts within the AbN are introduced to guide scalable TQA. RESULTS: We test the TQA methodology for a subject-based subtaxonomy for the Bleeding subhierarchy in SNOMED's large Clinical finding hierarchy. To test the error concentration hypotheses, three domain experts reviewed a sample of 300 concepts. A consensus-based evaluation identified 87 erroneous concepts. The subtaxonomy-based TQA methodology was shown to uncover statistically significantly more erroneous concepts when compared to a control sample. DISCUSSION: The scalability of TQA methodologies is a challenge for large standards systems like SNOMED. We demonstrated innovative subject-based TQA techniques by identifying groups of concepts with a higher likelihood of having errors within the subtaxonomy. Scalability is achieved by reviewing a large hierarchy by subject. CONCLUSIONS: An innovative methodology for scaling the derivation of AbNs and a TQA methodology was shown to perform successfully for the largest hierarchy of SNOMED.
Christopher Ochs, James Geller, Yehoshua Perl, Yan Chen 0009, Junchuan Xu, Hua Min, James T. Case, Zhi Wei 0001
J. Am. Medical Informatics Assoc.6
2014 Sharing behavioral data through a grid infrastructure using data standards
abstract
OBJECTIVE: In an effort to standardize behavioral measures and their data representation, the present study develops a methodology for incorporating measures found in the National Cancer Institute's (NCI) grid-enabled measures (GEM) portal, a repository for behavioral and social measures, into the cancer data standards registry and repository (caDSR). METHODS: The methodology consists of four parts for curating GEM measures into the caDSR: (1) develop unified modeling language (UML) models for behavioral measures; (2) create common data elements (CDE) for UML components; (3) bind CDE with concepts from the NCI thesaurus; and (4) register CDE in the caDSR. RESULTS: UML models have been developed for four GEM measures, which have been registered in the caDSR as CDE. New behavioral concepts related to these measures have been created and incorporated into the NCI thesaurus. Best practices for representing measures using UML models have been utilized in the practice (eg, caDSR). One dataset based on a GEM-curated measure is available for use by other systems and users connected to the grid. CONCLUSIONS: Behavioral and population science data can be standardized by using and extending current standards. A new branch of CDE for behavioral science was developed for the caDSR. It expands the caDSR domain coverage beyond the clinical and biological areas. In addition, missing terms and concepts specific to the behavioral measures addressed in this paper were added to the NCI thesaurus. A methodology was developed and refined for curation of behavioral and population science data.
Hua Min, Riki Ohira, Michael A. Collins, Jessica Bondy, Nancy E. Avis, Olga Tchuvatkina, Paul Courtney, Richard P. Moser, Abdul R. Shaikh, Bradford W. Hesse, Mary Cooper, Dianne M. Reeves, Bob Lanese, Cindy Helba, Suzanne M. Miller, Eric A. Ross
J. Am. Medical Informatics Assoc.1
2009 Integration of prostate cancer clinical data using an ontology
Hua Min, Frank J. Manion, Elizabeth Goralczyk, Yu-Ning Wong, Eric A. Ross, J. Robert Beck
J. Biomed. Informatics1
2008 Automated comparative auditing of NCIT genomic roles using NCBI
Barry Cohen, Marc Oren, Hua Min, Yehoshua Perl, Michael Halper
J. Biomed. Informatics3
2007 Analysis of Error Concentrations in SNOMED
Michael Halper, Yue Wang 0033, Hua Min, Yan Chen 0009, George Hripcsak, Yehoshua Perl, Kent A. Spackman
AMIA3
2007 Structural methodologies for auditing SNOMED
Yue Wang 0033, Michael Halper, Hua Min, Yehoshua Perl, Yan Chen 0009, Kent A. Spackman
J. Biomed. Informatics3
2006 Research Paper: Auditing as Part of the Terminology Design Life Cycle
abstract
OBJECTIVE: To develop and test an auditing methodology for detecting errors in medical terminologies satisfying systematic inheritance. This methodology is based on various abstraction taxonomies that provide high-level views of a terminology and highlight potentially erroneous concepts. DESIGN: Our auditing methodology is based on dividing concepts of a terminology into smaller, more manageable units. First, we divide the terminology's concepts into areas according to their relationships/roles. Then each multi-rooted area is further divided into partial-areas (p-areas) that are singly-rooted. Each p-area contains a set of structurally and semantically uniform concepts. Two kinds of abstraction networks, called the area taxonomy and p-area taxonomy, are derived. These taxonomies form the basis for the auditing approach. Taxonomies tend to highlight potentially erroneous concepts in areas and p-areas. Human reviewers can focus their auditing efforts on the limited number of problematic concepts following two hypotheses on the probable concentration of errors. RESULTS: A sample of the area taxonomy and p-area taxonomy for the Biological Process (BP) hierarchy of the National Cancer Institute Thesaurus (NCIT) was derived from the application of our methodology to its concepts. These views led to the detection of a number of different kinds of errors that are reported, and to confirmation of the hypotheses on error concentration in this hierarchy. CONCLUSION: Our auditing methodology based on area and p-area taxonomies is an efficient tool for detecting errors in terminologies satisfying systematic inheritance of roles, and thus facilitates their maintenance. This methodology concentrates a domain expert's manual review on portions of the concepts with a high likelihood of errors.
Hua Min, Yehoshua Perl, Yan Chen 0009, Michael Halper, James Geller, Yue Wang 0033
J. Am. Medical Informatics Assoc.1
2004 Auditing concept categorizations in the UMLS
Huanying Gu, Yehoshua Perl, Gai Elhanan, Hua Min, Li Zhang 0048
Artif. Intell. Medicine4
2003 Consistency across the hierarchies of the UMLS Semantic Network and Metathesaurus
James J. Cimino, Hua Min, Yehoshua Perl
J. Biomed. Informatics2
2002 Using the metaschema to audit UMLS classification errors
Huanying Gu, Hua Min, Li Zhang 0048, Yehoshua Perl
AMIA2