VLDB 2026 Research / reviewers in the wild / expert
Qing T. Zeng
dblp:26/4956 · also Qing Zeng-Treitler
· DBLP profile ↗
93ranked-venue papers
20as first author
5since 2021 · last 2024
0000-0002-8353-7473ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 92 · 20 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | More than the Sum of its Parts: Applying Topic Modeling and Explainable AI to Deep Learning in Understanding Problematic Opioid UseabstractProblematic opioid use is a major crisis, especially among Veterans of the U.S. Military. In previous work we developed a natural language processing tool to identify problematic opioid use in Veterans Affairs clinical notes. In this work, we developed an application that identifies clinical notes associated with patients who were documented as experiencing problematic opioid use only in clinical notes, as compared to patients who had received a relevant ICD code for problematic opioid use. The application used topic and topic word output from topic modeling to train deep neural network models. The models performed well, achieving area under the curve values of 82% or above, exceeding the performance of baseline models using only topics or topic words. We also computed impact scores, an explainable artificial intelligence method that identifies critical features in training deep learning models. The impact scores extended additional understanding to the data and outcomes. Terri Elizabeth Workman, Joel Kupersmith, Qing T. Zeng |
IEEE Big Data | 4 |
| 2022 | Medication-Wide Association Study Plus (MWAS+): A New Approach to Drug Repurposing
Yan Cheng 0004, Ali Ahmed 0008, Edward Zamrini, Qing T. Zeng |
AMIA | 5 |
| 2022 | A Novel Hybrid Value-Aware Transformer Architecture for Learning from Longitudinal Clinical Data
Yan Cheng 0004, Stuart J. Nelson, Peter Kokkinos, Edward Zamrini, Ali Ahmed 0008, Qing T. Zeng |
AMIA | 7 |
| 2021 | Enhancing Clinical Data Analysis by Explaining Interaction Effects Between Covariates in Deep Neural Network Models
Ali Ahmed 0008, Edward Zamrini, Yan Cheng 0004, Qing T. Zeng |
AMIA | 5 |
| 2021 | Healthy Lifestyle and Mood: A Biomedical Informatics Citizen Science Project in a High School Classroom
Jennifer Ushe, Doug Redd, Scarlyn Gutierrez Nunez, Eduardo Trujillo-Rivera, Senait Tekle, Stuart J. Nelson, Qing T. Zeng |
AMIA | 7 |
| 2020 | Clinical Sublanguage Trend and Usage Analysis from a Large Clinical CorpusabstractThe field of clinical natural language processing (NLP) has been built on the analysis of clinical sublanguage characteristics. It is well recognized that not only does clinical sublanguage differ from general English (or other languages) but also clinical sublanguage differs among clinical subspecialties and among corpora originating from different healthcare systems. A less recognized aspect is that clinical sublanguage, like all languages, evolves over time. This paper analyses the evolution of clinical sublanguage using a large, national clinical text corpus spanning 15 years. Through the analyses of document types, length, ngrams, and concepts, we found strong evidence that clinical sublanguage does evolve and such changes have implications for NLP development and maintenance. Although the analysis is performed on one corpus, our observations of sublanguage changes are generalizable. Guy Divita, Terri Elizabeth Workman, Doug Redd, Jennifer H. Garvin, Qing T. Zeng |
IEEE BigData | 6 |
| 2020 | A Prototype Application to Identify LGBT Patients in Clinical NotesabstractLGBT Patients bear a disproportional burden of health disparities. Sexual orientation and gender identity is of clinical relevance to healthcare providers and data scientists. However, little work has been done to identify LGBT patients, especially in data derived from electronic health record notes. We developed a prototype application that leverages machine learning and rule-based pattern matching methods to identify LGBT patients in a large data source, Veterans Health Administration electronic health record notes. This application achieved 88.2% sensitivity, 91.5% specificity, and 85.9% positive predictive value in a binary classification task for three random document test sets. This work has implications in both improved healthcare and data research. Terri Elizabeth Workman, Joseph L. Goulet, Cynthia Brandt, Melissa Skanderson, Allison R. Warren, Jacob Eleazer, Kirsha Gordon, Qing T. Zeng |
IEEE BigData | 9 |
| 2020 | A Proficient Spelling Analysis Method Applied to Herbal and Dietary Supplement Discovery in a Large Clinical CorpusabstractIrregular spellings in clinical free text present challenges to natural language processing. A number of spelling correction tools exist, but automated spelling correction of clinical text is not a routine practice due to the risk of introducing new errors. We developed a novel spelling analysis application that combines Word2Vec and Levenshtein Edit Distance Constraints to identify variant forms of words. The use case applied to this study was that of discovering herbal and dietary supplements that interact with prescription medications in clinical text. The prototype application processed a large corpus (approximately 1.6 million records), achieving a positive predictive value of 0.9322, in identifying spelling variants, outperforming two baseline methods that achieved positive predictive values of 0.0348 and 0.0067. Our findings suggest that this prototype application provides a more efficient method for researchers and clinicians to find valid misspellings of terms in clinical text. Terri Elizabeth Workman, Guy Divita, Qing T. Zeng |
IEEE BigData | 4 |
| 2019 | 2018 Salary Survey of AMIA Members: Factors Associated with Higher Salaries
Yan Cheng 0004, April F. Mohanty, Omolola Ogunyemi, Catherine Arnott Smith, Gondy Leroy, Qing T. Zeng |
AMIA | 6 |
| 2019 | Thriving in Your Biomedical Informatics Career While Balancing Work, Personal, and Family Life
Lesley Clack, Qing T. Zeng, Gretchen Purcell Jackson, William R. Hersh, April F. Mohanty |
AMIA | 2 |
| 2019 | Preliminary chart review for Natural Language Processing development among electronic health records with documentation of marijuana use
Termeh Feinberg, Joseph L. Goulet, Amy Justice, Samah Jamal Fodeh, Lori Bastian, Qing T. Zeng, Cynthia Brandt |
AMIA | 6 |
| 2019 | Interpretability and Statistical Inferences of the Prediction of Clinical Outcomes Using Deep Neural Networks
Qing T. Zeng, Cecilia Dao, Orna Intrator, Joseph L. Goulet |
AMIA | 1 |
| 2019 | Discovering Sublanguages in a Large Clinical Corpus through Unsupervised Machine Learning and Information GainabstractSublanguages are domain-centered subsets of general or colloquial language. Their identification drives several language analysis tasks, but it is difficult to discern separate sublanguages in large clinical corpora. We applied k-means clustering of semantic properties, and a novel implementation of relative entropy as an information gain indicator, to identify sublanguages within a large clinical corpus (~1.6 million documents), visualizing the results in a heat map. Patterns both within and across clusters reveal sublanguage trends. These findings are significant in sublanguage analysis, and have implications on both regional and international levels. Terri Elizabeth Workman, Guy Divita, Qing T. Zeng |
IEEE BigData | 3 |
| 2019 | Explainable Deep Learning Applied to Understanding Opioid Use Disorder and Its Risk FactorsabstractOpioid Use Disorder is an international crisis, affecting many populations. Deep learning models can potentially predict opioid use disorder, but provide little insight to how predictions are derived. Impact scores, a new development in explainable artificial intelligence, measure how individual features affect deep learning outcomes. We modeled clinical visits to predict opioid use disorder, computed impact scores, and compared them to odds log ratios from logistic regression. Impact scores were generally comparable to odds log ratios, in providing insight to opioid abuse risk, but from a better-performing method than logistic regression. Terri Elizabeth Workman, Qing T. Zeng, Joel Kupersmith, Friedhelm Sandbrink, Joseph L. Goulet, Nawar M. Shaar, Christopher Spevak, Cynthia Brandt, Marc R. Blackman |
IEEE BigData | 2 |
| 2018 | Clinical Text Classification with Word Embedding Features vs. Bag-of-Words FeaturesabstractWord embedding motivated by deep learning have shown promising results over traditional bag-of-words features for natural language processing. When trained on large text corpora, word embedding methods such as word2vec and doc2vec methods have the advantage of learning from unlabeled data and reduce the dimension of the feature space. In this study, we experimented with word2vec and doc2vec features for a set of clinical text classification tasks and compared the results with using the traditional bag-of-words (BOW) features. The study showed that the word2vec features performed better than the BOW-1-gram features. However, when 2-grams were added to BOW, comparison results were mixed. Stephanie Taylor, Nell J. Marshall, Craig A. Morioka, Qing T. Zeng |
IEEE BigData | 5 |
| 2018 | A Novel Deep Learning Pipeline to Analyze Temporal Clinical DataabstractAnalysis of clinical temporal data can be difficult due to natural properties that often characterize it. The large number of variables, missing values, and other characteristics lead to issues of sparsity and high dimensional complexity. We hypothesized that a pipeline application implementing relevant deep learning methods could sequentially address these difficulties, demonstrating their combined utility in a classification task. We implemented Word2Vec, t-distributed stochastic neighbor embedding, and a convolutional neural network in a pipeline application. To test the pipeline, we applied it to a simple, binary classification task to identify patient encounter care setting. In preliminary testing, the pipeline application achieved 92% accuracy. It also produced temporal data cubes indicative of clinical encounters in intensive care unit (ICU) and non-ICU care settings. A deep learning pipeline process combining multiple methods holds promise in improving analytical tasks of clinical temporal data. Terri Elizabeth Workman, Michael Hirezi, Eduardo Trujillo-Rivera, Anita K. Patel, Julia A. Heneghan, James E. Bost, Qing T. Zeng, Murray Pollack |
IEEE BigData | 7 |
| 2017 | Doodle Health: A Crowdsourcing Game for the Co-design and Testing of Pictographs to Reduce Disparities in Healthcare Communication
Carrie M. Christensen, Doug Redd, Erica Lake, Jean Shipman, Heather Aiono, Roger Altizer, Bruce E. Bray, Qing T. Zeng |
AMIA | 8 |
| 2017 | Physician Conception of Patient Frailty in Cardiac Care Decisions
Kristina Doing-Harris, Rashmee U. Shah, Bruce E. Bray, Qing T. Zeng, Jennifer H. Garvin, Charlene R. Weir |
AMIA | 5 |
| 2016 | Improving Pain Assessment in Medical Intensive Care Unit Through Natural Language Processing
Doug Redd, Qing T. Zeng, Cynthia Brandt, Kathleen Akgün |
AMIA | 2 |
| 2016 | Identification and Use of Frailty Indicators from Text to Examine Associations with Clinical Outcomes Among Patients with Heart Failure
April F. Mohanty, Ali Ahmed 0008, Charlene R. Weir, Bruce E. Bray, Rashmee U. Shah, Doug Redd, Qing T. Zeng |
AMIA | 8 |
| 2016 | Automated pictographic illustration of discharge instructions with Glyph: impact on patient recall and satisfactionabstractOBJECTIVES: First, to evaluate the effect of standard vs pictograph-enhanced discharge instructions on patients' immediate and delayed recall of and satisfaction with their discharge instructions. Second, to evaluate the effect of automated pictograph enhancement on patient satisfaction with their discharge instructions. MATERIALS AND METHODS: Glyph, an automated healthcare informatics system, was used to automatically enhance patient discharge instructions with pictographs. Glyph was developed at the University of Utah by our research team. Patients in a cardiovascular medical unit were randomized to receive pictograph-enhanced or standard discharge instructions. Measures of immediate and delayed recall and satisfaction with discharge instructions were compared between two randomized groups: pictograph (n = 71) and standard (n = 73). RESULTS: Study participants who received pictograph-enhanced discharge instructions recalled 35% more of their instructions at discharge than those who received standard discharge instructions. The ratio of instructions at discharge was: standard = 0.04 ± 0.03 and pictograph-enhanced = 0.06 ± 0.03. The ratio of instructions at 1 week post discharge was: standard = 0.04 ± 0.02 and pictograph-enhanced 0.04 ± 0.02. Additionally, study participants who received pictograph-enhanced discharge instructions were more satisfied with the understandability of their instructions at 1 week post-discharge than those who received standard discharge instructions. DISCUSSION: Pictograph-enhanced discharge instructions have the potential to increase patient understanding of and satisfaction with discharge instructions. CONCLUSION: It is feasible to automatically illustrate discharge instructions and provide them to patients in a timely manner without interfering with clinical work. Illustrations in discharge instructions were found to improve patients' short-term recall of discharge instructions and delayed satisfaction (1-week post hospitalization) with the instructions. Therefore, it is likely that patients' understanding of and interaction with their discharge instructions is improved by the addition of illustrations. Brent Hill, Seneca I. Perri, Jinqiu Kuang, Bruce E. Bray, Long H. Ngo, Alexa K. Doig, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 7 |
| 2016 | Assessing the readability of ClinicalTrials.govabstractOBJECTIVE: ClinicalTrials.gov serves critical functions of disseminating trial information to the public and helping the trials recruit participants. This study assessed the readability of trial descriptions at ClinicalTrials.gov using multiple quantitative measures. MATERIALS AND METHODS: The analysis included all 165,988 trials registered at ClinicalTrials.gov as of April 30, 2014. To obtain benchmarks, the authors also analyzed 2 other medical corpora: (1) all 955 Health Topics articles from MedlinePlus and (2) a random sample of 100,000 clinician notes retrieved from an electronic health records system intended for conveying internal communication among medical professionals. The authors characterized each of the corpora using 4 surface metrics, and then applied 5 different scoring algorithms to assess their readability. The authors hypothesized that clinician notes would be most difficult to read, followed by trial descriptions and MedlinePlus Health Topics articles. RESULTS: Trial descriptions have the longest average sentence length (26.1 words) across all corpora; 65% of their words used are not covered by a basic medical English dictionary. In comparison, average sentence length of MedlinePlus Health Topics articles is 61% shorter, vocabulary size is 95% smaller, and dictionary coverage is 46% higher. All 5 scoring algorithms consistently rated CliniclTrials.gov trial descriptions the most difficult corpus to read, even harder than clinician notes. On average, it requires 18 years of education to properly understand these trial descriptions according to the results generated by the readability assessment algorithms. DISCUSSION AND CONCLUSION: Trial descriptions at CliniclTrials.gov are extremely difficult to read. Significant work is warranted to improve their readability in order to achieve CliniclTrials.gov's goal of facilitating information dissemination and subject recruitment. Danny T. Y. Wu, David A. Hanauer, Qiaozhu Mei, Patricia M. Clark, Lawrence C. An, Joshua Proulx, Qing T. Zeng, V. G. Vinod Vydiswaran, Kevyn Collins-Thompson, Kai Zheng 0002 |
J. Am. Medical Informatics Assoc. | 7 |
| 2016 | Mining Big Data in biomedicine and health care
Samah Jamal Fodeh, Qing T. Zeng |
J. Biomed. Informatics | 2 |
| 2015 | Using social media data to analyze patient satisfaction of health care facilities
Katherine Doyon, Qing T. Zeng, Rebecca Morris, Catherine Arnott Smith |
AMIA | 2 |
| 2015 | Discharge Instructions: What Do Patients Remember?
Brent Hill, Qing T. Zeng |
AMIA | 2 |
| 2015 | Representation of Functional Status Concepts from Clinical Documents and Social Media Sources by Standard Terminologies
Jinqiu Kuang, April F. Mohanty, Rashmi V. H., Charlene R. Weir, Bruce E. Bray, Qing T. Zeng |
AMIA | 6 |
| 2015 | Ginkgo and Warfarin Interaction in a Large Veterans Administration Population
Gregory J. Stoddard, Melissa Archer, Laura Shane-McWhorter, Bruce E. Bray, Doug Redd, Joshua Proulx, Qing T. Zeng |
AMIA | 7 |
| 2015 | Regular expression-based learning to extract bodyweight values from clinical notes
Maureen A. Murtaugh, Bryan Smith Gibson, Doug Redd, Qing T. Zeng |
J. Biomed. Informatics | 4 |
| 2014 | Development of an Alert System to Detect Drug Interactions with Herbal Supplements using Medical Record Data
Melissa Archer, Joshua Proulx, Laura Shane-McWhorter, Bruce E. Bray, Qing T. Zeng |
AMIA | 5 |
| 2014 | Crowdsourcing and Development of Health-related Pictographs for Minority Groups by Gaming - A Focus Group Study
Carrie M. Christensen, Qing T. Zeng, Seneca I. Perri, Erica Lake, Bruce E. Bray, Heather Aiono, Marty Malheiro |
AMIA | 2 |
| 2014 | What women want? Expressing women's voice on contraception
Kavitha Damal, Rebecca Morris, Qing T. Zeng |
AMIA | 3 |
| 2014 | v3NLP Marshallers: Providing NLP Workflow Interoperability
Guy Divita, Brian R. Ivie, Qing T. Zeng |
AMIA | 3 |
| 2014 | Sophia: An Expedient UMLS Concept Extraction Annotator
Guy Divita, Qing T. Zeng, Adi V. Gundlapalli, Scott L. DuVall, Jonathan R. Nebeker, Matthew H. Samore |
AMIA | 2 |
| 2014 | Automatically Enhancing Discharge Instructions with Pictographs to Improve Patient Recall and Satisfaction
Brent Hill, Seneca I. Perri, Jinqiu Kuang, Rebecca Morris, Katherine Doyon, Bruce E. Bray, Qing T. Zeng |
AMIA | 7 |
| 2014 | Differences in Nationwide Cohorts of Acupuncture Users Identified Using Structured and Free Text Medical Records
Doug Redd, Qing T. Zeng |
AMIA | 2 |
| 2014 | Research and applications: Learning regular expressions for clinical text classificationabstractOBJECTIVES: Natural language processing (NLP) applications typically use regular expressions that have been developed manually by human experts. Our goal is to automate both the creation and utilization of regular expressions in text classification. METHODS: We designed a novel regular expression discovery (RED) algorithm and implemented two text classifiers based on RED. The RED+ALIGN classifier combines RED with an alignment algorithm, and RED+SVM combines RED with a support vector machine (SVM) classifier. Two clinical datasets were used for testing and evaluation: the SMOKE dataset, containing 1091 text snippets describing smoking status; and the PAIN dataset, containing 702 snippets describing pain status. We performed 10-fold cross-validation to calculate accuracy, precision, recall, and F-measure metrics. In the evaluation, an SVM classifier was trained as the control. RESULTS: The two RED classifiers achieved 80.9-83.0% in overall accuracy on the two datasets, which is 1.3-3% higher than SVM's accuracy (p<0.001). Similarly, small but consistent improvements have been observed in precision, recall, and F-measure when RED classifiers are compared with SVM alone. More significantly, RED+ALIGN correctly classified many instances that were misclassified by the SVM classifier (8.1-10.3% of the total instances and 43.8-53.0% of SVM's misclassifications). CONCLUSIONS: Machine-generated regular expressions can be effectively used in clinical text classification. The regular expression-based classifier can be combined with other classifiers, like SVM, to improve classification performance. Duy Duc An Bui, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Evaluation of a pictograph enhancement system for patient instruction: a recall studyabstractOBJECTIVE: We developed a novel computer application called Glyph that automatically converts text to sets of illustrations using natural language processing and computer graphics techniques to provide high quality pictographs for health communication. In this study, we evaluated the ability of the Glyph system to illustrate a set of actual patient instructions, and tested patient recall of the original and Glyph illustrated instructions. METHODS: We used Glyph to illustrate 49 patient instructions representing 10 different discharge templates from the University of Utah Cardiology Service. 84 participants were recruited through convenience sampling. To test the recall of illustrated versus non-illustrated instructions, participants were asked to review and then recall a set questionnaires that contained five pictograph-enhanced and five non-pictograph-enhanced items. RESULTS: The mean score without pictographs was 0.47 (SD 0.23), or 47% recall. With pictographs, this mean score increased to 0.52 (SD 0.22), or 52% recall. In a multivariable mixed effects linear regression model, this 0.05 mean increase was statistically significant (95% CI 0.03 to 0.06, p<0.001). DISCUSSION: In our study, the presence of Glyph pictographs improved discharge instruction recall (p<0.001). Education, age, and English as first language were associated with better instruction recall and transcription. CONCLUSIONS: Automated illustration is a novel approach to improve the comprehension and recall of discharge instructions. Our results showed a statistically significant in recall with automated illustrations. Subjects with no-colleague education and younger subjects appeared to benefit more from the illustrations than others. Qing T. Zeng, Seneca I. Perri, Carlos Nakamura, Jinqiu Kuang, Brent Hill, Duy Duc An Bui, Gregory J. Stoddard, Bruce E. Bray |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Improving Discharge Instructions: Perspectives from Providers and Patients
Brent Hill, Qing T. Zeng, Seneca I. Perri, Seraphine Kapsandoy, Jinqiu Kuang |
AMIA | 2 |
| 2013 | Personalized Shared Decision Making Model Support via Summary Statistics and Patient Stories
Joshua Proulx, Brent Hill, Qing T. Zeng |
AMIA | 3 |
| 2013 | Multi-Layered Annotation of Template Elements in VA Clinical Text
Shuying Shen, Tyler Forbush, Guy Divita, Dezon Finch, Qing T. Zeng |
AMIA | 5 |
| 2013 | Recall of Computer-Illustrated Patient Instructions: An Evaluation of the GLYPH System
Qing T. Zeng, Seneca I. Perri, Brent Hill, Duy Duc An Bui, Carlos Nakamura |
AMIA | 1 |
| 2012 | Automated Illustration of Patients Instructions
Duy Duc An Bui, Carlos Nakamura, Bruce E. Bray, Qing T. Zeng |
AMIA | 4 |
| 2012 | Not So Familiar Quotations: Content of Quoted Strings in Clinic Notes
Brent Hill, Qing T. Zeng, Doug Redd |
AMIA | 2 |
| 2012 | Synonym, Topic Model and Predicate-Based Query Expansion for Retrieving Clinical Documents
Qing T. Zeng, Doug Redd, Thomas C. Rindflesch, Jonathan R. Nebeker |
AMIA | 1 |
| 2012 | A taxonomy of representation strategies in iconic communication
Carlos Nakamura, Qing T. Zeng |
Int. J. Hum. Comput. Stud. | 2 |
| 2012 | Active learning for clinical text classification: is it better than random sampling?abstractOBJECTIVE: This study explores active learning algorithms as a way to reduce the requirements for large training sets in medical text classification tasks. DESIGN: Three existing active learning algorithms (distance-based (DIST), diversity-based (DIV), and a combination of both (CMB)) were used to classify text from five datasets. The performance of these algorithms was compared to that of passive learning on the five datasets. We then conducted a novel investigation of the interaction between dataset characteristics and the performance results. MEASUREMENTS: Classification accuracy and area under receiver operating characteristics (ROC) curves for each algorithm at different sample sizes were generated. The performance of active learning algorithms was compared with that of passive learning using a weighted mean of paired differences. To determine why the performance varies on different datasets, we measured the diversity and uncertainty of each dataset using relative entropy and correlated the results with the performance differences. RESULTS: The DIST and CMB algorithms performed better than passive learning. With a statistical significance level set at 0.05, DIST outperformed passive learning in all five datasets, while CMB was found to be better than passive learning in four datasets. We found strong correlations between the dataset diversity and the DIV performance, as well as the dataset uncertainty and the performance of the DIST algorithm. CONCLUSION: For medical text classification, appropriate active learning algorithms can yield performance comparable to that of passive learning with considerably smaller training sets. In particular, our results suggest that DIV performs better on data with higher diversity and DIST on data with lower uncertainty. Rosa L. Figueroa, Qing T. Zeng, Long H. Ngo, Sergey Goryachev, Eduardo P. Wiechmann |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | A bootstrapping algorithm to improve cohort identification using structured dataabstractCohort identification is an important step in conducting clinical research studies. Use of ICD-9 codes to identify disease cohorts is a common approach that can yield satisfactory results in certain conditions; however, for many use-cases more accurate methods are required. In this study, we propose a bootstrapping method that supplements ICD-9 codes with lab results, medications, etc. to build classification models that can be used to identify cohorts more accurately. The proposed method does not require prior information about the true class of the patients. We used the method to identify Diabetes Mellitus (DM) and Hyperlipidemia (HL) patient cohorts from a database of 800 thousand patients. Evaluation results show that the method identified 11,000 patients who did not have DM related ICD-9 codes as positive for DM and 52,000 patients without HL codes as positive for HL. A review of 400 patient charts (200 patients for each condition) by two clinicians shows that in both the conditions studied, the labeling assigned by the proposed approach is more consistent with that of the clinicians compared to labeling through ICD-9 codes. The method is reasonably automated and, we believe, holds potential for inexpensive, more accurate cohort identification. Sasikiran Kandula, Qing T. Zeng, Lingji Chen, William L. Salomon, Bruce E. Bray |
J. Biomed. Informatics | 2 |
| 2009 | Tailoring Vocabularies for NLP in Sub-Domains: A Method to Detect Unused Word Sense
Rosa L. Figueroa, Qing T. Zeng, Sergey Goryachev, Eduardo P. Wiechmann |
AMIA | 2 |
| 2008 | Identification and Extraction of Family History Information from Clinical Reports
Sergey Goryachev, Hyeon-Eui Kim, Qing T. Zeng |
AMIA | 3 |
| 2008 | Creating a Gold Standard for the Readability Measurement of Health Texts
Sasikiran Kandula, Qing T. Zeng |
AMIA | 2 |
| 2008 | Improving Patient Comprehension and Recall of Discharge Instructions by Supplementing Free Texts with Pictographs
Qing T. Zeng, Hyeon-Eui Kim, Martha Hunter |
AMIA | 1 |
| 2008 | Research Paper: Consumer Health Concepts That Do Not Map to the UMLS: Where Do They Fit?abstractOBJECTIVE: This study has two objectives: first, to identify and characterize consumer health terms not found in the Unified Medical Language System (UMLS) Metathesaurus (2007 AB); second, to describe the procedure for creating new concepts in the process of building a consumer health vocabulary. How do the unmapped consumer health concepts relate to the existing UMLS concepts? What is the place of these new concepts in professional medical discourse? DESIGN: The consumer health terms were extracted from two large corpora derived in the process of Open Access Collaboratory Consumer Health Vocabulary (OAC CHV) building. Terms that could not be mapped to existing UMLS concepts via machine and manual methods prompted creation of new concepts, which were then ascribed semantic types, related to existing UMLS concepts, and coded according to specified criteria. RESULTS: This approach identified 64 unmapped concepts, 17 of which were labeled as uniquely "lay" and not feasible for inclusion in professional health terminologies. The remaining terms constituted potential candidates for inclusion in professional vocabularies, or could be constructed by post-coordinating existing UMLS terms. The relationship between new and existing concepts differed depending on the corpora from which they were extracted. CONCLUSION: Non-mapping concepts constitute a small proportion of consumer health terms, but a proportion that is likely to affect the process of consumer health vocabulary building. We have identified a novel approach for identifying such concepts. Alla Keselman, Catherine Arnott Smith, Guy Divita, Hyeon-Eui Kim, Allen C. Browne, Gondy Leroy, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 7 |
| 2008 | White Paper: Developing Informatics Tools and Strategies for Consumer-centered Health CommunicationabstractAs the emphasis on individuals' active partnership in health care grows, so does the public's need for effective, comprehensible consumer health resources. Consumer health informatics has the potential to provide frameworks and strategies for designing effective health communication tools that empower users and improve their health decisions. This article presents an overview of the consumer health informatics field, discusses promising approaches to supporting health communication, and identifies challenges plus direction for future research and development. The authors' recommendations emphasize the need for drawing upon communication and social science theories of information behavior, reaching out to consumers via a range of traditional and novel formats, gaining better understanding of the public's health information needs, and developing informatics solutions for tailoring resources to users' needs and competencies. This article was written as a scholarly outreach and leadership project by members of the American Medical Informatics Association's Consumer Health Informatics Working Group. Alla Keselman, Catherine Arnott Smith, Gondy Leroy, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 5 |
| 2008 | Research Paper: Estimating Consumer Familiarity with Health Terminology: A Context-based ApproachabstractOBJECTIVES: Effective health communication is often hindered by a "vocabulary gap" between language familiar to consumers and jargon used in medical practice and research. To present health information to consumers in a comprehensible fashion, we need to develop a mechanism to quantify health terms as being more likely or less likely to be understood by typical members of the lay public. Prior research has used approaches including syllable count, easy word list, and frequency count, all of which have significant limitations. DESIGN: In this article, we present a new method that predicts consumer familiarity using contextual information. The method was applied to a large query log data set and validated using results from two previously conducted consumer surveys. MEASUREMENTS: We measured the correlation between the survey result and the context-based prediction, syllable count, frequency count, and log normalized frequency count. RESULTS: The correlation coefficient between the context-based prediction and the survey result was 0.773 (p < 0.001), which was higher than the correlation coefficients between the survey result and the syllable count, frequency count, and log normalized frequency count (p < or = 0.012). CONCLUSIONS: The context-based approach provides a good alternative to the existing term familiarity assessment methods. Qing T. Zeng, Sergey Goryachev, Tony Tse, Alla Keselman, Aziz A. Boxwala |
J. Am. Medical Informatics Assoc. | 1 |
| 2007 | Towards Consumer-Friendly PHRs: Patients' Experience with Reviewing Their Health Records
Alla Keselman, Laura A. Slaughter, Catherine Arnott Smith, Hyeon-Eui Kim, Guy Divita, Allen C. Browne, Christopher Tsai, Qing T. Zeng |
AMIA | 8 |
| 2007 | Beyond Surface Characteristics: A New Health Text-Specific Readability Measurement
Hyeon-Eui Kim, Sergey Goryachev, Graciela Rosemblat, Allen C. Browne, Alla Keselman, Qing T. Zeng |
AMIA | 6 |
| 2007 | Making Texts in Electronic Health Records Comprehensible to Consumers: A Prototype Translator
Qing T. Zeng, Sergey Goryachev, Hyeon-Eui Kim, Alla Keselman, Douglas Rosendale |
AMIA | 1 |
| 2007 | Feature-guided clustering of multi-dimensional flow cytometry datasets
Qing T. Zeng, Juan Pablo Pratt, Jane Pak, Dino Ravnic, Harold Huss, Steven J. Mentzer |
J. Biomed. Informatics | 1 |
| 2006 | A Suite of Natural Language Processing Tools Developed for the I2B2 Project
Sergey Goryachev, Margarita Sordo, Qing T. Zeng |
AMIA | 3 |
| 2006 | Relating Consumer Knowledge of Health Terms and Health Concepts
Alla Keselman, Tony Tse, Jonathan Crowell, Allen C. Browne, Long H. Ngo, Qing T. Zeng |
AMIA | 6 |
| 2006 | Analysis of Information Needs of Users of MEDLINEplus, 2002 - 2003
Alicia Scott-Wright, Jonathan Crowell, Qing T. Zeng, David W. Bates, Robert A. Greenes |
AMIA | 3 |
| 2006 | Exploring Lexical Forms: First-Generation Consumer Health Vocabularies
Qing T. Zeng, Tony Tse, Guy Divita, Alla Keselman, Jonathan Crowell, Allen C. Browne |
AMIA | 1 |
| 2006 | Research Paper: Assisting Consumer Health Information Retrieval with Query RecommendationsabstractOBJECTIVE: Health information retrieval (HIR) on the Internet has become an important practice for millions of people, many of whom have problems forming effective queries. We have developed and evaluated a tool to assist people in health-related query formation. DESIGN: We developed the Health Information Query Assistant (HIQuA) system. The system suggests alternative/additional query terms related to the user's initial query that can be used as building blocks to construct a better, more specific query. The recommended terms are selected according to their semantic distance from the original query, which is calculated on the basis of concept co-occurrences in medical literature and log data as well as semantic relations in medical vocabularies. MEASUREMENTS: An evaluation of the HIQuA system was conducted and a total of 213 subjects participated in the study. The subjects were randomized into 2 groups. One group was given query recommendations and the other was not. Each subject performed HIR for both a predefined and a self-defined task. RESULTS: The study showed that providing HIQuA recommendations resulted in statistically significantly higher rates of successful queries (odds ratio = 1.66, 95% confidence interval = 1.16-2.38), although no statistically significant impact on user satisfaction or the users' ability to accomplish the predefined retrieval task was found. CONCLUSION: Providing semantic-distance-based query recommendations can help consumers with query formation during HIR. Qing T. Zeng, Jonathan Crowell, Robert M. Plovnick, Long H. Ngo, Emily Dibble |
J. Am. Medical Informatics Assoc. | 1 |
| 2006 | Viewpoint Paper: Exploring and Developing Consumer Health VocabulariesabstractLaypersons ("consumers") often have difficulty finding, understanding, and acting on health information due to gaps in their domain knowledge. Ideally, consumer health vocabularies (CHVs) would reflect the different ways consumers express and think about health topics, helping to bridge this vocabulary gap. However, despite the recent research on mismatches between consumer and professional language (e.g., lexical, semantic, and explanatory), there have been few systematic efforts to develop and evaluate CHVs. This paper presents the point of view that CHV development is practical and necessary for extending research on informatics-based tools to facilitate consumer health information seeking, retrieval, and understanding. In support of the view, we briefly describe a distributed, bottom-up approach for (1) exploring the relationship between common consumer health expressions and professional concepts and (2) developing an open-access, preliminary (draft) "first-generation" CHV. While recognizing the limitations of the approach (e.g., not addressing psychosocial and cultural factors), we suggest that such exploratory research and development will yield insights into the nature of consumer health expressions and assist developers in creating tools and applications to support consumer health information seeking. Qing T. Zeng, Tony Tse |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | A Web Application to Support Consumer Health Vocabulary Development
Jonathan Crowell, Qing T. Zeng, Tony Tse |
AMIA | 2 |
| 2005 | Identifying Consumer-Friendly Display (CFD) Names for Health Concepts
Qing T. Zeng, Tony Tse, Jonathan Crowell, Guy Divita, Laura Roth, Allen C. Browne |
AMIA | 1 |
| 2004 | Research Paper: A Frequency-based Technique to Improve the Spelling Suggestion Rank in Medical QueriesabstractOBJECTIVE: There is an abundance of health-related information online, and millions of consumers search for such information. Spell checking is of crucial importance in returning pertinent results, so the authors propose a technique for increasing the effectiveness of spell-checking tools used for health-related information retrieval. DESIGN: A sample of incorrectly spelled medical terms was submitted to two different spell-checking tools, and the resulting suggestions, derived under two different dictionary configurations, were re-sorted according to how frequently each term appeared in log data from a medical search engine. MEASUREMENTS: Univariable analysis was carried out to assess the effect of each factor (spell-checking tool, dictionary type, re-sort, or no re-sort) on the probability of success. The factors that were statistically significant in the univariable analysis were then used in multivariable analysis to evaluate the independent effect of each of the factors. RESULTS: The re-sorted suggestions proved to be significantly more accurate than the original list returned by the spell-checking tool. The odds of finding the correct suggestion in the number one rank were increased by 63% after re-sorting using the authors' method. This effect was independent of both the dictionary and the spell-checking tools that were used. CONCLUSION: Using knowledge about the frequency of a given word's occurrence in the medical domain can significantly improve spelling correction for medical queries. Jonathan Crowell, Qing T. Zeng, Long H. Ngo, Eve-Marie Lacroix |
J. Am. Medical Informatics Assoc. | 2 |
| 2004 | Review Paper: The InterMed Approach to Sharable Computer-interpretable Guidelines: A ReviewabstractInterMed is a collaboration among research groups from Stanford, Harvard, and Columbia Universities. The primary goal of InterMed has been to develop a sharable language that could serve as a standard for modeling computer-interpretable guidelines (CIGs). This language, called GuideLine Interchange Format (GLIF), has been developed in a collaborative manner and in an open process that has welcomed input from the larger community. The goals and experiences of the InterMed project and lessons that the authors have learned may contribute to the work of other researchers who are developing medical knowledge-based tools. The lessons described include (1) a work process for multi-institutional research and development that considers different viewpoints, (2) an evolutionary lifecycle process for developing medical knowledge representation formats, (3) the role of cognitive methodology to evaluate and assist in the evolutionary development process, (4) development of an architecture and (5) design principles for sharable medical knowledge representation formats, and (6) a process for standardization of a CIG modeling language. Mor Peleg, Aziz A. Boxwala, Samson W. Tu, Qing T. Zeng, Omolola Ogunyemi, Dongwen Wang, Vimla L. Patel, Robert A. Greenes, Edward H. Shortliffe |
J. Am. Medical Informatics Assoc. | 4 |
| 2004 | GLIF3: a representation format for sharable computer-interpretable clinical practice guidelines
Aziz A. Boxwala, Mor Peleg, Samson W. Tu, Omolola Ogunyemi, Qing T. Zeng, Dongwen Wang, Vimla L. Patel, Robert A. Greenes, Edward H. Shortliffe |
J. Biomed. Informatics | 5 |
| 2004 | Design and implementation of the GLIF3 guideline execution engine
Dongwen Wang, Mor Peleg, Samson W. Tu, Aziz A. Boxwala, Omolola Ogunyemi, Qing T. Zeng, Robert A. Greenes, Vimla L. Patel, Edward H. Shortliffe |
J. Biomed. Informatics | 6 |
| 2003 | Coverage of patient safety terms in the UMLS Metathesaurus
Aziz A. Boxwala, Qing T. Zeng, Anthony Chamberas, Luke Sato, Meghan Dierks |
AMIA | 2 |
| 2003 | A Technique to Improve the Spelling Suggestion Rank in Medical Queries
Jonathan Crowell, Qing T. Zeng, Sandra Kogan |
AMIA | 2 |
| 2003 | Visual Representation of Cell Subpopulation from Flow Cytometry Data
Qing T. Zeng, Frank C. Kuo, James Rawn, Steven J. Mentzer |
AMIA | 2 |
| 2002 | Applying Axiomatic Design Methodology to Create Guidelines That Are Locally Adaptable
Aziz A. Boxwala, Qing T. Zeng, Derrick Tate, Robert A. Greenes, David G. Fairchild |
AMIA | 2 |
| 2002 | Using a neural network with flow cytometry histograms to recognize cell surface protein binding patterns
Qing T. Zeng, James Rawn, Matthew P. Wand, Alan J. Young, Edgar L. Milford, Steven J. Mentzer, Robert A. Greenes |
AMIA | 2 |
| 2002 | Health Information Retrieval Tool (HIRT)
Mra Thinzar Nyun, Omolola Ogunyemi, Qing T. Zeng |
AMIA | 3 |
| 2002 | Matching of flow-cytometry histograms using information theory in feature space
Qing T. Zeng, Matthew P. Wand, Alan J. Young, James Rawn, Edgar L. Milford, Steven J. Mentzer, Robert A. Greenes |
AMIA | 1 |
| 2002 | Poster Abstract: Identification of Special Patterns of Numerical Typographic Errors Increases the Likelihood of Finding a Misplaced Patient FileabstractWhen a typographic error of a patient identification number occurs on a patient document such as an envelope for radiology films or the cover of a patient record, it will result in misplacement of the document. Once misplaced, such documents are often extremely difficult to recover. After analyzing 290 numerical typos, we found that errors do not occur randomly. Instead, many of the typos share certain specific patterns. Six major types of non-random numeral typographic error patterns have been identified and their frequency characterized. Knowing these patterns and their odds increases the likelihood of finding a misplaced file. In addition, awareness of these patterns during transcribing or writing a patient ID may decrease the chance of typographic errors. Ying-Chou Sun, Dah-Dian Tang, Qing T. Zeng, Robert A. Greenes |
J. Am. Medical Informatics Assoc. | 3 |
| 2002 | Research Paper: Providing Concept-oriented Views for Clinical Data Using a Knowledge-based System: An EvaluationabstractOBJECTIVE: Clinical information systems typically present patient data in chronologic order, organized by the source of the information (e.g., laboratory, radiology). This study evaluates the functionality and utility of a knowledge-based system that generates concept-oriented views (organized around clinical concepts such as disease or organ system) of clinical data. DESIGN: The authors have developed a system that uses a knowledge base of interrelationships between medical concepts to infer relationships between data in electronic medical records. They use these inferences to produce summaries, or views, of the data that are relevant to a specific concept of interest. They evaluated the ability of the system to select relevant information, reduce information overload, and support physician information retrieval. MEASUREMENTS: The sensitivity and specificity of the system for identifying relevant patient information were calculated. Effect on information overload was assessed by comparing the amount of information in each view with the amount of information in the entire record. Information retrieval accuracy and cost (time) were used to measure the effect of using concept-oriented views on the efficiency and effectiveness of retrievals. RESULTS: The sensitivity and specificity of the system for identifying relevant clinical information were generally in the range of 70 to 80 percent. Concept-oriented views are effective in reducing the amount of information retrieved (over 80 percent reduction) and, compared with source-oriented views, are able to improve physician retrieval accuracy (p=0.04). CONCLUSION: Computer-generated, concept-oriented views can be used to reduce clinician information overload and improve the accuracy of clinical data retrieval. Qing T. Zeng, James J. Cimino, Kelly H. Zou |
J. Am. Medical Informatics Assoc. | 1 |
| 2001 | Finding appropriate clinical trials: evaluating encoded eligibility criteria with incomplete data
Nachman Ash, Omolola Ogunyemi, Qing T. Zeng, Lucila Ohno-Machado |
AMIA | 3 |
| 2001 | Facilitate the Delivery of Tailored Materials to Asthmatic Pregnant Women
Sandra Kogan, Lynda M. Cristiano, Barrett T. Kitch, Qing T. Zeng |
AMIA | 5 |
| 2001 | Problems and challenges in patient information retrieval: a descriptive study
Sandra Kogan, Qing T. Zeng, Nachman Ash, Robert A. Greenes |
AMIA | 2 |
| 2001 | Using features of Arden Syntax with object-oriented medical data models for guideline modeling
Mor Peleg, Omolola Ogunyemi, Samson W. Tu, Aziz A. Boxwala, Qing T. Zeng, Robert A. Greenes, Edward H. Shortliffe |
AMIA | 5 |
| 2001 | Identification of Special Patterns of Numerical Typographic Errors to Increases the Likelihood of Finding a Misplaced Patient File
Ying-Chou Sun, Dah-Dian Tang, Qing T. Zeng, Robert A. Greenes |
AMIA | 3 |
| 2001 | Molecular identification using flow cytometry histograms and information theory
Qing T. Zeng, Alan J. Young, Aziz A. Boxwala, James Rawn, W. Long, Matthew P. Wand, Mikhail Salganik, Edgar L. Milford, Steven J. Mentzer, Robert A. Greenes |
AMIA | 1 |
| 2001 | Toward a Representation Format for Sharable Clinical Guidelines
Aziz A. Boxwala, Samson W. Tu, Mor Peleg, Qing T. Zeng, Omolola Ogunyemi, Robert A. Greenes, Edward H. Shortliffe, Vimla L. Patel |
J. Biomed. Informatics | 4 |
| 2001 | A Knowledge-Based, Concept-Oriented View Generation System for Clinical Data
Qing T. Zeng, James J. Cimino |
J. Biomed. Informatics | 1 |
| 2000 | GLIF3: the evolution of a guideline representation format
Mor Peleg, Aziz A. Boxwala, Omolola Ogunyemi, Qing T. Zeng, Samson W. Tu, Ronilda C. Lacson, Elmer V. Bernstam, Nachman Ash, Kris Mork, Lucila Ohno-Machado, Edward H. Shortliffe, Robert A. Greenes |
AMIA | 4 |
| 2000 | A Three-layer Domain Ontology for Guideline Representation and Sharing
Qing T. Zeng, Samson W. Tu, Aziz A. Boxwala, Mor Peleg, Robert A. Greenes, Edward H. Shortliffe |
AMIA | 1 |
| 1999 | Evaluation of a system to identify relevant patient information and its impact on clinical information retrieval
Qing T. Zeng, James J. Cimino |
AMIA | 1 |
| 1998 | Automated knowledge extraction from the UMLS
Qing T. Zeng, James J. Cimino |
AMIA | 1 |
| 1997 | Supporting infobuttons with terminological knowledge
James J. Cimino, Gai Elhanan, Qing T. Zeng |
AMIA | 3 |
| 1997 | Linking a clinical system to heterogeneous information resources
Qing T. Zeng, James J. Cimino |
AMIA | 1 |