EDBT 2026 Demo / reviewers in the wild / expert
Gondy Leroy
dblp:43/3357
· DBLP profile ↗
49ranked-venue papers
21as first author
10since 2021 · last 2026
0000-0003-4751-6680ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 35 · 15 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep learning for autism detection using clinical notes: A comparison of transfer learning for a transparent and black-box approachabstractAutism spectrum disorder (ASD) is a complex neurodevelopmental condition whose rising prevalence places increasing demands on a lengthy diagnostic process. Machine learning (ML) has shown promise in automating ASD diagnosis, but most existing models operate as black boxes and are typically trained on a single dataset, limiting their generalizability. In this study, we introduce a transparent and interpretable ML approach that leverages BioBERT, a state-of-the-art language model, to analyze unstructured clinical text. The model is trained to label descriptions of behaviors and map them to diagnostic criteria, which are then used to assign a final label (ASD or not). We evaluate transfer learning, the ability to transfer knowledge to new data, using two distinct real-world datasets. We trained on datasets sequentially and mixed together and compared the performance of the best models and their ability to transfer to new data. We also created a black-box approach and repeated this transfer process for comparison. Our transparent model demonstrated robust performance, with the mixed-data training strategy yielding the best results (97 % sensitivity, 98 % specificity). Sequential training across datasets led to a slight drop in performance, highlighting the importance of training data order. The black-box model performed worse (90 % sensitivity, 96 % specificity) when trained sequentially or with mixed data. Overall, our transparent approach outperformed the black-box approach. Mixing datasets during training resulted in slightly better performance and should be the preferred approach when practically possible. This work paves the way for more trustworthy, generalizable, and clinically actionable AI tools in neurodevelopmental diagnostics. Gondy Leroy, Prakash Bisht, Sai Madhuri Kandula, Nell Maltman, Sydney A Rice |
Artif. Intell. Medicine | 1 |
| 2026 | Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation
Yue Guo 0007, Jae Ho Sohn, Gondy Leroy, Trevor Cohen |
J. Biomed. Informatics | 3 |
| 2025 | Enhancing Text Datasets With Scaling and Targeting Data Augmentation to Improve BERT-Based Machine LearnersabstractSynthetic data is used to increase a dataset's size for machine learning when acquiring new data is difficult. However, this is difficult for text data due to its symbolic nature. Large language models have decreased this difficulty. Using autism spectrum disorder as a use case, we analyze how synthetic data chosen based on descriptive metrics impacts the performance of a downstream classifier. We leverage a finetuned multilabel, bidirectional encoder model to label textual descriptions of children's behaviors (N = 10892) with seven diagnostic criteria. We measure precision, recall, and F1 per label to compare the impact of augmentation schemes. We evaluate performance without augmentation (baseline), then compare data source (original versus synthetic), amount of data added (50% or 100% of baseline count), and method of augmentation (adding to one class via Data Targeting or to the entire dataset via Data Scaling). The data points were selected based on scores from our white-box metrics: type-token ratio, cosine similarity, and perplexity. We also conducted a qualitative evaluation of the data using expert feedback. This resulted in a consistent increase in recall (approximately 8 %) but a similarly consistent decrease in precision (approximately 10 %). Neither the white-box metrics nor the following standard-deviation-based stability analysis provided a clear relationship to our results in our model. Cost analysis showed that data targeting could lower the BioBERT model's cost. Overall, this study shows that different schemes should be favored depending on the intent of use, e.g., screening or diagnosing in medicine. Chancellor R. Woolsey, Gondy Leroy, Nell Maltman |
Expert Syst. Appl. | 2 |
| 2025 | Coherence and comprehensibility: Large language models predict lay understanding of health-related content
Trevor Cohen, Weizhe Xu, Yue Guo 0007, Serguei V. S. Pakhomov, Gondy Leroy |
J. Biomed. Informatics | 5 |
| 2024 | APPLS: Evaluating Evaluation Metrics for Plain Language SummarizationabstractWhile there has been significant development of models for Plain Language Summarization (PLS), evaluation remains a challenge. PLS lacks a dedicated assessment metric, and the suitability of text generation evaluation metrics is unclear due to the unique transformations involved (e.g., adding background explanations, removing jargon). To address these questions, our study introduces a granular meta-evaluation testbed, APPLS, designed to evaluate metrics for PLS. We identify four PLS criteria from previous work-informativeness, simplification, coherence, and faithfulness-and define a set of perturbations corresponding to these criteria that sensitive metrics should be able to detect. We apply these perturbations to the texts of two PLS datasets to create our testbed. Using APPLS, we assess performance of 14 metrics, including automated scores, lexical features, and LLM prompt-based evaluations. Our analysis reveals that while some current metrics show sensitivity to specific criteria, no single method captures all four criteria simultaneously. We therefore recommend a suite of automated metrics be used to capture PLS quality along all relevant criteria. This work contributes the first meta-evaluation testbed for PLS and a comprehensive evaluation of existing metrics. Yue Guo 0007, Tal August, Gondy Leroy, Trevor Cohen, Lucy Lu Wang |
EMNLP | 3 |
| 2024 | Introduction to the special issue on smart and connected health
Zhijun Yan, Gondy Leroy, Qiuju Yin, Nicholas R. Hardiker, Dongsong Zhang |
Inf. Manag. | 2 |
| 2024 | Transparent deep learning to identify autism spectrum disorders (ASD) in EHR using clinical notesabstractOBJECTIVE: Machine learning (ML) is increasingly employed to diagnose medical conditions, with algorithms trained to assign a single label using a black-box approach. We created an ML approach using deep learning that generates outcomes that are transparent and in line with clinical, diagnostic rules. We demonstrate our approach for autism spectrum disorders (ASD), a neurodevelopmental condition with increasing prevalence. METHODS: We use unstructured data from the Centers for Disease Control and Prevention (CDC) surveillance records labeled by a CDC-trained clinician with ASD A1-3 and B1-4 criterion labels per sentence and with ASD cases labels per record using Diagnostic and Statistical Manual of Mental Disorders (DSM5) rules. One rule-based and three deep ML algorithms and six ensembles were compared and evaluated using a test set with 6773 sentences (N = 35 cases) set aside in advance. Criterion and case labeling were evaluated for each ML algorithm and ensemble. Case labeling outcomes were compared also with seven traditional tests. RESULTS: Performance for criterion labeling was highest for the hybrid BiLSTM ML model. The best case labeling was achieved by an ensemble of two BiLSTM ML models using a majority vote. It achieved 100% precision (or PPV), 83% recall (or sensitivity), 100% specificity, 91% accuracy, and 0.91 F-measure. A comparison with existing diagnostic tests shows that our best ensemble was more accurate overall. CONCLUSIONS: Transparent ML is achievable even with small datasets. By focusing on intermediate steps, deep ML can provide transparent decisions. By leveraging data redundancies, ML errors at the intermediate level have a low impact on final outcomes. Gondy Leroy, Jennifer G. Andrews, Madison Kealohi-Preece, Ajay Jaswani, Hyunju Song, Maureen K. Galindo, Sydney A Rice |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Retrieval augmentation of large language models for lay language generation
Yue Guo 0007, Wei Qiu 0006, Gondy Leroy, Trevor Cohen |
J. Biomed. Informatics | 3 |
| 2021 | Incidence and Impact of Missing Functional Elements on Information Comprehension using Audio and Text
Gondy Leroy, David Kauchak, Nick Kloehn |
AMIA | 1 |
| 2021 | Comparison of women and men in biomedical informatics scientific dissemination: retrospective observational case study of the AMIA Annual Symposium: 2017-2020abstractOBJECTIVE: Although the representation of women in science has improved, women remain underrepresented in scientific publications. This study compares women and men in scholarly dissemination through the AMIA Annual Symposium. MATERIALS AND METHODS: Through a retrospective observational study, we analyzed 2017-2020 AMIA submissions for differences in panels, papers, podium abstracts, posters, workshops, and awards for men compared with women. We assigned a label of woman or man to authors and reviewers using Genderize.io, and then compared submission and acceptance rates, performed regression analyses to evaluate the impact of the assumed gender, and performed sentiment analysis of reviewer comments. RESULTS: Of the 4687 submissions for which Genderize.io could predict man or woman based on first name, 40% were led by women and 60% were led by men. The acceptance rate was smilar. Although submission and acceptance rates for women increased over the 4 years, women-led podium abstracts, panels, and workshops were underrepresented. Men reviewers increased the odds of rejection. Men provided longer reviews and lower reviewer scores, but women provided reviews that had more positive words. DISCUSSION: Overall, our findings reflect significant gains for women in the 4 years of conference data analyzed. However, there remain opportunities to improve representation of women in workshop submissions, panel and podium abstract speakers, and balanced peer reviews. Future analyses could be strengthened by collecting gender directly from authors, including diverse genders such as non-binary. CONCLUSION: We found little evidence of major bias against women in submission, acceptance, and awards associated with the AMIA Annual Symposium from 2017 to 2020. Our study is unique because of the analysis of both authors and reviewers. The encouraging findings raise awareness of progress and remaining opportunities in biomedical informatics scientific dissemination. Andrea L. Hartzler, Gondy Leroy, Brenda Daurelle, Magali Ochoa, Jeffrey Williamson, Dasha Cohen, Carole H. Stipelman |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | AutoMeTS: The Autocomplete for Medical Text SimplificationabstractThe goal of text simplification (TS) is to transform difficult text into a version that is easier to understand and more broadly accessible to a wide variety of readers.In some domains, such as healthcare, fully automated approaches cannot be used since information must be accurately preserved.Instead, semi-automated approaches can be used that assist a human writer in simplifying text faster and at a higher quality.In this paper, we examine the application of autocomplete to text simplification in the medical domain.We introduce a new parallel medical data set consisting of aligned English Wikipedia with Simple English Wikipedia sentences and examine the application of pretrained neural language models (PNLMs) on this dataset.We compare four PNLMs (BERT, RoBERTa, XLNet, and GPT-2), and show how the additional context of the sentence to be simplified can be incorporated to achieve better results (6.17% absolute improvement over the best individual model).We also introduce an ensemble model that combines the four PNLMs and outperforms the best individual model by 2.1%, resulting in an overall word prediction accuracy of 64.52%. Hoang Van, David Kauchak, Gondy Leroy |
COLING | 3 |
| 2019 | 2018 Salary Survey of AMIA Members: Factors Associated with Higher Salaries
Yan Cheng 0004, April F. Mohanty, Omolola Ogunyemi, Catherine Arnott Smith, Gondy Leroy, Qing T. Zeng |
AMIA | 5 |
| 2019 | Predicting Transition Words Between Sentences for English and Spanish Medical Text
David Kauchak, Gondy Leroy, Menglu Pei, Sonia Colina |
AMIA | 2 |
| 2019 | Evaluating Artifacts Using Experiments as Part of the Design Science FrameworkabstractDesign science is of increasing importance in information systems (IS) with top journals and funding agencies recognizing it as an important component of IS research. One critical element of good design science is a solid evaluation of the artifacts, i.e., of individual algorithms and completed information systems. This tutorial focuses on the design and execution of experiments to evaluate such artifacts in an efficient manner. The artifact's life cycle is taken into account to ensure the evaluation is appropriate, effective and efficient through optimal use of resources. Evaluations can range from automated evaluations of algorithms using gold standards to user studies of an entire information system with representative users in situ. The tutorial will focus on the controlled experiment for artifacts, appropriate statistical analysis, and tips on common errors that can be avoided. Gondy Leroy |
RCIS | 1 |
| 2019 | Using Lexical Chains to Identify Text Difficulty: A Corpus Statistics and Classification StudyabstractOur goal is data-driven discovery of features for text simplification. In this paper, we investigate three types of lexical chains: exact, synonymous, and semantic. A lexical chain links semantically related words in a document. We examine their potential with a document-level corpus statistics study (914 texts) to estimate their overall capacity to differentiate between easy and difficult text and a classification task (11 000 sentences) to determine usefulness of features at sentence-level for simplification. For the corpus statistics study we tested five document-level features for each chain type: total number of chains, average chain length, average chain span, number of crossing chains, and the number of chains longer than half the document length. We found significant differences between easy and difficult text for average chain length and the average number of cross chains. For the sentence classification study, we compared the lexical chain features to standard bag-of-words features on a range of classifiers: logistic regression, naïve Bayes, decision trees, linear and RBF kernel SVM, and random forest. The lexical chain features performed significantly better than the bag-of-words baseline across all classifiers with the best classifier achieving an accuracy of ∼90% (compared to 78% for bag-of-words). Overall, we find several lexical chain features provide specific information useful for identifying difficult sentences of text, beyond what is available from standard lexical features. Partha Mukherjee, Gondy Leroy, David Kauchak |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Optimizing Corpus Creation for Training Word Embedding in Low Resource Domains: A Case Study in Autism Spectrum Disorder (ASD)
Gondy Leroy, Sydney Pettygrove, Maureen K. Galindo, Margaret Kurzius-Spencer |
AMIA | 2 |
| 2017 | When synonyms are not enough: Optimal parenthetical insertion for text simplification
Gondy Leroy, David Kauchak |
AMIA | 2 |
| 2017 | Spanish Text Simplification Using Term Familiarity: Applying Principles from English Text Simplification
Gondy Leroy, Brianda Armenta Navarrete, Sonia Colina, David Kauchak |
AMIA | 1 |
| 2017 | The Role of Surface, Semantic and Grammatical Features on Simplification of Spanish Medical Texts: A User Study
Partha Mukherjee, Gondy Leroy, David Kauchak, Brianda Armenta Navarrete, Damian Y. Romero Diaz, Sonia Colina |
AMIA | 2 |
| 2017 | Creating a Corpus Resource for Text Simplification R & D
Debra Revere, Partha Mukherjee, David Kauchak, Gondy Leroy |
AMIA | 4 |
| 2017 | Automated Lexicon and Feature Construction Using Word Embedding and Clustering for Classification of ASD Diagnoses Using EHR
Gondy Leroy, Sydney Pettygrove, Margaret Kurzius-Spencer |
NLDB | 1 |
| 2017 | Measuring text difficulty using parse-tree frequencyabstractText simplification often relies on dated, unproven readability formulas. As an alternative and motivated by the success of term familiarity, we test a complementary measure: grammar familiarity. Grammar familiarity is measured as the frequency of the 3rd level sentence parse tree and is useful for evaluating individual sentences. We created a database of 140K unique 3rd level parse structures by parsing and binning all 5.4M sentences in English Wikipedia. We then calculated the grammar frequencies across the corpus and created 11 frequency bins. We evaluate the measure with a user study and corpus analysis. For the user study, we selected 20 sentences randomly from each bin, controlling for sentence length and term frequency, and recruited 30 readers per sentence (N = 6,600) on Amazon Mechanical Turk. We measured actual difficulty (comprehension) using a Cloze test, perceived difficulty using a 5‐point Likert scale, and time taken. Sentences with more frequent grammatical structures, even with very different surface presentations, were easier to understand, perceived as easier, and took less time to read. Outcomes from readability formulas correlated with perceived but not with actual difficulty. Our corpus analysis shows how the metric can be used to understand grammar regularity in a broad range of corpora. David Kauchak, Gondy Leroy, Alan Hogue |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | NegAIT: A new parser for medical text simplification using morphological, sentential and double negation
Partha Mukherjee, Gondy Leroy, David Kauchak, Srinidhi Rajanarayanan, Damian Y. Romero Diaz, Nicole P. Yuan, T. Gail Pritchard, Sonia Colina |
J. Biomed. Informatics | 2 |
| 2016 | Grammar frequency and simplification: when intuition fails
David Kauchak, Gondy Leroy, Melissa Just |
AMIA | 2 |
| 2016 | Reviewing Asthma-related Grey Literature and Personal Opinions on Twitter using LDA and CTM Clustering
Gondy Leroy, Joe Koolippurackal, Shikha Swami, Philip Harber |
AMIA | 1 |
| 2014 | Using Natural Language Processing for Autism Trigger Extraction
Gondy Leroy, Margaret Kurzius-Spencer, Sydney Pettygrove |
AMIA | 1 |
| 2014 | Does Query Expansion Limit Our Learning? A Comparison of Social-Based Expansion to Content-Based Expansion for Medical Queries on the Internet
Christopher Pentoney, Jeffrey Harwell, Gondy Leroy |
AMIA | 3 |
| 2013 | Development and evaluation of a biomedical search engine using a predicate-based vector space model
Myungjae Kwak, Gondy Leroy, Jesse D. Martinez, Jeffrey Harwell |
J. Biomed. Informatics | 2 |
| 2012 | Comparison of Lay and Professional Search Diagrams with Implications for a Biomedical Search Engine Design
Jeffrey Harwell, Gondy Leroy, Myungjae Kwak, Jesse D. Martinez |
AMIA | 2 |
| 2012 | A Systematic Grammatical Analysis of Easy and Difficult Medical Text
David Kauchak, William Coster, Gondy Leroy |
AMIA | 3 |
| 2012 | Improving Perceived and Actual Text Difficulty for Health Information Consumers using Semi-Automated Methods
Gondy Leroy, James E. Endicott, Obay Mouradi, David Kauchak, Melissa Just |
AMIA | 1 |
| 2012 | Evidence of Using a Mobile Device for Communication with Children with Autism: Lessons Learned and Critical Obstacles
Gondy Leroy, Juliette Gutierrez, HyeKyeung Seung |
AMIA | 1 |
| 2011 | Eliciting user requirements using Appreciative inquiry
Carol K. Gonzales, Gondy Leroy |
Empir. Softw. Eng. | 2 |
| 2011 | A crime reports analysis system to identify related crimesabstractAbstract The popularity of online and anonymous options to report crimes, such as tips websites and text messaging, has led to an increasing amount of textual information available to law enforcement personnel. However, locating, filtering, extracting, and combining information to solve crimes is a time‐consuming task. In response, we are developing entity and document similarity algorithms to automatically identify overlapping and complementary information. These are essential components for systems that combine and contrast crime information. The entity similarity algorithm integrates a domain‐specific hierarchical lexicon with Jaccard coefficients. The document similarity algorithm combines the entity similarity scores using a Dice coefficient. We describe the evaluation of both components. To evaluate the entity similarity algorithm, we compared the new algorithm and four generic algorithms with a gold standard. The strongest correlation with the gold standard, r = 0.710, was found with our entity similarity algorithm. To evaluate the document similarity algorithm, we first developed a test bed containing witness reports for 17 crimes shown in video clips. We evaluated five versions of the algorithm that differ in how much importance is assigned to different entity types. Cosine similarity is then used as a baseline comparison to evaluate the performance of the document similarity algorithms for accuracy in recognizing reports describing the same crime and distinguishing them from reports on different crimes. The best version achieved 92% accuracy. Chih Hao Ku 0001, Gondy Leroy |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2010 | Perils of providing visual health information overviews for consumers with low health literacy or high stressabstractThis pilot study explores the impact of a health topics overview (HTO) on reading comprehension. The HTO is generated automatically based on the presence of Unified Medical Language System terms. In a controlled setting, we presented health texts and posed 15 questions for each. We compared performance with and without the HTO. The answers were available in the text, but not always in the HTO. Our study (n=48) showed that consumers with low health literacy or high stress performed poorly when the HTO was available without linking directly to the answer. They performed better with direct links in the HTO or when the HTO was not available at all. Consumers with high health literacy or low stress performed better regardless of the availability of the HTO. Our data suggests that vulnerable consumers relied solely on the HTO when it was available and were misled when it did not provide the answer. Gondy Leroy, Trudi Miller |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Persuading Consumers to Form Precise Search Engine Queries
Gondy Leroy |
AMIA | 1 |
| 2008 | Smartphones to facilitate communication and improve social skills of children with severe autism spectrum disorder: special education teachers as proxiesabstractWe present an overview of the approach we used and the challenges we encountered while designing software for smartphones to facilitate communication and improve social skills of children with severe autism spectrum disorder (ASD). We employed participatory design, using special education teachers of children with ASD as proxies for our target population. Gianluca De Leo, Gondy Leroy |
IDC | 2 |
| 2008 | Evaluating Online Health Information: Beyond Readability Formulas
Gondy Leroy, Stephen Helmreich, James R. Cowie, Trudi Miller |
AMIA | 1 |
| 2008 | Research Paper: Consumer Health Concepts That Do Not Map to the UMLS: Where Do They Fit?abstractOBJECTIVE: This study has two objectives: first, to identify and characterize consumer health terms not found in the Unified Medical Language System (UMLS) Metathesaurus (2007 AB); second, to describe the procedure for creating new concepts in the process of building a consumer health vocabulary. How do the unmapped consumer health concepts relate to the existing UMLS concepts? What is the place of these new concepts in professional medical discourse? DESIGN: The consumer health terms were extracted from two large corpora derived in the process of Open Access Collaboratory Consumer Health Vocabulary (OAC CHV) building. Terms that could not be mapped to existing UMLS concepts via machine and manual methods prompted creation of new concepts, which were then ascribed semantic types, related to existing UMLS concepts, and coded according to specified criteria. RESULTS: This approach identified 64 unmapped concepts, 17 of which were labeled as uniquely "lay" and not feasible for inclusion in professional health terminologies. The remaining terms constituted potential candidates for inclusion in professional vocabularies, or could be constructed by post-coordinating existing UMLS terms. The relationship between new and existing concepts differed depending on the corpora from which they were extracted. CONCLUSION: Non-mapping concepts constitute a small proportion of consumer health terms, but a proportion that is likely to affect the process of consumer health vocabulary building. We have identified a novel approach for identifying such concepts. Alla Keselman, Catherine Arnott Smith, Guy Divita, Hyeon-Eui Kim, Allen C. Browne, Gondy Leroy, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 6 |
| 2008 | White Paper: Developing Informatics Tools and Strategies for Consumer-centered Health CommunicationabstractAs the emphasis on individuals' active partnership in health care grows, so does the public's need for effective, comprehensible consumer health resources. Consumer health informatics has the potential to provide frameworks and strategies for designing effective health communication tools that empower users and improve their health decisions. This article presents an overview of the consumer health informatics field, discusses promising approaches to supporting health communication, and identifies challenges plus direction for future research and development. The authors' recommendations emphasize the need for drawing upon communication and social science theories of information behavior, reaching out to consumers via a range of traditional and novel formats, gaining better understanding of the public's health information needs, and developing informatics solutions for tailoring resources to users' needs and competencies. This article was written as a scholarly outreach and leadership project by members of the American Medical Informatics Association's Consumer Health Informatics Working Group. Alla Keselman, Catherine Arnott Smith, Gondy Leroy, Qing T. Zeng |
J. Am. Medical Informatics Assoc. | 4 |
| 2008 | A balanced approach to health information evaluation: A vocabulary-based naïve Bayes classifier and readability formulasabstractAbstract Since millions seek health information online, it is vital for this information to be comprehensible. Most studies use readability formulas, which ignore vocabulary, and conclude that online health information is too difficult. We developed a vocabularly‐based, naïve Bayes classifier to distinguish between three difficulty levels in text. It proved 98% accurate in a 250‐document evaluation. We compared our classifier with readability formulas for 90 new documents with different origins and asked representative human evaluators, an expert and a consumer, to judge each document. Average readability grade levels for educational and commercial pages was 10th grade or higher, too difficult according to current literature. In contrast, the classifier showed that 70–90% of these pages were written at an intermediate, appropriate level indicating that vocabulary usage is frequently appropriate in text considered too difficult by readability formula evaluations. The expert considered the pages more difficult for a consumer than the consumer did. Gondy Leroy, Trudi Miller, Graciela Rosemblat, Allen C. Browne |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2007 | Introduction to the special issue on decision support in medicine
Gondy Leroy, Hsinchun Chen |
Decis. Support Syst. | 1 |
| 2006 | Health Information Text Characteristics
Gondy Leroy, Evren Eryilmaz, Benjamin T. Laroya |
AMIA | 1 |
| 2006 | Dynamic Generation of a Table of Contents with Consumer-Friendly Labels
Trudi Miller, Gondy Leroy, Elizabeth Wood |
AMIA | 2 |
| 2005 | Communication Software using Pictures for use with Pocket PCs
Gondy Leroy, John Huang, Serena Chuang, Marjorie H. Charlop-Christy |
AMIA | 1 |
| 2005 | Genescene: An ontology-enhanced integration of linguistic and co-occurrence based relations in biomedical textsabstractAbstract The increasing amount of publicly available literature and experimental data in biomedicine makes it hard for biomedical researchers to stay up‐to‐date. Genescene is a toolkit that will help alleviate this problem by providing an overview of published literature content. We combined a linguistic parser with Concept Space, a co‐occurrence based semantic net. Both techniques extract complementary biomedical relations between noun phrases from MEDLINE abstracts. The parser extracts precise and semantically rich relations from individual abstracts. Concept Space extracts relations that hold true for the collection of abstracts. The Gene Ontology, the Human Genome Nomenclature, and the Unified Medical Language System, are also integrated in Genescene. Currently, they are used to facilitate the integration of the two relation types, and to select the more interesting and high‐quality relations for presentation. A user study focusing on p53 literature is discussed. All MEDLINE abstracts discussing p53 were processed in Genescene. Two researchers evaluated the terms and relations from several abstracts of interest to them. The results show that the terms were precise (precision 93%) and relevant, as were the parser relations (precision 95%). The Concept Space relations were more precise when selected with ontological knowledge (precision 78%) than without (60%). Gondy Leroy, Hsinchun Chen |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | A shallow parser based on closed-class words to capture relations in biomedical text
Gondy Leroy, Hsinchun Chen, Jesse D. Martinez |
J. Biomed. Informatics | 1 |
| 2003 | The use of dynamic context to improve casual internet searchingabstractResearch has shown that most users' online information searches are suboptimal. Query optimization based on a relevance feedback or genetic algorithm using dynamic query contexts can help casual users search the Internet. These algorithms can draw on implicit user feedback based on the surrounding links and text in a search engine result set to expand user queries with a variable number of keywords in two manners. Positive expansion adds terms to a user's keywords with a Boolean "and," negative expansion adds terms to the user's keywords with a Boolean "not." Each algorithm was examined for three user groups, high, middle, and low achievers, who were classified according to their overall performance. The interactions of users with different levels of expertise with different expansion types or algorithms were evaluated. The genetic algorithm with negative expansion tripled recall and doubled precision for low achievers, but high achievers displayed an opposed trend and seemed to be hindered in this condition. The effect of other conditions was less substantial. Gondy Leroy, Ann M. Lally, Hsinchun Chen |
ACM Trans. Inf. Syst. | 1 |
| 2001 | Meeting medical terminology needs-the ontology-enhanced Medical Concept MapperabstractThis paper describes the development and testing of the Medical Concept Mapper, a tool designed to facilitate access to online medical information sources by providing users with appropriate medical search terms for their personal queries. Our system is valuable for patients whose knowledge of medical vocabularies is inadequate to find the desired information, and for medical experts who search for information outside their field of expertise. The Medical Concept Mapper maps synonyms and semantically related concepts to a user's query. The system is unique because it integrates our natural language processing tool, i.e., the Arizona (AZ) Noun Phraser, with human-created ontologies, the Unified Medical Language System (UMLS) and WordNet, and our computer generated Concept Space, into one system. Our unique contribution results from combining the UMLS Semantic Net with Concept Space in our deep semantic parsing (DSP) algorithm. This algorithm establishes a medical query context based on the UMLS Semantic Net, which allows Concept Space terms to be filtered so as to isolate related terms relevant to the query. We performed two user studies in which Medical Concept Mapper terms were compared against human experts' terms. We conclude that the AZ Noun Phraser is well suited to extract medical phrases from user queries, that WordNet is not well suited to provide strictly medical synonyms, that the UMLS Metathesaurus is well suited to provide medical synonyms, and that Concept Space is well suited to provide related medical terms, especially when these terms are limited by our DSP algorithm. Gondy Leroy, Hsinchun Chen |
IEEE Trans. Inf. Technol. Biomed. | 1 |