William R. Hersh

dblp:59/3255 · DBLP profile ↗
← Back
118ranked-venue papers
43as first author
9since 2021 · last 2026
0000-0002-4114-5148ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 93 · 28 first-author · 9 since 2021Databases, data management, data science and information retrieval · 22 · 15 first-authorArtificial intelligence and machine learning · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Clinical document metadata extraction: A scoping review
abstract
OBJECTIVES: Clinical document metadata, such as document type, structure, author role, medical specialty, and encounter setting, is essential for accurate interpretation of information captured in clinical documents. However, vast documentation heterogeneity and drift over time challenge harmonization of document metadata. Automated extraction methods have emerged to coalesce metadata from disparate practices into target schema. This scoping review aims to catalog research on clinical document metadata extraction, identify methodological trends and applications, and highlight gaps warranting further investigation. METHODS: We followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) guidelines to identify articles from Ovid MEDLINE, Ovid EMBASE, Scopus, Web of Science and external sources that perform clinical document metadata extraction, either primarily as a methodology study, secondarily as a feature for a downstream application, or for analysis. We initially identified and screened 342 articles published between 2011 and 2025, then comprehensively reviewed 77 we deemed relevant to our study. RESULTS: Among the 77 articles included in our full text review, 49 were methodological, 22 used document metadata as features in a downstream application, and 6 analyzed document metadata composition. We observe myriad purposes for methodological study and application types. Available labelled public data remains sparse except for structural section datasets. Methods for extracting document metadata have progressed from largely rule-based and traditional machine learning with ample feature engineering to transformer-based architectures with minimal feature engineering. DISCUSSION AND CONCLUSION: Clinical document metadata extraction research has accelerated over recent years. The emergence of large language models has enabled broader exploration of generalizability across tasks and datasets, allowing the possibility of advanced clinical text processing systems. We anticipate that research will continue to expand into richer document metadata representations and integrate further into clinical applications and workflows.
Kurt Miller, Qiuhao Lu, William R. Hersh, Kirk Roberts, Steven Bedrick, Andrew Wen
J. Biomed. Informatics3
2025 Dynamic few-shot prompting for clinical note section classification using lightweight, open-source large language models
abstract
OBJECTIVE: Unlocking clinical information embedded in clinical notes has been hindered to a significant degree by domain-specific and context-sensitive language. Identification of note sections and structural document elements has been shown to improve information extraction and dependent downstream clinical natural language processing (NLP) tasks and applications. This study investigates the viability of a dynamic example selection prompting method to section classification using lightweight, open-source large language models (LLMs) as a practical solution for real-world healthcare clinical NLP systems. MATERIALS AND METHODS: We develop a dynamic few-shot prompting approach to classifying sections where section samples are first embedded using a transformer-based model and deposited in a vector store. During inference, the embedded samples with the most similar contextual embeddings to a given input section text are retrieved from the vector store and inserted into the LLM prompt. We evaluate this technique on two datasets comprising two section schemas, including varying levels of context. We compare the performance to baseline zero-shot and randomly selected few-shot scenarios. RESULTS: The dynamic few-shot prompting experiments yielded the highest F1 scores in each of the classification tasks and datasets for all seven of the LLMs included in the evaluation, averaging a macro F1 increase of 39.3% and 21.1% in our primary section classification task over the zero-shot and static few-shot baselines, respectively. DISCUSSION AND CONCLUSION: Our results showcase substantial performance improvements imparted by dynamically selecting examples for few-shot LLM prompting, and further improvement by including section context, demonstrating compelling potential for clinical applications.
Kurt Miller, Steven Bedrick, Qiuhao Lu, Andrew Wen, William R. Hersh, Kirk Roberts
J. Am. Medical Informatics Assoc.5
2024 Automatically pre-screening patients for the rare disease aromatic l-amino acid decarboxylase deficiency using knowledge engineering, natural language processing, and machine learning on a large EHR population
abstract
OBJECTIVES: Electronic health record (EHR) data may facilitate the identification of rare diseases in patients, such as aromatic l-amino acid decarboxylase deficiency (AADCd), an autosomal recessive disease caused by pathogenic variants in the dopa decarboxylase gene. Deficiency of the AADC enzyme results in combined severe reductions in monoamine neurotransmitters: dopamine, serotonin, epinephrine, and norepinephrine. This leads to widespread neurological complications affecting motor, behavioral, and autonomic function. The goal of this study was to use EHR data to identify previously undiagnosed patients who may have AADCd without available training cases for the disease. MATERIALS AND METHODS: A multiple symptom and related disease annotated dataset was created and used to train individual concept classifiers on annotated sentence data. A multistep algorithm was then used to combine concept predictions into a single patient rank value. RESULTS: Using an 8000-patient dataset that the algorithms had not seen before ranking, the top and bottom 200 ranked patients were manually reviewed for clinical indications of performing an AADCd diagnostic screening test. The top-ranked patients were 22.5% positively assessed for diagnostic screening, with 0% for the bottom-ranked patients. This result is statistically significant at P < .0001. CONCLUSION: This work validates the approach that large-scale rare-disease screening can be accomplished by combining predictions for relevant individual symptoms and related conditions which are much more common and for which training data is easier to create.
Aaron M. Cohen, Jolie Kaner, Ryan Miller, Jeffrey W. Kopesky, William R. Hersh
J. Am. Medical Informatics Assoc.5
2024 Search still matters: information retrieval in the era of generative AI
abstract
OBJECTIVE: Information retrieval (IR, also known as search) systems are ubiquitous in modern times. How does the emergence of generative artificial intelligence (AI), based on large language models (LLMs), fit into the IR process? PROCESS: This perspective explores the use of generative AI in the context of the motivations, considerations, and outcomes of the IR process with a focus on the academic use of such systems. CONCLUSIONS: There are many information needs, from simple to complex, that motivate use of IR. Users of such systems, particularly academics, have concerns for authoritativeness, timeliness, and contextualization of search. While LLMs may provide functionality that aids the IR process, the continued need for search systems, and research into their improvement, remains essential.
William R. Hersh
J. Am. Medical Informatics Assoc.1
2022 Beyond Wrangling and Modeling: Data Science and Machine Learning Competencies and Curricula for The Rest of Us
William R. Hersh, Jessica Ancker, Robert Hoyt
AMIA1
2022 A debate on the extension of the Practice Pathway for ABMS clinical informatics board certification for physicians in the United States
Ellen Kim, Christoph U. Lehmann, William R. Hersh, Clifton D. Fuller, Bruce P. Levy
AMIA3
2021 Career Development Issues for Women in Biomedical Informatics within Professional Organizations
Donghua Tao, Duo Helen Wei, Rubina F. Rizvi, Deepti Pandita, Bushra Alghamdi, Polina V. Kukhareva, Margarita Sordo, William R. Hersh, Omolola Ogunyemi, Kelly Evans, Gretchen Purcell Jackson
AMIA8
2021 A comparative analysis of system features used in the TREC-COVID information retrieval challenge
Jimmy S. Chen, William R. Hersh
J. Biomed. Informatics2
2021 Searching for scientific evidence in a pandemic: An overview of TREC-COVID
abstract
We present an overview of the TREC-COVID Challenge, an information retrieval (IR) shared task to evaluate search on scientific literature related to COVID-19. The goals of TREC-COVID include the construction of a pandemic search test collection and the evaluation of IR methods for COVID-19. The challenge was conducted over five rounds from April to July 2020, with participation from 92 unique teams and 556 individual submissions. A total of 50 topics (sets of related queries) were used in the evaluation, starting at 30 topics for Round 1 and adding 5 new topics per round to target emerging topics at that state of the still-emerging pandemic. This paper provides a comprehensive overview of the structure and results of TREC-COVID. Specifically, the paper provides details on the background, task structure, topic structure, corpus, participation, pooling, assessment, judgments, results, top-performing systems, lessons learned, and benchmark datasets.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Biomed. Informatics9
2020 Translational Research of Machine Learning and Artificial Intelligence Advances in Clinical Settings - Experiences and Challenges
William R. Hersh, Gretchen Purcell Jackson, Marc S. Williams, Colin G. Walsh, David A. Dorr
AMIA1
2020 TREC-COVID: rationale and structure of an information retrieval shared task for COVID-19
abstract
TREC-COVID is an information retrieval (IR) shared task initiated to support clinicians and clinical research during the COVID-19 pandemic. IR for pandemics breaks many normal assumptions, which can be seen by examining 9 important basic IR research questions related to pandemic situations. TREC-COVID differs from traditional IR shared task evaluations with special considerations for the expected users, IR modality considerations, topic development, participant requirements, assessment process, relevance criteria, evaluation metrics, iteration process, projected timeline, and the implications of data use as a post-task test collection. This article describes how all these were addressed for the particular requirements of developing IR systems under a pandemic situation. Finally, initial participation numbers are also provided, which demonstrate the tremendous interest the IR community has in this effort.
Kirk Roberts, Tasmeer Alam, Steven Bedrick, Dina Demner-Fushman, Kyle Lo, Ian Soboroff, Ellen M. Voorhees, Lucy Lu Wang, William R. Hersh
J. Am. Medical Informatics Assoc.9
2019 Thriving in Your Biomedical Informatics Career While Balancing Work, Personal, and Family Life
Lesley Clack, Qing T. Zeng, Gretchen Purcell Jackson, William R. Hersh, April F. Mohanty
AMIA4
2019 Educational Communities in the Academic Forum: Sharing Knowledge and Promoting Standards
Stephen B. Johnson, Saif S. Khairat, Suzanne Austin Boren, Vishnu Mohan, Eva LaVerne Manos, William R. Hersh
AMIA6
2018 Benchmarking Information Retrieval for Precision Oncology: the TREC Precision Medicine Track
Kirk Roberts, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh, Steven Bedrick, Alexander J. Lazar, Shubham Pant
AMIA4
2018 Panel: Collaborative Science Within Academic Medical Centers: Opportunities and Challenges for Informatics
Justin Starren, William R. Hersh, Christopher A. Longhurst, Philip R. O. Payne
AMIA2
2017 Evaluation of Clinical Text Segmentation to Facilitate Cohort Retrieval
Tracy Edinger, Dina Demner-Fushman, Aaron M. Cohen, Steven Bedrick, William R. Hersh
AMIA5
2017 Clinical Informatics in Medical Education: Innovations from the AMA Accelerating Change in Medical Education Initiative
William R. Hersh, Susan E. Skochelak, Anderson Spickard, Blaine Y. Takesue, Paul N. Gorman
AMIA1
2017 Information Retrieval for Biomedical Datasets: The 2016 bioCADDIE Challenge
Kirk Roberts, Anupama E. Gururaj, Saeid Pournejati, Trevor Cohen, William R. Hersh, Dina Demner-Fushman, Lucila Ohno-Machado, Hua Xu 0001
AMIA6
2017 Intrainstitutional EHR collections for patient-level information retrieval
abstract
Research in clinical information retrieval has long been stymied by the lack of open resources. However, both clinical information retrieval research innovation and legitimate privacy concerns can be served by the creation of intrainstitutional, fully protected resources. In this article, we provide some principles and tools for information retrieval resource‐building in the unique problem setting of patient‐level information retrieval, following the tradition of the Cranfield paradigm. We further include an analysis of parallel information retrieval resources at Oregon Health & Science University and Mayo Clinic that were built on these principles.
Stephen T. Wu, Sijia Liu 0002, Yanshan Wang, Tamara Timmons, Harsha Uppili, Steven Bedrick, William R. Hersh
J. Assoc. Inf. Sci. Technol.7
2017 Ethics, Medicine, and Information Technology, Kenneth W. Goodman. Cambridge University Press (2016). 192 pp. ISBN:1316432661
William R. Hersh
J. Biomed. Informatics1
2016 Development of Test Topics for Cohort Identification
Tamara Timmons, Stephen T. Wu, William R. Hersh
AMIA3
2016 On Developing Resources for Patient-level Information Retrieval
Stephen T. Wu, Tamara Timmons, Amy Yates, Meikun Wang, Steven Bedrick, William R. Hersh
LREC6
2016 State-of-the-art in biomedical literature retrieval for clinical cases: a survey of the TREC 2014 CDS track
Kirk Roberts, Matthew S. Simpson, Dina Demner-Fushman, Ellen M. Voorhees, William R. Hersh
Inf. Retr. J.5
2016 Early experiences of accredited clinical informatics fellowships
abstract
Since the launch of the clinical informatics subspecialty for physicians in 2013, over 1100 physicians have used the practice and education pathways to become board-certified in clinical informatics. Starting in 2018, only physicians who have completed a 2-year clinical informatics fellowship program accredited by the Accreditation Council on Graduate Medical Education will be eligible to take the board exam. The purpose of this viewpoint piece is to describe the collective experience of the first four programs accredited by the Accreditation Council on Graduate Medical Education and to share lessons learned in developing new fellowship programs in this novel medical subspecialty.
Christopher A. Longhurst, Natalie M. Pageler, Jonathan P. Palma, John T. Finnell, Bruce P. Levy, Thomas R. Yackel, Vishnu Mohan, William R. Hersh
J. Am. Medical Informatics Assoc.8
2014 Informatics without Borders: International Outreach of US-based Training Programs
Cynthia S. Gadd, Lucila Ohno-Machado, William R. Hersh, Rebecca S. Jacobson
AMIA3
2014 Evolving Career Landscapes in Biomedical and Health Informatics
Rui Zhang 0028, William R. Hersh, Genevieve B. Melton, Laura K. Wiley, Julie Doberne, Nawanan Theera-Ampornpunt
AMIA2
2014 Design and evaluation of the ONC health information technology curriculum
abstract
OBJECTIVE: As part of the Heath Information Technology for Economic and Clinical Health (HITECH) Act, the Office of the National Coordinator for Health Information Technology (ONC) implemented its Workforce Development Program, which included initiatives to train health information technology (HIT) professionals in 12 workforce roles, half of them in community colleges. To achieve this, the ONC tasked five universities with established informatics programs with creating curricular materials that could be used by community colleges. The five universities created 20 components that were made available for downloading from the National Training and Dissemination Center (NTDC) website. This paper describes an evaluation of the curricular materials by its intended audience of educators. METHODS: We measured the quantity of downloads from the NTDC site and administered a survey about the curricular materials to its registered users to determine use patterns and user characteristics. The survey was evaluated using mixed methods. Registered users downloaded nearly half a million units or components from the NTDC website. We surveyed these 9835 registered users. RESULTS: 1269 individuals completed all or part of the survey, of whom 339 identified themselves as educators (26.7% of all respondents). This paper addresses the survey responses of educators. DISCUSSION: Successful aspects of the curriculum included its breadth, convenience, hands-on and course planning capabilities. Several areas were identified for potential improvement. CONCLUSIONS: The ONC HIT curriculum met its goals for community college programs and will likely continue to be a valuable resource for the larger informatics community in the future.
Vishnu Mohan, Patricia A. Abbott, Shelby Acteson, Eta S. Berner, Corkey Devlin, William Edward Hammond, Rita Kukafka, William R. Hersh
J. Am. Medical Informatics Assoc.8
2013 Training the Informatics Research Workforce, Part 1
Valerie Florance, John F. Hurdle, Rita Kukafka, William R. Hersh, Cynthia S. Gadd, Alexa T. McCray
AMIA4
2013 Training the Informatics Research Workforce, Part 2
Valerie Florance, Perry L. Miller, Lucila Ohno-Machado, George Hripcsak, William R. Hersh, George Demiris
AMIA5
2013 Workshop on health search and discovery: helping users and advancing medicine
abstract
This workshop brings together researchers and practitioners from industry and academia to discuss search and discovery in the medi-cal domain. The event focuses on ways to make medical and health information more accessible to laypeople (including enhancements to ranking algorithms and search interfaces), and how we can dis-cover new medical facts and phenomena from information sought online, as evidenced in query streams and other sources such as social media. This domain also offers many opportunities for appli-cations that monitor and improve quality of life of those affected by medical conditions, by providing tools to support their health-related information behavior.
Ryen W. White, Elad Yom-Tov, Eric Horvitz, Eugene Agichtein, William R. Hersh
SIGIR5
2012 Barriers to Retrieving Patient Information from Electronic Health Record Data: Failure Analysis from the TREC Medical Records Track
Tracy Edinger, Aaron M. Cohen, Steven Bedrick, Kyle H. Ambert, William R. Hersh
AMIA5
2010 A systematic literature review of automated clinical coding and classification systems
abstract
Clinical coding and classification processes transform natural language descriptions in clinical text into data that can subsequently be used for clinical care, research, and other purposes. This systematic literature review examined studies that evaluated all types of automated coding and classification systems to determine the performance of such systems. Studies indexed in Medline or other relevant databases prior to March 2009 were considered. The 113 studies included in this review show that automated tools exist for a variety of coding and classification purposes, focus on various healthcare specialties, and handle a wide variety of clinical document types. Automated coding and classification systems themselves are not generalizable, nor are the results of the studies evaluating them. Published research shows these systems hold promise, but these data must be considered in context, with performance relative to the complexity of the task and the desired outcome.
Mary H. Stanfill, Margaret Williams, Susan H. Fenton, Robert A. Jenders, William R. Hersh
J. Am. Medical Informatics Assoc.5
2010 Unintended consequences of health information technology: A need for biomedical informatics
Elmer V. Bernstam, William R. Hersh, Ida Sim, David Eichmann, Jonathan C. Silverstein, Jack W. Smith, Michael J. Becich
J. Biomed. Informatics2
2009 Evaluation of a gene information summarization system by users during the analysis process of microarray datasets
abstract
BACKGROUND: Summarization of gene information in the literature has the potential to help genomics researchers translate basic research into clinical benefits. Gene expression microarrays have been used to study biomarkers for disease and discover novel types of therapeutics and the task of finding information in journal articles on sets of genes is common for translational researchers working with microarray data. However, manually searching and scanning the literature references returned from PubMed is a time-consuming task for scientists. We built and evaluated an automatic summarizer of information on genes studied in microarray experiments. The Gene Information Clustering and Summarization System (GICSS) is a system that integrates two related steps of the microarray data analysis process: functional gene clustering and gene information gathering. The system evaluation was conducted during the process of genomic researchers analyzing their own experimental microarray datasets. RESULTS: The clusters generated by GICSS were validated by scientists during their microarray analysis process. In addition, presenting sentences in the abstract provided significantly more important information to the users than just showing the title in the default PubMed format. CONCLUSION: The evaluation results suggest that GICSS can be useful for researchers in genomic area. In addition, the hybrid evaluation method, partway between intrinsic and extrinsic system evaluation, may enable researchers to gauge the true usefulness of the tool for the scientists in their natural analysis workflow and also elicit suggestions for future enhancements. AVAILABILITY: GICSS can be accessed online at: http://ir.ohsu.edu/jianji/index.html.
Jianji Yang, Aaron M. Cohen, William R. Hersh
BMC Bioinform.3
2009 TREC genomics special issue overview
William R. Hersh, Ellen M. Voorhees
Inf. Retr.1
2009 Tasks, topics and relevance judging for the TREC Genomics Track: five years of experience evaluating biomedical text information retrieval systems
Phoebe M. Roberts, Aaron M. Cohen, William R. Hersh
Inf. Retr.3
2009 Research Paper: Usability Testing Finds Problems for Novice Users of Pediatric Portals
abstract
OBJECTIVE: Patient portals may improve pediatric chronic disease outcomes, but few have been rigorously evaluated for usability by parents. Using scenario-based testing with think-aloud protocols, we evaluated the usability of portals for parents of children with cystic fibrosis, diabetes or arthritis. DESIGN Sixteen parents used a prototype and test data to complete 14 tasks followed by a validated satisfaction questionnaire. Three iterations of the prototype were used. MEASUREMENTS: During the usability testing, we measured the time it took participants to complete or give up on each task. Sessions were videotaped and content-analyzed for common themes. Following testing, participants completed the Computer Usability Satisfaction Questionnaire which measured their opinions on the efficiency of the system, its ease of use, and the likeability of the system interface. A 7-point Likert scale was used, with seven indicating the highest possible satisfaction. RESULTS: Mean task completion times ranged from 73 (+/- 61) seconds to locate a document to 431 (+/- 286) seconds to graph laboratory results. Tasks such as graphing, location of data, requesting access, and data interpretation were challenging. Satisfaction was greatest for interface pleasantness (5.9 +/- 0.7) and likeability (5.8 +/- 0.6) and lowest for error messages (2.3 +/- 1.2) and clarity of information (4.2 +/- 1.4). Overall mean satisfaction scores improved between iteration one and three. CONCLUSIONS: Despite parental involvement and prior heuristic testing, scenario-based testing demonstrated difficulties in navigation, medical language complexity, error recovery, and provider-based organizational schema. While such usability testing can be expensive, the current study demonstrates that it can assist in making healthcare system interfaces for laypersons more user-friendly and potentially more functional for patients and their families.
Maria T. Britto, Holly Brügge Jimison, Jennifer Knopf Munafo, Jennifer Wissman, Michelle L. Rogers, William R. Hersh
J. Am. Medical Informatics Assoc.6
2009 Research Paper: Predictors of Student Success in Graduate Biomedical Informatics Training: Introductory Course and Program Success
abstract
OBJECTIVE: To predict student performance in an introductory graduate-level biomedical informatics course from application data. DESIGN: A predictive model built through retrospective review of student records using hierarchical binary logistic regression with half of the sample held back for cross-validation. The model was also validated against student data from a similar course at a second institution. MEASUREMENTS: Earning an A grade (Mastery) or a C grade (Failure) in an introductory informatics course. RESULTS: The authors analyzed 129 student records at the University of Texas School of Health Information Sciences at Houston (SHIS) and 106 at Oregon Health and Science University Department of Medical Informatics and Clinical Epidemiology (DMICE). In the SHIS cross-validation sample, the Graduate Record Exam verbal score (GRE-V) correctly predicted Mastery in 69.4%. Undergraduate grade point average (UGPA) and underrepresented minority status (URMS) predicted 81.6% of Failures. At DMICE, GRE-V, UGPA, and prior graduate degree significantly correlated with Mastery. Only GRE-V was a significant independent predictor of Mastery at both institutions. There were too few URMS students and Failures at DMICE to analyze. Course Mastery strongly predicted program performance defined as final cumulative GPA at SHIS (n=19, r=0.634, r2=0.40, p=0.0036) and DMICE (n=106, r=0.603, r2=0.36, p<0.001). CONCLUSIONS: The authors identified predictors of performance in an introductory informatics course including GRE-V, UGPA and URMS. Course performance was a very strong predictor of overall program performance. Findings may be useful for selecting students for admission and identifying students at risk for Failure as early as possible.
Irmgard Willcockson, Craig W. Johnson, William R. Hersh, Elmer V. Bernstam
J. Am. Medical Informatics Assoc.3
2008 Evaluating the AMIA-OHSU 10x10 Program to Train Healthcare Professionals in Medical Informatics
Sue S. Feldman, William R. Hersh
AMIA2
2008 What Workforce is Needed to Implement the Health Information Technology Agenda? Analysis from the HIMSS Analytics™ Database
William R. Hersh, Adam Wright
AMIA1
2008 Survival Prediction Models for Estimating the Benefit of Post-Operative Radiation Therapy for Gallbladder Cancer and Lung Cancer
Jayashree Kalpathy-Cramer, William R. Hersh, Jong Song Kim, Charles R. Thomas, Samuel J. Wang
AMIA2
2008 Effectiveness of global features for automatic medical image classification and retrieval - The experiences of OHSU at ImageCLEFmed
Jayashree Kalpathy-Cramer, William R. Hersh
Pattern Recognit. Lett.2
2007 A comparative analysis of retrieval features used in the TREC 2006 Genomics Track passage retrieval task
Hari Krishna Rekapalli, Aaron M. Cohen, William R. Hersh
AMIA3
2007 Automatic Summarization of Mouse Gene Information by Clustering and Sentence Extraction from MEDLINE Abstracts
Jianji Yang, Aaron M. Cohen, William R. Hersh
AMIA3
2007 Reviewer merits
Gary Marchionini, Tefko Saracevic, John M. Carroll 0001, Donald H. Kraft, William R. Hersh, Josiane Mothe, Justin Zobel, Peter Hernon, Candy Schwartz
Inf. Process. Manag.5
2007 Research paper: A Day in the Life of PubMed: Analysis of a Typical Day's Query Log
abstract
OBJECTIVE: To characterize PubMed usage over a typical day and compare it to previous studies of user behavior on Web search engines. DESIGN: We performed a lexical and semantic analysis of 2,689,166 queries issued on PubMed over 24 consecutive hours on a typical day. MEASUREMENTS: We measured the number of queries, number of distinct users, queries per user, terms per query, common terms, Boolean operator use, common phrases, result set size, MeSH categories, used semantic measurements to group queries into sessions, and studied the addition and removal of terms from consecutive queries to gauge search strategies. RESULTS: The size of the result sets from a sample of queries showed a bimodal distribution, with peaks at approximately 3 and 100 results, suggesting that a large group of queries was tightly focused and another was broad. Like Web search engine sessions, most PubMed sessions consisted of a single query. However, PubMed queries contained more terms. CONCLUSION: PubMed's usage profile should be considered when educating users, building user interfaces, and developing future biomedical information retrieval systems.
Jorge R. Herskovic, Len Y. Tanaka, William R. Hersh, Elmer V. Bernstam
J. Am. Medical Informatics Assoc.3
2006 Who are the Informaticians?
William R. Hersh
AMIA1
2006 Adopting e-Learning Standards in Health Care: Competency-based Learning in the Medical Informatics Domain
William R. Hersh, Ravi Teja Bhupatiraju, Peter S. Greene, Valerie Smothers, Cheryl Cohen
AMIA1
2006 Functional Gene Group Summarization by Clustering MEDLINE Abstract Sentences
Jianji Yang, Aaron M. Cohen, William R. Hersh
AMIA3
2006 Research Paper: Using Citation Data to Improve Retrieval from MEDLINE
abstract
OBJECTIVE: To determine whether algorithms developed for the World Wide Web can be applied to the biomedical literature in order to identify articles that are important as well as relevant. DESIGN AND MEASUREMENTS A direct comparison of eight algorithms: simple PubMed queries, clinical queries (sensitive and specific versions), vector cosine comparison, citation count, journal impact factor, PageRank, and machine learning based on polynomial support vector machines. The objective was to prioritize important articles, defined as being included in a pre-existing bibliography of important literature in surgical oncology. RESULTS Citation-based algorithms were more effective than noncitation-based algorithms at identifying important articles. The most effective strategies were simple citation count and PageRank, which on average identified over six important articles in the first 100 results compared to 0.85 for the best noncitation-based algorithm (p < 0.001). The authors saw similar differences between citation-based and noncitation-based algorithms at 10, 20, 50, 200, 500, and 1,000 results (p < 0.001). Citation lag affects performance of PageRank more than simple citation count. However, in spite of citation lag, citation-based algorithms remain more effective than noncitation-based algorithms. CONCLUSION Algorithms that have proved successful on the World Wide Web can be applied to biomedical information retrieval. Citation-based algorithms can help identify important articles within large sets of relevant results. Further studies are needed to determine whether citation-based algorithms can effectively meet actual user information needs.
Elmer V. Bernstam, Jorge R. Herskovic, Yindalon Aphinyanagphongs, Constantin F. Aliferis, Madurai G. Sriram, William R. Hersh
J. Am. Medical Informatics Assoc.6
2006 Research Paper: Reducing Workload in Systematic Review Preparation Using Automated Citation Classification
abstract
OBJECTIVE: To determine whether automated classification of document citations can be useful in reducing the time spent by experts reviewing journal articles for inclusion in updating systematic reviews of drug class efficacy for treatment of disease. DESIGN: A test collection was built using the annotated reference files from 15 systematic drug class reviews. A voting perceptron-based automated citation classification system was constructed to classify each article as containing high-quality, drug class-specific evidence or not. Cross-validation experiments were performed to evaluate performance. MEASUREMENTS: Precision, recall, and F-measure were evaluated at a range of sample weightings. Work saved over sampling at 95% recall was used as the measure of value to the review process. RESULTS: A reduction in the number of articles needing manual review was found for 11 of the 15 drug review topics studied. For three of the topics, the reduction was 50% or greater. CONCLUSION: Automated document citation classification could be a useful tool in maintaining systematic reviews of the efficacy of drug therapy. Further work is needed to refine the classification system and determine the best manner to integrate the system into the production of systematic reviews.
Aaron M. Cohen, William R. Hersh, K. Peterson, Po-Yin Yen
J. Am. Medical Informatics Assoc.2
2006 Viewpoint Paper: Who are the Informaticians? What We Know and Should Know
abstract
The beginning of the 21st century has seen a surge in interest and enthusiasm for health care information technology based on its ability to demonstrate improvements in the quality, safety, and cost-efficiency of health care. One question, however, for which we have fewer answers is "who will be the individuals that develop, implement, and evaluate these systems?" In particular, while most attention has been paid to the exemplar leaders in health information technology, less has been focused on the issue of the workforce necessary to sustain the systems to achieve their vision. The discipline of medical informatics must pay sufficient attention to the professional workforce that will deploy systems outside the informatics research setting so their benefits may more widely accrue.
William R. Hersh
J. Am. Medical Informatics Assoc.1
2006 Model Formulation: Advancing Biomedical Image Retrieval: Development and Analysis of a Test Collection
abstract
OBJECTIVE: Develop and analyze results from an image retrieval test collection. METHODS: After participating research groups obtained and assessed results from their systems in the image retrieval task of Cross-Language Evaluation Forum, we assessed the results for common themes and trends. In addition to overall performance, results were analyzed on the basis of topic categories (those most amenable to visual, textual, or mixed approaches) and run categories (those employing queries entered by automated or manual means as well as those using visual, textual, or mixed indexing and retrieval methods). We also assessed results on the different topics and compared the impact of duplicate relevance judgments. RESULTS: A total of 13 research groups participated. Analysis was limited to the best run submitted by each group in each run category. The best results were obtained by systems that combined visual and textual methods. There was substantial variation in performance across topics. Systems employing textual methods were more resilient to visually oriented topics than those using visual methods were to textually oriented topics. The primary performance measure of mean average precision (MAP) was not necessarily associated with other measures, including those possibly more pertinent to real users, such as precision at 10 or 30 images. CONCLUSIONS: We developed a test collection amenable to assessing visual and textual methods for image retrieval. Future work must focus on how varying topic and run types affect retrieval performance. Users' studies also are necessary to determine the best measures for evaluating the efficacy of image retrieval systems.
William R. Hersh, Henning Müller, Jeffery R. Jensen, Jianji Yang, Paul N. Gorman, Patrick Ruch
J. Am. Medical Informatics Assoc.1
2005 A Standards-Based Approach for Facilitating Discovery of Learning Objectsat the Point of Care
William R. Hersh, Ravi Teja Bhupatiraju, Peter S. Greene, Valerie Smothers, Cheryl Cohen
AMIA1
2005 Evaluation axes for medical image retrieval systems: the imageCLEF experience
abstract
Content--based image retrieval in the medical domain is an extremely hot topic in medical imaging as it promises to help better managing the large amount of medical images being produced. Applications are mainly expected in the field of medical teaching files and for research projects, where performance issues and speed are less critical than in the field of diagnostic aid. Final goal with most impact will be the use as a diagnostic aid in a real--world clinical setting.Other applications of image retrieval and image classification can be the automatic annotation of images with basic concepts or the control of DICOM header information.ImageCLEF is part of the Cross Language Evaluation Forum (CLEF). Since 2004, a medical image retrieval task has been added. Goal is to create databases of a realistic and useful size and also query topics that are based on real--world needs in the medical domain but still correspond to the limited capabilities of purely visual retrieval at the moment. Goal is to direct the research onto real applications and towards real clinical problems to give researchers who are not directly linked to medical facilities a possibility to work on the interesting problem of medical image retrieval based on real data sets and problems. The missing link between computer science research departments and clinical routine is one of the biggest problems that becomes evident when reading much of the current literature on medical image retrieval. Most databases are extremely small, the treated problems often far from clinical reality, and there is no integration of the prototypes into a hospital infrastructure. Only few retrieval articles specifically mention problems related to the DICOM format (Digital Imaging and Communications in Medicine) and the sheer amount of data that needs to be treated in an image archive ( > 30.000 images per day in the Geneva radiology).This article develops the various axes that can be taken into account for medical image retrieval system evaluation. First, the axes are developed based on current challenges and experiences from ImageCLEF. Then, the resources developed for ImageCLEF are listed and finally, the application of the axes is explained to show the bases of the ImageCLEFmed evaluation campaign. This article will only concentrate on the medical retrieval tasks, the non-medical tasks will only shortly be mentioned.
Henning Müller, Paul D. Clough, William R. Hersh, Thomas Deselaers, Thomas M. Deserno, Antoine Geissbühler
ACM Multimedia3
2005 A survey of current work in biomedical text mining
abstract
The volume of published biomedical research, and therefore the underlying biomedical knowledge base, is expanding at an increasing rate. Among the tools that can aid researchers in coping with this information overload are text mining and knowledge extraction. Significant progress has been made in applying text mining to named entity recognition, text classification, terminology extraction, relationship extraction and hypothesis generation. Several research groups are constructing integrated flexible text-mining systems intended for multiple uses. The major challenge of biomedical text mining over the next 5-10 years is to make these systems useful to biomedical researchers. This will require enhanced access to full text, better understanding of the feature space of biomedical literature, better methods for measuring the usefulness of systems to users, and continued cooperation with the biomedical research community to ensure that their needs are addressed.
Aaron M. Cohen, William R. Hersh
Briefings Bioinform.2
2005 Evaluation of biomedical text-mining systems: Lessons learned from information retrieval
abstract
Biomedical text-mining systems have great promise for improving the efficiency and productivity of biomedical researchers. However, such systems are still not in routine use. One impediment to their development is the lack of systematic and rigorous evaluation, comparable to the approaches developed for information retrieval systems. The developers of text-mining systems need to improve both test collections for system-oriented evaluation and undertake user-oriented evaluations to determine the most effective use of their systems for their intended audience.
William R. Hersh
Briefings Bioinform.1
2005 Using co-occurrence network structure to extract synonymous gene and protein names from MEDLINE abstracts
abstract
BACKGROUND: Text-mining can assist biomedical researchers in reducing information overload by extracting useful knowledge from large collections of text. We developed a novel text-mining method based on analyzing the network structure created by symbol co-occurrences as a way to extend the capabilities of knowledge extraction. The method was applied to the task of automatic gene and protein name synonym extraction. RESULTS: Performance was measured on a test set consisting of about 50,000 abstracts from one year of MEDLINE. Synonyms retrieved from curated genomics databases were used as a gold standard. The system obtained a maximum F-score of 22.21% (23.18% precision and 21.36% recall), with high efficiency in the use of seed pairs. CONCLUSION: The method performs comparably with other studied methods, does not rely on sophisticated named-entity recognition, and requires little initial seed knowledge.
Aaron M. Cohen, William R. Hersh, Christopher Dubay, Kent A. Spackman
BMC Bioinform.2
2004 Research Paper: Computerized Physician Order Entry in U.S. Hospitals: Results of a 2002 Survey
abstract
OBJECTIVE: To determine the availability of inpatient computerized physician order entry in U.S. hospitals and the degree to which physicians are using it. DESIGN: Combined mail and telephone survey of 964 randomly selected hospitals, contrasting 2002 data and results of a survey conducted in 1997. AVAILABILITY: computerized order entry has been installed and is available for use by physicians; inducement: the degree to which use of computers to enter orders is required of physicians; participation: the proportion of physicians at an institution who enter orders by computer; and saturation: the proportion of total orders at an institution entered by a physician using a computer. RESULTS: The response rate was 65%. Computerized order entry was not available to physicians at 524 (83.7%) of 626 hospitals responding, whereas 60 (9.6%) reported complete availability and 41 (6.5%) reported partial availability. Of 91 hospitals providing data about inducement/requirement to use the system, it was optional at 31 (34.1%), encouraged at 18 (19.8%), and required at 42 (46.2%). At 36 hospitals (45.6%), more than 90% of physicians on staff use the system, whereas six (7.6%) reported 51-90% participation and 37 (46.8%) reported participation by fewer than half of physicians. Saturation was bimodal, with 25 (35%) hospitals reporting that more than 90% of all orders are entered by physicians using a computer and 20 (28.2%) reporting that less than 10% of all orders are entered this way. CONCLUSION: Despite increasing consensus about the desirability of computerized physician order entry (CPOE) use, these data indicate that only 9.6% of U.S. hospitals presently have CPOE completely available. In those hospitals that have CPOE, its use is frequently required. In approximately half of those hospitals, more than 90% of physicians use CPOE; in one-third of them, more than 90% of orders are entered via CPOE.
Joan S. Ash, Paul N. Gorman, Veena Seshadri, William R. Hersh
J. Am. Medical Informatics Assoc.4
2003 Development Methodology for a "Next Generation" Medical Informatics Curriculum for Clinicians
Eric Rose, Roni F. Zeiger, Sarah Corley, Paul N. Gorman, Thomas R. Yackel, William R. Hersh
AMIA6
2003 Research Paper: A Pilot Study of Contextual UMLS Indexing to Improve the Precision of Concept-based Representation in XML-structured Clinical Radiology Reports
abstract
OBJECTIVE: Despite the advantages of structured data entry, much of the patient record is still stored as unstructured or semistructured narrative text. The issue of representing clinical document content remains problematic. The authors' prior work using an automated UMLS document indexing system has been encouraging but has been affected by the generally low indexing precision of such systems. In an effort to improve precision, the authors have developed a context-sensitive document indexing model to calculate the optimal subset of UMLS source vocabularies used to index each document section. This pilot study was performed to evaluate the utility of this indexing approach on a set of clinical radiology reports. DESIGN: A set of clinical radiology reports that had been indexed manually using UMLS concept descriptors was indexed automatically by the SAPHIRE indexing engine. Using the data generated by this process the authors developed a system that simulated indexing, at the document section level, of the same document set using many permutations of a subset of the UMLS constituent vocabularies. MEASUREMENTS: The precision and recall scores generated by simulated indexing for each permutation of two or three UMLS constituent vocabularies were determined. RESULTS: While there was considerable variation in precision and recall values across the different subtypes of radiology reports, the overall effect of this indexing strategy using the best combination of two or three UMLS constituent vocabularies was an improvement in precision without significant impact of recall. CONCLUSION: In this pilot study a contextual indexing strategy improved overall precision in a set of clinical radiology reports.
Henry J. Lowe, William R. Hersh
J. Am. Medical Informatics Assoc.3
2002 Delivering bioinformatics training: bridging the gaps between computer science and biomedicine
Christopher Dubay, James M. Brundege, William R. Hersh, Kent A. Spackman
AMIA3
2002 Professional's Information Link (PiL): a web-based asynchronous consultation service
William R. Hersh, Robin Miller, Daniel Olson, Lynetta Sacherek, Ping Cross
AMIA1
2002 SmartQuery: context-sensitive links to medical knowledge sources from the electronic patient record
Susan Price, William R. Hersh, Daniel Olson, Peter J. Embí
AMIA2
2002 Distributed medical informatics education using internet2
Patricia Tidmarsh, Joseph Cummings, William R. Hersh, Charles P. Friedman
AMIA3
2002 User interface effects in past batch versus user experiments
abstract
No abstract available.
Andrew Turpin, William R. Hersh
SIGIR2
2002 Research Paper: Factors Associated with Success in Searching MEDLINE and Applying Evidence to Answer Clinical Questions
abstract
OBJECTIVES: This study sought to assess the ability of medical and nurse practitioner students to use MEDLINE to obtain evidence for answering clinical questions and to identify factors associated with the successful answering of questions. METHODS: A convenience sample of medical and nurse practitioner students was recruited. After completing instruments measuring demographic variables, computer and searching attitudes and experience, and cognitive traits, the subjects were given a brief orientation to MEDLINE searching and the techniques of evidence-based medicine. The subjects were then given 5 questions (from a pool of 20) to answer in two sessions using the Ovid MEDLINE system and the Oregon Health & Science University library collection. Each question was answered using three possible responses that reflected the quality of the evidence. All actions capable of being logged by the Ovid system were captured. Statistical analysis was performed using a model based on generalized estimating equations. The relevance-based measures of recall and precision were measured by defining end queries and having relevance judgments made by physicians who were not associated with the study. RESULTS: Forty-five medical and 21 nurse practitioner students provided usable answers to 324 questions. The rate of correctness increased from 32.3 to 51.6 percent for medical students and from 31.7 to 34.7 percent for nurse practitioner students. Ability to answer questions correctly was most strongly associated with correctness of the answer before searching, user experience with MEDLINE features, the evidence-based medicine question type, and the spatial visualization score. The spatial visualization score showed multi-colinearity with student type (medical vs. nurse practitioner). Medical and nurse practitioner students obtained comparable recall and precision, neither of which was associated with correctness of the answer. CONCLUSIONS: Medical and nurse practitioner students in this study were at best moderately successful at answering clinical questions correctly with the assistance of literature searching. The results confirm the importance of evaluating both search ability and the ability to use the resulting information to accomplish a clinical task.
William R. Hersh, M. Katherine Crabtree, David H. Hickam, Lynetta Sacherek, Charles P. Friedman, Patricia Tidmarsh, Craig Mosbaek, Dale Kraemer
J. Am. Medical Informatics Assoc.1
2002 Telehealth: The Need for Evaluation Redux
abstract
Generally, we do not publish papers describing the background and methods for a research project until the results are available for inclusion. In the case of the papers by Shea and Starren in this issue, we decided to make an exception. Our decision reflects the following factors. First, the literature does not contain examples of adequate evaluation of telemedicine despite years of application of the technology and several calls to action. Second, this trial is a major effort, and the results will not be available for some time. Third, it is unlikely that multiple large-scale trials are underway in this area. Accordingly, we decided that access to the methods would inform the community of the type of research that is needed, regardless of the outcome of this specific trial. The urgent need for more evaluation argues against publication delay. The lack of competing parallel efforts limits the chance of introducing bias by early publication.—William W. Stead, MD Five years ago, the journal published a set of five articles describing telehealth applications. Accompanying those papers was an editorial lamenting the lack of adequate evaluation for these studies.1 This editorial appeared shortly after an Institute of Medicine report was published that reached the same conclusions about telemedicine in general.2 Unfortunately, this situation has not changed in the ensuing half-decade. Two years ago, concerned with political pressure to reimburse telemedicine services through Medicare despite an unclear picture about efficacy and cost-effectiveness, the Health Care Financing Authority, along with the Agency for Healthcare Research and Quality, awarded a contract to the Evidence-based Practice Center at Oregon Health & Science University, to prepare an evidence review on the efficacy of telemedicine interventions in terms of diagnosis, clinical outcomes, satisfaction, access to care, and cost. The original review assessed telehealth applications for the Medicare population,3 while a supplemental study analyzed pediatric and obstetric population applications.4 Our conclusions from these reviews, which were exhaustive analyses of the peer-reviewed literature in telemedicine, echoed the previous observations—that while telemedicine research has led to novel and creative uses of the technology, the quality of the evaluation studies is poor. It is important to note that the major problems we found were with the methodologies of the studies. Thus, we were careful not to conclude that telemedicine technologies were not efficacious but rather that the low quality of studies assessing them precluded any conclusions about their efficacy. Examples of the problems we found included: Diagnostic efficacy studies in which the telemedicine and in-person assessments were performed by the same individuals A paucity of clinical outcome studies in clinical areas in which telemedicine is widely used Satisfaction studies with extremely low response rates and use of nonstandardized instruments Access studies that failed to use appropriate measures of access to care Cost studies that focused solely on the tradeoff of system cost vs. patient travel or emergency transport cost, ignoring the effects of adverse outcomes or the cost of the whole episode of care In our reviews, we noted that medical informatics investigators have demonstrated for many years the capability to carry out well-designed studies assessing the application of information technology in the health care setting.5–7 To the potential concern that the technology is changing too rapidly to achieve adequate research control conditions, we also stated that techniques (e.g., “tracker trials”) have been developed to cope with clinical trials of changing technologies.8 We also took journal editors to task for publishing papers that had well-written descriptions of systems and issues in the use of technologies but were marred by poor evaluation designs. Thus, we concluded that the telemedicine community still had not met the a challenge of defining the efficacious use and cost-effectiveness of their technologies. This issue of the Journal features two papers describing a large-scale telemedicine project in New York.9,10 The ongoing evaluation study is naturally of great interest to us. We are pleased that a large-scale evaluation has been incorporated into this project from the outset and note that its methodology is vastly superior to those of most of the studies we reviewed for our evidence report. However, we do have two concerns about the methodology that we hope the authors will address in this or future studies. The first is that the control intervention appears to consist of no intervention at all. Since the experimental intervention consists of both a telemedicine intervention and intensive nursing case management, a positive outcome of the study will not enable us to discern whether benefit accrued from the telemedicine intervention, the extensive case management, or both. Somewhat ironically, this issue has also plagued the various studies of the Diabetes Control and Complications Trial.11 Although these experiments were purported to demonstrate the beneficial effects of tight blood sugar control in diabetes mellitus, they may in reality have demonstrated the value of nursing case management. A better control group in the Shea et al. study would be one in which comparable nursing case management was also delivered by non-telemedicine means. This would enable us to assess the added value of telemedicine per se. The second concern is a hope that Shea et al. will look beyond intermediate patient outcome measures of glycosylated hemoglobin and blood pressure to actual outcomes, such as development of complications, morbidity, and mortality. Although the measurement of intermediate outcomes provides more statistical power, the ultimate aim of health care interventions is to directly benefit patients, not improve their test results. A related problem with the intermediate outcome measures is their use in cost-effectiveness calculations. Although the current evaluation will allow this intervention to be compared with other interventions that affect the intermediate measures (i.e., other diabetes interventions), it will not allow comparisons that a policy maker might wish to make, such as comparison with treatments for osteoporosis. A better cost-effectiveness measure would be quality-adjusted life years. Nonetheless, we applaud this large-scale trial and eagerly await the results. We hope the findings will not only enable us to get a better picture of the value of this form of telemedicine application for diabetic care but also serve as a means of expanding our knowledge about the contribution of control conditions for future evaluations. Such high-quality studies should also serve as a standard and raise the bar for the methodology of future evaluations in telemedicine and telehealth applications. Drs. Hersh, Patterson, and Kraemer appropriately highlight two limitations of the design for the IDEATel project. The first, boiled down, is whether the medium can be evaluated separately from the message. In designing IDEATel, we could not develop a practical, realistic way to separate the electronic medium for delivering diabetes care from the care itself. To do so would have required an artificial simulation of what electronically delivered care would be and then to have found an effective and methodologically convincing way to deliver this care non-electronically. This has never been done, it is not clear what this really means, and we did not believe we could do it successfully. For example, would each home telecare visit in the intervention group need to be matched by an in-person house call or by an in-person visit to a diabetes center in northern Manhattan or in Syracuse (nearly 800 miles for the most distant of the upstate region participants)? These two diabetes centers are where the intervention case managers are located. Our design choices were conditioned by what we believed we could do on a very large scale, with very short start-up time, and with very complex demands in terms of mounting the intervention technology. With these demands on the intervention side, it was desirable to keep the control side as simple as possible. The second point is that an optimal design would focus on “actual” outcomes, such as cardiovascular events, amputations, and death, rather than intermediate outcomes such as blood pressure, glucose control, and lipid levels. This was not a viable option without a much larger sample size and longer follow-up time. Many clinicians would accept that improvement in these intermediate variables would clearly indicate benefit to patients. We believe current and emerging technologies in medical informatics and telecommunication will alter not only the way care is delivered but what care is delivered, and that these two changes will occur in tight linkage. Telemedicine makes it possible to provide diabetes center-based case management to people who are not otherwise getting it. The most important factor driving the adoption of these technologies is that patients want to have access to health information, self-care resources, and the health care system electronically, remotely, and 24/7, which implies asynchronism. We hope that the IDEATel project will demonstrate that people now on the far side of the digital divide are there because of resources, not because of unwillingness or inability to use these technologies, and that these changes will be beneficial in terms of health outcomes and cost. Correspondence and reprints: Steven Shea, MD, Division of General Medicine, 622 W. 168th Street, New York, NY 10032; e-mail:
William R. Hersh, Patricia K. Patterson, Dale Kraemer
J. Am. Medical Informatics Assoc.1
2002 Forum Paper: Does National Regulatory Mandate of Provider Order Entry Portend Greater Benefit Than Risk for Health Care Delivery?: The 2001 ACMI Debate
abstract
The 2001 debate of the American College of Medical Informatics focused on the proposition that national regulatory mandate of computer-based provider order entry (CPOE), to take effect by the end of 2005, portends greater benefit than risk for health care delivery. Both sides accepted that provider order entry offers potential benefit. Those supporting the proposition emphasized public safety, noting that payers have little economic incentive to pay for quality and that a mandate would force vendors to improve the usability and value of their systems. They argued that the mandate would align the economic incentives to finally allow CPOE to be widely adopted. Those opposing the proposition emphasized the risks resulting from a mandate, including the direct implementation costs, the logistic issues of implementation, and the cost of failed implementations. They also noted the potential for errors introduced by the systems themselves and the fact that the safety and utility of commercially available CPOE products have yet to be proved.
J. Marc Overhage, Blackford Middleton, Randolph A. Miller, Rita D. Zielstorff, William R. Hersh
J. Am. Medical Informatics Assoc.5
2001 Distributed Medical Informatics Education Using Internet2
Joseph Cummings, Patricia Tidmarsh, William R. Hersh, Charles P. Friedman
AMIA3
2001 A Preliminary Trial of Tagging Online Documents Using "Medical Core Metadata"
Yang Gong, William R. Hersh
AMIA2
2001 The Role of Computer Science and Computing Skills in a Medical Informatics Curriculum
Susan Price, Judith R. Logan, William R. Hersh
AMIA3
2001 Why Batch and User Evaluations Do Not Give the Same Results
abstract
Much system-oriented evaluation of information retrieval systems has used the Cranfield approach based upon queries run against test collections in a batch mode. Some researchers have questioned whether this approach can be applied to the real world, but little data exists for or against that assertion. We have studied this question in the context of the TREC Interactive Track. Previous results demonstrated that improved performance as measured by relevance-based metrics in batch studies did not correspond with the results of outcomes based on real user searching tasks. The experiments in this paper analyzed those results to determine why this occurred. Our assessment showed that while the queries entered by real users into systems yielding better results in batch studies gave comparable gains in ranking of relevant documents for those users, they did not translate into better performance on specific tasks. This was most likely due to users being able to adequately find and utilize relevant documents ranked further down the output list.
Andrew Turpin, William R. Hersh
SIGIR2
2001 Interactivity at the Text Retrieval Conference (TREC)
William R. Hersh, Paul Over
Inf. Process. Manag.1
2001 Challenging conventional assumptions of automated information retrieval with real users: Boolean searching and batch retrieval evaluations
William R. Hersh, Andrew Turpin, Susan Price, Dale Kraemer, Daniel Olson, Benjamin Chan, Lynetta Sacherek
Inf. Process. Manag.1
2001 Managing Gigabytes - Compressing and Indexing Documents and Images (Second Edition)
William R. Hersh
Inf. Retr.1
2001 Application of Information Technology: Implementation and Evaluation of a Medical Informatics Distance Education Program
abstract
OBJECTIVE: Given the need for continuing education in medical informatics for mid-career professionals, the authors aimed to implement and evaluate distance learning courses in this area. DESIGN: The authors performed a needs assessment, content and technology planning, implementation, and student evaluation. MEASUREMENTS: The needs assessment and student evaluations were assessed using a combination of Likert scale and free-form questions. RESULTS: The needs assessment indicated much interest in a medical informatics distance learning program, with electronic medical records and outcome research the subject areas of most interest. The courses were implemented by means of streaming audio plus slides for lectures and threaded discussion boards for student interaction. Students were assessed by multiple-choice tests, a term paper, and a take-home final examination. In their course evaluations, student expressed strong satisfaction with the teaching modalities, course content, and system performance. Although not assessed experimentally, the performance of distance learning students was superior to that of on-campus students. CONCLUSION: Medical informatics education can be successfully implemented by means of distance learning technologies, with favorable student satisfaction and demonstrated learning. A graduate certificate program is now being implemented.
William R. Hersh, Katherine Junium, Mark Mailhot, Patricia Tidmarsh
J. Am. Medical Informatics Assoc.1
2001 Selective Automated Indexing of Findings and Diagnoses in Radiology Reports
William R. Hersh, Mark Mailhot, Catherine Arnott Smith, Henry J. Lowe
J. Biomed. Informatics1
2000 Distance learning techniques for medical informatics
William R. Hersh
AMIA1
2000 Assessing thesaurus-based query expansion using the UMLS Metathesaurus
William R. Hersh, Susan Price, Larry Donohoe
AMIA1
2000 Representing Clinical Information in an Internal Medicine Teaching Image Database
Jason A. Lyman, William R. Hersh, Kent A. Spackman
AMIA2
2000 Do batch and user evaluation give the same results?
abstract
Do improvements in system performance demonstrated by batch evaluations confer the same benefit for real users? We carried out experiments designed to investigate this question. After identifying a weighting scheme that gave maximum improvement over the baseline in a non-interactive evaluation, we used it with real users searching on an instance recall task. Our results showed the weighting scheme giving beneficial results in batch studies did not do so with real users. Further analysis did identify other factors predictive of instance recall, including number of documents saved by the user, document recall, and number of documents seen by the user.
William R. Hersh, Andrew Turpin, Susan Price, Benjamin Chan, Dale Kraemer, Lynetta Sacherek, Daniel Olson
SIGIR1
2000 Electronic Publishing of Scholarly Communication in the Biomedical Sciences
abstract
The main purpose of biomedical publishing is to convey the evolving scientific understanding of living organisms and the use of that understanding as the basis of health care among researchers, practitioners, and (increasingly) consumers. Published papers also serve as an archive of scientific knowledge, documenting both successful and unsuccessful results. Biomedical journals serve other purposes as well, such as providing benchmarks for academic promotion and revenue for professional societies and publishers. While MEDLINE has become ubiquitous and an increasing number of journals are available electronically, the fundamental model of publishing is unchanged. However, a number of challenges to that model are emerging. Boyd and Herkovic1 describe four challenges to scholarly publication: Cost. Publishers are charging increasing amounts of money for universities and academicians to access work that the latter created. Access. This high cost in turn is leading universities to decrease the size of their library collections, which in turn reduces access. Peer review. As more new journals, conference proceedings, and other forms of publication are introduced, the standards for peer review are diluted. Archiving. The growing use of electronic formats raises challenges to the long-term record of scientific publication. It is not altogether clear in what ways electronic publishing will improve or worsen this situation. On the one hand, the use of Internet-based management of the submission and dissemination process may reduce costs and increase the efficiency of peer review. On the other hand, if copyright laws and electronic protections hinder access, then costs could continue to go up, leading researchers to bypass peer review or reducing the incentive for high-quality scientific publication. We agree with Roberts2 that the World Wide Web provides at the same time both the possibilities for un-precedented dissemination of the fruits of scientific research and a brave new world where the price of obtaining such fruits remains high. As academicians who collaborate increasingly over the Internet, publish, read the works of others, and use such works in research, care, and teaching, we are attracted to the potential benefits of digital technologies. We idealize a world of free-flowing information adequately protected by reasonable intellectual property laws. But we fear the downside, whereby information becomes an expensive commodity surrounded by barriers that lock out individuals with lesser resources, such as students and those from developing countries. In summary, we share Roberts' “cautious optimism” about this new medium. This special issue of JAMIA is devoted to the topic of electronic publishing of scholarly information in the biomedical sciences. We sought both original research as well as perspective papers that addressed how electronic publishing affects the use of health and medical information by researchers, clinicians, consumers, and publishers. We have assembled a collection of seven papers that sample some of the important issues confronting the medical informatics community with regard to electronic publishing. Three papers present viewpoints on the economic and scientific ramifications of this new approach to disseminating scientific information. Two papers provide case studies, high-lighting specific experiences with this new medium. Another two papers present formulations for models for specific aspects of electronic publishing that may become prominent in the future. Coiera's viewpoint paper frames electronic publishing on the Internet in an economic light.3 He demonstrates the uncertainty we alluded to from Roberts' paper, noting that we cannot yet know how effective this new medium will be. Markovitz4 and Jacobson5 provide point and counterpoint views on the value of PubMed Central, the new initiative at the National Institutes of Health to create a freely available archive of scientific research papers built on top of the National Library of Medicine's PubMed system. Their papers highlight what are really two distinct issues in publishing: How do we handle preliminary and non-peer-reviewed research reports in this age of rapid information dissemination? Can we achieve a model of archive that still allows publishers reasonable revenues from a business standpoint? Anderson6 and D'Alessandro et al.7 provide two case studies of electronic publishing and share their results and perspectives. Anderson draws on the experience of the journal Pediatrics, which has been among the innovators in electronic publishing in medicine. He provides insight into the economic challenges facing academic publishers and concludes that while the Internet is a “disruptive technology,” there are feasible economic models for sustainable revenues. D'Alessandro et al. provide an examination of issues challenging the development of a well-known medical Web site, the University of Iowa Virtual Hospital. Their paper describes how they handle issues of content ownership, access, and archiving. Lehmann and Goodman8 and Tarczy-Hornoch et al.9 give us a glimpse of the future, showing how the Web has the potential to enhance access to and use of health information beyond being a store and dissemination medium. Lehmann, well known for advocating the use of Bayesian decision making by practicing clinicians, presents a model of how this technique could be made available in a practical way. Tarczy-Hornoch et al. focus on access to information on genetic testing. As we uncover more of the genome and the results of these studies make their way into clinical information, it will be increasingly important for frontline clinicians to have access to that information in synthesized forms. Clearly, these papers are but two examples of a great deal of work in progress that explores the potential of information in a digital world. We unfortunately did not have room for more. We conclude that an exciting era of publishing of scientific information is at hand. We do recognize that information quality control and dissemination have real costs and truly can never be “free.” But we also share the alarm over the growing cost and inaccessibility of information for economic reasons. The new digital technologies are profoundly changing the relationships between producers, middlemen, and consumers of information in all aspects of our lives—including commerce, finance, education, health care, government, and entertainment. Biomedicine is no different, and we are optimistic that models will emerge that allow effective symbiotic relationships among the players and stakeholders. This should, in turn, usher in a new era of unprecedented means of facilitating scientific collaborations and access and interaction with health information that will benefit society overall.
William R. Hersh, Thomas C. Rindfleisch
J. Am. Medical Informatics Assoc.1
2000 Review Paper: Integration and Beyond: Linking Information from Disparate Sources and into Workflow
abstract
The vision of integrating information-from a variety of sources, into the way people work, to improve decisions and process-is one of the cornerstones of biomedical informatics. Thoughts on how this vision might be realized have evolved as improvements in information and communication technologies, together with discoveries in biomedical informatics, and have changed the art of the possible. This review identified three distinct generations of "integration" projects. First-generation projects create a database and use it for multiple purposes. Second-generation projects integrate by bringing information from various sources together through enterprise information architecture. Third-generation projects inter-relate disparate but accessible information sources to provide the appearance of integration. The review suggests that the ideas developed in the earlier generations have not been supplanted by ideas from subsequent generations. Instead, the ideas represent a continuum of progress along the three dimensions of workflow, structure, and extraction.
William W. Stead, Randolph A. Miller, Mark A. Musen, William R. Hersh
J. Am. Medical Informatics Assoc.4
2000 Discussion Forum: Integration and Beyond: Panel Discussion
abstract
This is the edited transcript of a discussion, among the authors and audience, that followed the presentation that led to the paper “Integration and Beyond: Linking Information from Disparate Sources and into Workflow,” which appears on p. 135. Mark Musen: Bill, when you presented the three generations of integration, the implication was that the third generation is at hand. It was all present tense. I think all of us agree that architectures that allow us to encapsulate knowledge and data in ways that permit reuse are quite exciting. But are we really in the present tense? Have we really achieved these kinds of architectures and, in particular, when you go to the vendor demonstrations, what do you see of this? Bill Stead: That is a very interesting question, because I think the third generation is more in hand than the second. I think “generation” may be the wrong word, because it suggests that the third supplants the second. Instead, techniques from each of the generations coexist in equilibrium. For example, the UMLS provides us with mapping between codes for the kind of things you order for a patient (diagnosis, tests, medications) and the literature. At Vanderbilt we use this mapping to let you ask, “What are the references relevant to the things that have been ordered?” So, to that degree, third generation exists. It is the second generation that is really hard, because it requires regularization. There's a difference between the Vanderbilt vegetable and the Columbia MED, in that the Columbia MED relates the source vocabularies of the various feed systems (third generation), whereas the Vanderbilt vegetable tries to build an enterprise-wide source vocabulary that is then reflected back into the source systems, prealigning their vocabularies. Back to your question, 3M is an example of a vendor that has pursued an architectural strategy. Member of the audience: I think one of the toughest things we all have to deal with is updating our dictionaries. In the simplest cases, the name of an organism is changed and we just have to do the maintenance. It is tougher, when, as with Citrobacter, they do genetic studies and say, “Oh, it's really six different organisms, not one.” We have the human genome project coming very quickly. Even that is just the tip of the iceberg. We're not only going to see all the genes; we're then going to see clinical tests based on gene expression. Essentially, you'll be able to look at something on the order of 180,000 gene products and whether they're up or down regulated. How are we going to integrate such an incredible amount of data at a time when we're going to also be changing how we think about these processes? Classification and simple mapping are not going to work, because the lumpers and splitters are going to be arguing furiously on a daily basis. Randy Miller: The problems you mentioned are clearly on the horizon and very important. But at a simple level, people are people and all of what you're talking about doesn't change how people will present to their primary care providers. At least that part of what exists will not get torn apart. I think what you're talking about is very rich, very vast information overlays on top of what we already have. We don't have to throw out what we have, we need to be ready to extend the linkages. How that will be done is an unanswered question that will result in multiple research grants. Bill Hersh: I think you allude to one of the key points, which is structuring the metadata with the right levels of granularity. Clearly, when we find an organism that can't fit in the existing framework, then that's problem. But if we find that an organism just represents a subcategory of others, and if there's a good hierarchic structure, it can be fit in. The same goes, for example, for diabetes. People classify diabetes with this complication and that complication, but often we just want to know whether the patient has diabetes. Again, a good hierarchic metadata structure can overcome some of those problems. I think we also need to recognize some of the practical limitations that face us. There are limits to the accuracy of the information that's in medical records; there are limits to the consistency in which people apply vocabulary terms. Computers can be completely precise in terms of mapping from this to that, but people will continue to have different conceptions of what a “grade II systolic murmur” is. Bill Stead: I agree with both answers, but I want to continue to clarify what we are talking about. We get in trouble because people use words to reduce concepts to something that we can manage in our heads. So we lump, and person A lumps differently from person B. So we are each a “legacy system,” and our information resources have grown from this starting point. I think we need to work at two ends of the spectrum. Whenever possible, capture data according to granular definitions. If we have an organism and we discover that it splits into six organisms, that's actually a very easy problem to solve, as you said. What you've got to do is say, “A is now B, C, and D and it mapped here.” That is straightforward. That's the end of the spectrum where we can stay granular. For example, never store a doctor and the doctor's service as one piece of information. At the other end of the spectrum, where the granular definitions are not obvious, do not try to classify the data. Instead, tag a “clump” of information with metadata. This tagging, together with increasingly sophisticated extraction techniques, will be used to approximate meaning. Over time, we will get to a complete set of coded data by working from the two ends. Mark Musen: I'm not sure that everything will ever be completely coded. Given the fact that the world is continuously changing, I don't think we can assume that Aristotle was correct that eventually there will be a classification that we will all accept. For example, I do not know whether gastric ulcer is an infectious disease or a gastrointestinal disease, and maybe it is both. As we continue to learn more about medicine and as our organizations change out from under us, I think we're going to be in the situation where the way we categorize the world is going to change. This is very hard stuff. Instead of working on the ultimate classification that will have all of the problems of the International Classification of Diseases, we need to build structures that not only allow us to enumerate the kinds of data that our programs operate on, but attempt as best as we can to enumerate the assumptions that we're making about our data and about the world. Then, as things change, we can, as human beings, try to update our ontologies. I think we have to be able to deal with changing worlds and with the fact that people and computers each need different views on the data, and that means different assumptions as well. Bill Hersh: To reiterate Mark's point, some people have heard this quote, that “perfect is the enemy of good.” We, especially us academic types, strive for perfection, but in reality the world is not perfect, and I don't know that everything will be perfectly coded. But we can reach compromises, such that we can code bits of information that enable us to do useful things. Bill Stead: I think human beings are each different, but we have an underlying genetic code that we are in the process of discovering. Next, we are going to have to work out the problem of going from genotype to phenotype. When I say that I think in the end things will be coded, I think we're going to discover something that is to information what DNA is to people. It will be a very granular base set of building blocks, which will be rolled up into concepts much as genes produce proteins. So I do not want to go to one ontology or one classification. Still, I like having ontologies, particularly ones that clearly represent the difference between themselves and the others. Member of the audience: I'd like to ask a question about capturing ontologies from multiple people. Imagine for a moment that knowledge freezes long enough for us to try to catch it. Do you have a vision of a tool that will allow multiple knowledge-domain people to act at once? To work out discrepancies in their visions? Mark Musen: Put differently, the question was how do we deal with the fact that there is no overarching ontology? How do we build the tools that will allow us to try to achieve consensus in ontologies? I think the answer to that question is that we do not know. I'm being a little bit facetious, but philosophers have been trying to deal with that problem for 2,000 to 3,000 years. I think you see two different approaches in the computer science community. You see the approach that Doug Lenat has taken. He is trying to create an ontology that he believes will provide all the knowledge that one needs to read the Encyclopaedia Britannica. Such an overarching ontology would need to capture most of human existence. The real problem, though, is how you ever validate the distinctions made in that ontology and have confidence that things have been captured in a way that is consistent and understandable? How do you record all the assumptions that you make while constructing the ontology? When you have concepts like “semi-tangible object” and “semi-intangible object,” it's very hard to know for sure whether what one records about those distinctions really makes sense. At the other end of the spectrum, you see people who really want a thousand flowers to bloom and who are not trying to achieve that kind of perfect alignment among views of the world. For example, the Knowledge Systems Laboratory at Stanford is trying to make constrained ontologies that deal with very narrow domains, so that the kinds of problems that you allude to do not happen, because the number of concepts in the ontology is relatively small. The answer lies somewhere between Doug Lenat's view of the world, that all we have to do is work hard enough and everything will fall into place, and the view that we can't possibly do this, so we have to have just a small number of constrained ontologies. We need to elucidate a set of principles that will provide the basis for tools that will help us try to, if not merge small ontologies, at least create the kinds of alignments that will allow us to bring them together in ways that make them useful. Randy Miller: One of the things that I learned from my mentor, Jack Myers, is that as an informatician, as opposed to a philosopher or a computer scientist, you do not need to represent everything. If you have a problem at hand, you represent it at a level that is tractable and doable. If you do what Doug Lenat's doing, you can spend your entire career representing stuff that is not ever going to be used in a real system, because there is no way to apply it. While that may sound harsh, the reality is that we do not know how to represent time, severity of finding, and severity of illness well at all, but we can still build systems that do diagnosis or a good job of making recommendations for therapy. So you do not have to capture the world in all its infinite detail. The trick is to understand what the critical information is and represent things at that level. Otherwise, you get mired in detail. Mark Musen: Let me underscore your last point. Doug Lenat actually felt pretty confident that his ontology covered all the areas that one would want to deal with, until last year, when HotBot contracted to use CYC as the basis for indexing Web pages. This contract showed, first of all, that ontologies have incredible commercial potential, but it also pointed out to Doug Lenat that there was a whole realm of human experience that was not well represented in the ontology. Specifically, there was a need to categorize different kinds of pornography, which Lenat had not thought about previously. Member of the audience: Health Level Seven's development of a set of reference information models is one of the major efforts for creating a structure for ontologies in the United States. Can you talk about how your organizations are participating in the development of that reference information model (RIM) and how you are using your academic experiences to contribute to that effort among providers, academics, and vendors? Bill Stead: Vanderbilt is an institutional member and a strong advocate of HL7. The central core of our communication subsystem uses HL7, and we build middle ware as needed to bridge between the core and legacy products. We have not put direct energy into the process for defining the reference information model. We use the HL7 model as a starting point, but we extend it as needed. In this way we incorporate it into immediate solutions to real problems, while providing useful information about future directions. Bill Hersh: None of us has been involved directly in that effort. However, our research into the nature of ontologies and the vocabulary projects such as the Cannon Grouping should useful to the effort. Mark Musen: I will just add that I think the vendor community is in the best position to work on ontology content, because they have the most direct connection with the needs of end users. I think that academicians need to follow this work very carefully. We are, we hope, in the best position to be developing the kinds of tools that will help us examine ontologies, relate them to each other, and allow them to evolve as our understanding of the world changes. Randy Miller: I have a slightly contrary view, partly out of ignorance about HL7 RIM. The key question is what problems it is trying to solve. That should drive what the content is. If you can state the problems it is going to be used to solve, then you can say whether it should clinically rich. In that case it will require lots of input from academic clinicians. If it is to solve the problem of interchange of data among vendors, then it needs vendor input. But until you explicitly state what it's going to be used for, just building it for the sake of building it is not useful. I know that the HL7 RIM is not being built that way. I am just saying that I think that's the way to address your question, to seek the specific purpose before giving an answer.
William W. Stead, Randolph A. Miller, Mark A. Musen, William R. Hersh
J. Am. Medical Informatics Assoc.4
1999 Perceptions of house officers who use physician order entry
Joan S. Ash, Paul N. Gorman, William R. Hersh, Mary Lavelle, Susan B. Poulsen
AMIA3
1999 Implementation and evaluation of a virtual learning center for distributed education
Katherine A. Caton, William R. Hersh
AMIA2
1999 Maintaining a catalog of manually-indexed, clinically-oriented World Wide Web content
William R. Hersh, Andrea Ball, Bikram Day, Mary Masterson, Lynetta Sacherek
AMIA1
1999 Teaching English Medical Terminology Using the UMLS Metathesaurus and World Wide Web
William R. Hersh
AMIA1
1999 Videophones for Wellness Coaching in an Overweight Population
Patricia K. Patterson, Catherine Salveson, William R. Hersh, P. Barton Duell, Judith Carlson
AMIA3
1999 Filtering Web pages for quality indicators: an empirical approach to finding high quality consumer health information on the World Wide Web
Susan Price, William R. Hersh
AMIA2
1999 Model Formulation: A Model for Enhancing Internet Medical Document Retrieval with "Medical Core Metadata"
abstract
OBJECTIVE: Finding documents on the World Wide Web relevant to a specific medical information need can be difficult. The goal of this work is to define a set of document content description tags, or metadata encodings, that can be used to promote disciplined search access to Internet medical documents. DESIGN: The authors based their approach on a proposed metadata standard, the Dublin Core Metadata Element Set, which has recently been submitted to the Internet Engineering Task Force. Their model also incorporates the National Library of Medicine's Medical Subject Headings (MeSH) vocabulary and MEDLINE-type content descriptions. RESULTS: The model defines a medical core metadata set that can be used to describe the metadata for a wide variety of Internet documents. CONCLUSIONS: The authors propose that their medical core metadata set be used to assign metadata to medical documents to facilitate document retrieval by Internet search engines.
Gary Malet, Felix Munoz, Richard Appleyard, William R. Hersh
J. Am. Medical Informatics Assoc.4
1998 Physician order entry in U.S. hospitals
Joan S. Ash, Paul N. Gorman, William R. Hersh
AMIA3
1998 Information retrieval at the millenium
William R. Hersh
AMIA1
1998 SAPHIRE International: a tool for cross-language information retrieval
William R. Hersh, Larry Donohoe
AMIA1
1998 DynamiQuest: A System for Asking Clinical Examination Questions Based on Content Viewed
William R. Hersh, Susan Malveau
AMIA1
1998 A Needs Assessment of a Web-Based Smoking Cessation Program
Calvin G. Huey, Linda Lucas, Karen B. Eden, William R. Hersh
AMIA4
1998 Towards knowledge-based retrieval of medical images. The role of semantic indexing, image content representation and knowledge-based retrieval
Henry J. Lowe, Ilya Antipov, William R. Hersh, Catherine Arnott Smith
AMIA3
1998 MCM generator: a Java-based tool for generating medical metadata
Felix Munoz, William R. Hersh
AMIA2
1998 Factors Influencing Successful Use of Information Retrieval Systems by Nurse Practitioner Students
Linda Rose, M. Katherine Crabtree, William R. Hersh
AMIA3
1998 Developing search strategies for detecting high quality reviews in a hypertext test collection
Michael P. Zacks, William R. Hersh
AMIA2
1998 Book Review: Rationalizing Medical Work: Decision-Support Techniques and Medical Practices by M. Berg
William R. Hersh
Inf. Process. Manag.1
1997 ORCAS: Network-Based Management of Cancer Clinical Trials
Richard Brodner, William R. Hersh, Mark Helfand, Elizabeth Brown
AMIA2
1997 MedWeaver: integrating decision support, literature searching, and Web exploration using the UMLS Metathesaurus
William M. Detmer, G. Octo Barnett, William R. Hersh
AMIA3
1997 Assessing the feasibility of large-scale natural language processing in a corpus of ordinary medical records: a lexical analysis
William R. Hersh, Emily M. Campbell, Susan Malveau
AMIA1
1996 Applications of Technology: Clini Web: Managing Clinical Information on the World Wide Web
abstract
The World Wide Web is a powerful new way to deliver on-line clinical information, but several problems limit its value to health care professionals: content is highly distributed and difficult to find, clinical information is not separated from non-clinical information, and the current Web technology is unable to support some advanced retrieval capabilities. A system called CliniWeb has been developed to address these problems. CliniWeb is an index to clinical information on the World Wide Web, providing a browsing and searching interface to clinical content at the level of the health care student or provider. Its database contains a list of clinical information resources on the Web that are indexed by terms from the Medical Subject Headings disease tree and retrieved with the assistance of SAPHIRE. Limitations of the processes used to build the database are discussed, together with directions for future research.
William R. Hersh, Keven E. Brown, Larry Donohoe, Emily M. Campbell, Ashley E. Horacek
J. Am. Medical Informatics Assoc.1
1996 A Task-Oriented Approach to Information Retrieval Evaluation
abstract
As retrieval systems become more oriented towards end-users, there is an increasing need for improved methods to evaluate their effectiveness. We performed a task-oriented assessment of two MEDLINE searching systems, one which promotes traditional Boolean searching on human-indexed thesaurus terms and the other natural language searching on words in the title, abstract, and indexing terms. Medical students were randomized to one of the two systems and given clinical questions to answer. The students were able to use each system successfully, with no significant differences in questions correctly answered, time taken, relevant articles retrieved, or user satisfaction between the systems. This approach to evaluation was successful in measuring effectiveness of system use and demonstrates that both types of systems can be used equally well with minimal training. © 1996 John Wiley & Sons, Inc.
William R. Hersh, Jeffrey Pentecost, David H. Hickam
J. Am. Soc. Inf. Sci.1
1995 Towards New Measures of Information Retrieval Evaluation
abstract
All of the methods cumently used to evaluate information retrieval (Ill) systems have limitations in their ability to measure how well users are able to acquire information.We utilized on approach to assessing information obtained based on the user's ability to answer qtrcstions from a shortanswer test.Senior medical students took the ten-question test and then searched one of (WO IR systems on the five questions for which they were least ccrtuin of their answer.Our results showed thtit pre-searching scores on the test were low but that searching yielded a high proportion of answers with both systems.These methods are able to measure information obtained, and will be used in subsequent studies to assess differences among IR systems.provide a starting point at determining the quantity of useful information ob[ained from an IR system.they say Pemlission to nmke digil:il/h:lrd copies u 1'all or port i) 1'tl)is m:ttcri:ll without fee is grontcd provided [Il:it the co}>ics :irc not ll]:l(jc or distributed for profit or commercini ~LiV:lllhgC, lllc ACM copyriglll/ server notice, the tillc of lhc pu[)]ic?:ltion Jnd i(s d;i[u :tppc:tr, :111'j notice is given that copyright is by permission of tllc Assi)ci:ltiol] for Computin
William R. Hersh, Diane L. Elliot, David H. Hickam, Stephanie L. Wolf, Anna Molnar, Christine Leichtenstien
SIGIR1
1995 Research Paper: The Canon Group's Effort: Working Toward a Merged Model
abstract
OBJECTIVE: To develop a representational schema for clinical data for use in exchanging data and applications, using a collaborative approach. DESIGN: Representational models for clinical radiology were independently developed manually by several Canon Group members who had diverse application interests, using sample reports. These models were merged into one common model through an iterative process by means of workshops, meetings, and electronic mail. RESULTS: A core merged model for radiologic findings present in a set of reports that subsumed the models that were developed independently. CONCLUSIONS: The Canon Group's modeling effort focused on a collaborative approach to developing a representational schema for clinical concepts, using chest radiography reports as the initial experiment. This effort resulted in a core model that represents a consensus. Further efforts in modeling will extend the representational coverage and will also address issues such as scalability, automation, evaluation, and support of the collaborative effort.
Carol Friedman, Stanley M. Huff, William R. Hersh, Edward Pattison-Gordon, James J. Cimino
J. Am. Medical Informatics Assoc.3
1995 The Electronic Medical Record: Promises and Problems
abstract
Despite the growth of computer technology in medicine, most medical encounters are still documented on paper medical records. The electronic medical record has numerous documented benefits, yet its use is still sparse. This article describes the state of electronic medical records, their advantage over existing paper records, the problems impeding their implementation, and concerns over their security and confidentiality. © 1995 John Wiley & Sons, Inc.
William R. Hersh
J. Am. Soc. Inf. Sci.1
1995 An Evaluation of Interactive Boolean and Natural Language Searching with an Online Medical Textbook
abstract
Few studies have compared the interactive use of Boolean and natural language searching systems. We studied the use of three retrieval systems by senior medical students searching on queries generated by actual physicians in a clinical setting. The searchers were randomized to search on two or three different retrieval systems: a Boolean system, a word-based natural language system, and a concept-based natural language system. Our results showed no statistically significant differences in recall or precision among the three systems. Likewise, we found no user preference for any system over the others. In the course of this study we did find, however, a number of problems with traditional measures of retrieval evaluation when applied to the interactive search setting. © 1995 John Wiley & Sons, Inc.
William R. Hersh, David H. Hickam
J. Am. Soc. Inf. Sci.1
1995 Erratum: An Evaluation of Interactive Boolean and Natural Language Searching with an Online Medical Textbook. August 1995 Issue, pages 478-489
William R. Hersh, David H. Hickam
J. Am. Soc. Inf. Sci.1
1995 Information Retrieval in Medicine: The SAPHIRE Experience
abstract
Information retrieval systems are being used increasingly in biomedical settings, but many problems still exist in indexing, retrieval, and evaluation. The SAPHIRE Project was undertaken to seek solutions for these problems. This article summarizes the evaluation studies that have been done with SAPHIRE, highlighting the lessons learned and laying out the challenges ahead to all medical information retrieval efforts. © 1995 John Wiley & Sons, Inc.
William R. Hersh, David H. Hickam
J. Am. Soc. Inf. Sci.1
1995 Introduction and Overview
William R. Hersh, Lois F. Lunin
J. Am. Soc. Inf. Sci.1
1994 OHSUMED: An Interactive Retrieval Evaluation and New Large Test Collection for Research
William R. Hersh, Chris Buckley, T. J. Leone, David H. Hickam
SIGIR1
1994 Position Paper: Toward a Medical-concept Representation Language
abstract
The Canon Group is an informal organization of medical informatics researchers who are working on the problem of developing a "deeper" representation formalism for use in exchanging data and developing applications. Individuals in the group represent experts in such areas as knowledge representation and computational linguistics, as well as in a variety of medical subdisciplines. All share the view that current mechanisms for the characterization of medical phenomena are either inadequate (limited or rigid) or idiosyncratic (useful for a specific application but incapable of being generalized or extended). The Group proposes to focus on the design of a general schema for medical-language representation including the specification of the resources and associated procedures required to map language (including standard terminologies) into representations that make all implicit relations "visible," reveal "hidden attributes," and generally resolve ambiguous or vague references. The Group is proceeding by examining large numbers of texts (records) in medical sub-domains to identify candidate "concepts" and by attempting to develop general rules and representations for elements such as attributes and values so that all concepts may be expressed uniformly.
David A. Evans 0001, James J. Cimino, William R. Hersh, Stanley M. Huff, Douglas S. Bell
J. Am. Medical Informatics Assoc.3
1994 Research Paper: A Performance and Failure Analysis of SAPHIRE with a MEDLINE Test Collection
abstract
OBJECTIVE: Assess the performance of the SAPHIRE automated information retrieval system. DESIGN: Comparative study of automated and human searching of a MEDLINE test collection. MEASUREMENTS: Recall and precision of SAPHIRE were compared with those attributes of novice physicians, expert physicians, and librarians for a test collection of 75 queries and 2,334 citations. Failure analysis assessed the efficacy of the Metathesaurus as a concept vocabulary; the reasons for retrieval of nonrelevant articles and nonretrieval of relevant articles; and the effect of changing the weighting formula for relevance ranking of retrieved articles. RESULTS: Recall and precision of SAPHIRE were comparable to those of both physician groups, but less than those of librarians. CONCLUSION: The current version of the Metathesaurus, as utilized by SAPHIRE, was unable to represent the conceptual content of one-fourth of physician-generated MEDLINE queries. The most likely cause for retrieval of nonrelevant articles was the presence of some or all of the search terms in the article, with frequencies high enough to lead to retrieval. The most likely cause for nonretrieval of relevant articles was the absence of the actual terms from the query, with synonyms or hierarchically related terms present instead. There were significant variations in performance when SAPHIRE's concept-weighing formulas were modified.
William R. Hersh, David H. Hickam, Robert Brian Haynes, K. Ann McKibbon
J. Am. Medical Informatics Assoc.1
1994 Relevance and Retrieval Evaluation: Perspectives from Medicine
abstract
The traditional notion of topical relevance has allowed much useful work to be done in the evaluation of retrieval systems, but has limitations for complete assessment of retrieval systems. While topical relevance can be effective in evaluating various indexing and retrieval approaches, it is ineffective for measuring the impact that systems have on users. An alternative is to use a more situational definition of relevance, which takes account of the impact of the system on the user. Both types of relevance are examined from the standpoint of the medical domain, concluding that each have their appropriate use. But in medicine there is increasing emphasis on outcomes-oriented research which, when applied to information science, requires that the impact of an information system on the activities which prompt its use be assessed. An iterative model of retrieval evaluation is proposed, starting first with the use of topical relevance to insure documents on the subject can be retrieved. This is followed by the use of situational relevance to show the user can interact positively with the system. The final step is to study how the system impacts the user in the purpose for which the system was consulted, which can be done by methods such as protocol analysis and simulation. These diverse types of studies are necessary to increase our understanding of the nature of retrieval systems. © 1994 John Wiley & Sons, Inc.
William R. Hersh
J. Am. Soc. Inf. Sci.1