Michael G. Kahn

dblp:59/256 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0003-4786-6875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 30 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorSecurity and privacy · 2Databases, data management, data science and information retrieval · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Linkability measures to assess the data characteristics for record linkage
abstract
OBJECTIVES: Accurate record linkage (RL) enables consolidation and de-duplication of data from disparate datasets, resulting in more comprehensive and complete patient data. However, conducting RL with low quality or unfit data can waste institutional resources on poor linkage results. We aim to evaluate data linkability to enhance the effectiveness of record linkage. MATERIALS AND METHODS: We describe a systematic approach using data fitness ("linkability") measures, defined as metrics that characterize the availability, discriminatory power, and distribution of potential variables for RL. We used the isolation forest algorithm to detect abnormal linkability values from 188 sites in Indiana and Colorado, and manually reviewed the data to understand the cause of anomalies. RESULT: We calculated 10 linkability metrics for 11 potential linkage variables (LVs) across 188 sites for a total of 20 680 linkability metrics. Potential LVs such as first name, last name, date of birth, and sex have low missing data rates, while Social Security Number vary widely in completeness among all sites. We investigated anomalous linkability values to identify the cause of many records having identical values in certain LVs, issues with placeholder values disguising data missingness, and orphan records. DISCUSSION: The fitness of a variable for RL is determined by its availability and its discriminatory power to uniquely identify individuals. These results highlight the need for awareness of placeholder values, which inform the selection of variables and methods to optimize RL performance. CONCLUSION: Evaluating linkability measures using the isolation forest algorithm to highlight anomalous findings can help identify fitness-for-use issues that must be addressed before initiating the RL process to ensure high-quality linkage outcomes.
Toan Ong, Michael G. Kahn, Lauren R. Lembcke, Lisa M. Schilling, Shaun J. Grannis
J. Am. Medical Informatics Assoc.3
2022 Characterizing Patient Representations for Computational Phenotyping
Tiffany Callahan, Adrianne L. Stefanski, Danielle Ostendorf, Jordan M. Wyrwa, Sara J. Deakyne Davies, George Hripcsak, Lawrence Hunter, Michael G. Kahn
AMIA8
2022 Migrating a research data warehouse to a public cloud: challenges and opportunities
abstract
OBJECTIVE: Clinical research data warehouses (RDWs) linked to genomic pipelines and open data archives are being created to support innovative, complex data-driven discoveries. The computing and storage needs of these research environments may quickly exceed the capacity of on-premises systems. New RDWs are migrating to cloud platforms for the scalability and flexibility needed to meet these challenges. We describe our experience in migrating a multi-institutional RDW to a public cloud. MATERIALS AND METHODS: This study is descriptive. Primary materials included internal and public presentations before and after the transition, analysis documents, and actual billing records. Findings were aggregated into topical categories. RESULTS: Eight categories of migration issues were identified. Unanticipated challenges included legacy system limitations; network, computing, and storage architectures that realize performance and cost benefits in the face of hyper-innovation, complex security reviews and approvals, and limited cloud consulting expertise. DISCUSSION: Cloud architectures enable previously unavailable capabilities, but numerous pitfalls can impede realizing the full benefits of a cloud environment. Rapid changes in cloud capabilities can quickly obsolete existing architectures and associated institutional policies. Touchpoints with on-premise networks and systems can add unforeseen complexity. Governance, resource management, and cost oversight are critical to allow rapid innovation while minimizing wasted resources and unnecessary costs. CONCLUSIONS: Migrating our RDW to the cloud has enabled capabilities and innovations that would not have been possible with an on-premises environment. Notwithstanding the challenges of managing cloud resources, the resulting RDW capabilities have been highly positive to our institution, research community, and partners.
Michael G. Kahn, Joyce Y. Mui, Michael Ames, Anoop K. Yamsani, Nikita Pozdeyev, Nicholas Rafaels, Ian M. Brooks
J. Am. Medical Informatics Assoc.1
2021 Quality assessment of real-world data repositories across the data life cycle: A literature review
abstract
OBJECTIVE: Data quality (DQ) must be consistently defined in context. The attributes, metadata, and context of longitudinal real-world data (RWD) have not been formalized for quality improvement across the data production and curation life cycle. We sought to complete a literature review on DQ assessment frameworks, indicators and tools for research, public health, service, and quality improvement across the data life cycle. MATERIALS AND METHODS: The review followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines. Databases from health, physical and social sciences were used: Cinahl, Embase, Scopus, ProQuest, Emcare, PsycINFO, Compendex, and Inspec. Embase was used instead of PubMed (an interface to search MEDLINE) because it includes all MeSH (Medical Subject Headings) terms used and journals in MEDLINE as well as additional unique journals and conference abstracts. A combined data life cycle and quality framework guided the search of published and gray literature for DQ frameworks, indicators, and tools. At least 2 authors independently identified articles for inclusion and extracted and categorized DQ concepts and constructs. All authors discussed findings iteratively until consensus was reached. RESULTS: The 120 included articles yielded concepts related to contextual (data source, custodian, and user) and technical (interoperability) factors across the data life cycle. Contextual DQ subcategories included relevance, usability, accessibility, timeliness, and trust. Well-tested computable DQ indicators and assessment tools were also found. CONCLUSIONS: A DQ assessment framework that covers intrinsic, technical, and contextual categories across the data life cycle enables assessment and management of RWD repositories to ensure fitness for purpose. Balancing security, privacy, and FAIR principles requires trust and reciprocity, transparent governance, and organizational cultures that value good documentation.
Siaw-Teng Liaw, Jason Guan Nan Guo, Sameera Ansari, Jitendra Jonnagaddala, Myron Anthony Godinho, Alder Jose Borelli, Simon de Lusignan, Daniel Capurro, Harshana Liyanage, Navreet Bhattal, Vicki Bennett, Jaclyn Chan, Michael G. Kahn
J. Am. Medical Informatics Assoc.13
2020 Novel Application of Data Quality Metrics to Tailor Standardization of Patient Matching Fields
Shaun J. Grannis, Huiping Xu, Toan Ong, Michael G. Kahn, Lauren R. Lembcke, Suranga Nath Kasthurirathne
AMIA5
2020 An Autocorrelation-based LSTM-Autoencoder for Anomaly Detection on Time-Series Data
abstract
Data quality significantly impacts the results of data analytics. Researchers have proposed machine learning based anomaly detection techniques to identify incorrect data. Existing approaches fail to (1) identify the underlying domain constraints violated by the anomalous data, and (2) generate explanations of these violations in a form comprehensible to domain experts. We propose IDEAL, which is an LSTM-Autoencoder based approach that detects anomalies in multivariate time-series data, generates domain constraints, and reports subsequences that violate the constraints as anomalies. We propose an automated autocorrelation-based windowing approach to adjust the network input size, thereby improving the correctness and performance of constraint discovery over manual and brute-force approaches. The anomalies are visualized in a manner comprehensible to domain experts in the form of decision trees extracted from a random forest classifier. Domain experts can then provide feedback to retrain the learning model and improve the accuracy of the process. We evaluate the effectiveness of IDEAL using datasets from Yahoo servers, NASA Shuttle, and Colorado State University Energy Institute. We demonstrate that IDEAL can detect previously known anomalies from these datasets. Using mutation analysis, we show that IDEAL can detect different types of injected faults. We also demonstrate that the accuracy improves after incorporating domain expert feedback.
Hajar Homayouni, Sudipto Ghosh 0001, Indrakshi Ray, Shlok Gondalia, Jerry Duggan, Michael G. Kahn
IEEE BigData6
2020 A hybrid approach to record linkage using a combination of deterministic and probabilistic methodology
abstract
OBJECTIVE: The disjointed healthcare system and the nonexistence of a universal patient identifier across systems necessitates accurate record linkage (RL). We aim to describe the implementation and evaluation of a hybrid record linkage method in a statewide surveillance system for congenital heart disease. MATERIALS AND METHODS: Clear-text personally identifiable information on individuals in the Colorado Congenital Heart Disease surveillance system was obtained from 5 electronic health record and medical claims data sources. Two deterministic methods and 1 probabilistic RL method using first name, last name, social security number, date of birth, and house number were initially implemented independently and then sequentially in a hybrid approach to assess RL performance. RESULTS: 16 480 nonunique individuals with congenital heart disease were ascertained. Deterministic linkage methods, when performed independently, yielded 4505 linked pairs (consisting of 2 records linked together within or across data sources). Probabilistic RL, using 3 initial characters of last name and gender for blocking, yielded 6294 linked pairs when executed independently. Using a hybrid linkage routine resulted in 6451 linkages and an additional 18%-24% correct linked pairs as compared to the independent methods. A hybrid linkage routine resulted in higher recall and F-measure scores compared to probabilistic and deterministic methods performed independently. DISCUSSION: The hybrid approach resulted in increased linkage accuracy and identified pairs of linked record that would have otherwise been missed when using any independent linkage technique. CONCLUSION: When performing RL within and across disparate data sources, the hybrid RL routine outperformed independent deterministic and probabilistic methods.
Toan Ong, Lindsey M. Duca, Michael G. Kahn, Tessa L. Crume
J. Am. Medical Informatics Assoc.3
2019 An Interactive Data Quality Test Approach for Constraint Discovery and Fault Detection
abstract
Data quality tests validate heterogeneous data to detect violations of syntactic and semantic constraints. The specification of these constraints can be incomplete because domain experts typically specify them in an ad hoc manner. Existing automated test approaches can generate false alarms and do not explain the constraint violations while reporting faulty data records. In previous work, we proposed ADQuaTe, which is an automated data quality test approach that uses an unsupervised deep learning techni que (1) to discover constraints from big datasets that may have been missed by experts, and (2) to label as suspicious those records that violate the constraints. These records are grouped and explanations for constraint violations are presented to domain experts who determine whether or not the groups are actually faulty. This paper presents ADQuaTe2, which extends ADQuaTe to use an interactive learning technique that incorporates expert feedback to retrain the learning model and improve the accuracy of constraint discovery and fault detection. We evaluate the effectiveness of the approach on real-world datasets from a health data warehouse and a plant diagnosis database. We also use datasets with known faults from the UCI repository to evaluate the improvement in the accuracy of the approach after incorporating ground truth knowledge.
Hajar Homayouni, Sudipto Ghosh 0001, Indrakshi Ray, Michael G. Kahn
IEEE BigData4
2017 A longitudinal analysis of data quality in a large pediatric data research network
abstract
OBJECTIVE: PEDSnet is a clinical data research network (CDRN) that aggregates electronic health record data from multiple children's hospitals to enable large-scale research. Assessing data quality to ensure suitability for conducting research is a key requirement in PEDSnet. This study presents a range of data quality issues identified over a period of 18 months and interprets them to evaluate the research capacity of PEDSnet. MATERIALS AND METHODS: Results were generated by a semiautomated data quality assessment workflow. Two investigators reviewed programmatic data quality issues and conducted discussions with the data partners' extract-transform-load analysts to determine the cause for each issue. RESULTS: The results include a longitudinal summary of 2182 data quality issues identified across 9 data submission cycles. The metadata from the most recent cycle includes annotations for 850 issues: most frequent types, including missing data (>300) and outliers (>100); most complex domains, including medications (>160) and lab measurements (>140); and primary causes, including source data characteristics (83%) and extract-transform-load errors (9%). DISCUSSION: The longitudinal findings demonstrate the network's evolution from identifying difficulties with aligning the data to a common data model to learning norms in clinical pediatrics and determining research capability. CONCLUSION: While data quality is recognized as a critical aspect in establishing and utilizing a CDRN, the findings from data quality assessments are largely unpublished. This paper presents a real-world account of studying and interpreting data quality findings in a pediatric CDRN, and the lessons learned could be used by other CDRNs.
Ritu Khare, Levon Utidjian, Byron Ruth, Michael G. Kahn, Evanette Burrows, Keith Marsolo, Nandan Patibandla, Hanieh Razzaghi, Ryan Colvin, Daksha Ranade, Melody Kitzmiller, Daniel Eckrich, L. Charles Bailey
J. Am. Medical Informatics Assoc.4
2017 Pragmatic (trial) informatics: a perspective from the NIH Health Care Systems Research Collaboratory
abstract
Pragmatic clinical trials (PCTs) are research investigations embedded in health care settings designed to increase the efficiency of research and its relevance to clinical practice. The Health Care Systems Research Collaboratory, initiated by the National Institutes of Health Common Fund in 2010, is a pioneering cooperative aimed at identifying and overcoming operational challenges to pragmatic research. Drawing from our experience, we present 4 broad categories of informatics-related challenges: (1) using clinical data for research, (2) integrating data from heterogeneous systems, (3) using electronic health records to support intervention delivery or health system change, and (4) assessing and improving data capture to define study populations and outcomes. These challenges impact the validity, reliability, and integrity of PCTs. Achieving the full potential of PCTs and a learning health system will require meaningful partnerships between health system leadership and operations, and federally driven standards and policies to ensure that future electronic health record systems have the flexibility to support research.
Rachel L. Richesson, Beverly Green, Reesa Laws, Jon Puro, Michael G. Kahn, Alan Bauck, Michelle Smerek, Erik G. Van Eaton, Meredith Nahm, William Edward Hammond, Kari A. Stephens, Greg E. Simon
J. Am. Medical Informatics Assoc.5
2016 PEDSnet: from building a high-quality CDRN to conducting science
L. Charles Bailey, Michael G. Kahn, Sara J. Deakyne Davies, Ritu Khare, Katherine Deans
AMIA2
2016 Privacy Preserving Probabilistic Record Linkage Using Locality Sensitive Hashes
Ibrahim Lazrig, Toan Ong, Indrajit Ray, Indrakshi Ray, Michael G. Kahn
DBSec5
2015 Data Quality in Clinical Data Research Networks (CDRNs)
Allison B. McCoy, Michael G. Kahn, Lemuel R. Waitman, Jason N. Doctor
AMIA2
2015 Using National Database for Autism Research (NDAR) privacy-preserving record linkage protocol in the PEDSnet CDRN
Toan Ong, Michael G. Kahn, Dan J. Hall
AMIA2
2015 Comparing Weight Redistribution and Distance Imputation Methods for Missing Data in Clear-text and Encrypted Record Linkage
Toan Ong, Lisa M. Schilling, Michael G. Kahn
AMIA3
2015 Privacy Preserving Record Matching Using Automated Semi-trusted Broker
Ibrahim Lazrig, Tarik Moataz, Indrajit Ray, Indrakshi Ray, Toan Ong, Michael G. Kahn, Frédéric Cuppens, Nora Cuppens
DBSec6
2014 Brief communication: PEDSnet: a National Pediatric Learning Health System
abstract
A learning health system (LHS) integrates research done in routine care settings, structured data capture during every encounter, and quality improvement processes to rapidly implement advances in new knowledge, all with active and meaningful patient participation. While disease-specific pediatric LHSs have shown tremendous impact on improved clinical outcomes, a national digital architecture to rapidly implement LHSs across multiple pediatric conditions does not exist. PEDSnet is a clinical data research network that provides the infrastructure to support a national pediatric LHS. A consortium consisting of PEDSnet, which includes eight academic medical centers, two existing disease-specific pediatric networks, and two national data partners form the initial partners in the National Pediatric Learning Health System (NPLHS). PEDSnet is implementing a flexible dual data architecture that incorporates two widely used data models and national terminology standards to support multi-institutional data integration, cohort discovery, and advanced analytics that enable rapid learning.
Christopher B. Forrest, Peter A. Margolis, L. Charles Bailey, Keith Marsolo, Mark A. Del Beccaro, Jonathan A. Finkelstein, David E. Milov, Veronica J. Vieland, Bryan A. Wolf, Feliciano B. Yu, Michael G. Kahn
J. Am. Medical Informatics Assoc.11
2014 Brief communication: Developing a data infrastructure for a learning health system: the PORTAL network
abstract
The Kaiser Permanente & Strategic Partners Patient Outcomes Research To Advance Learning (PORTAL) network engages four healthcare delivery systems (Kaiser Permanente, Group Health Cooperative, HealthPartners, and Denver Health) and their affiliated research centers to create a new national network infrastructure that builds on existing relationships among these institutions. PORTAL is enhancing its current capabilities by expanding the scope of the common data model, paying particular attention to incorporating patient-reported data more systematically, implementing new multi-site data governance procedures, and integrating the PCORnet PopMedNet platform across our research centers. PORTAL is partnering with clinical research and patient experts to create cohorts of patients with a common diagnosis (colorectal cancer), a rare diagnosis (adolescents and adults with severe congenital heart disease), and adults who are overweight or obese, including those with pre-diabetes or diabetes, to conduct large-scale observational comparative effectiveness research and pragmatic clinical trials across diverse clinical care settings.
Elizabeth A. McGlynn, Tracy A. Lieu, Mary L. Durham, Alan Bauck, Reesa Laws, Alan S. Go, Jersey Chen, Heather Spencer Feigelson, Douglas A. Corley, Deborah Rohm Young, Andrew F. Nelson, Arthur J. Davidson, Leo S. Morales, Michael G. Kahn
J. Am. Medical Informatics Assoc.14
2014 Improving record linkage performance in the presence of missing linkage data
abstract
INTRODUCTION: Existing record linkage methods do not handle missing linking field values in an efficient and effective manner. The objective of this study is to investigate three novel methods for improving the accuracy and efficiency of record linkage when record linkage fields have missing values. METHODS: By extending the Fellegi-Sunter scoring implementations available in the open-source Fine-grained Record Linkage (FRIL) software system we developed three novel methods to solve the missing data problem in record linkage, which we refer to as: Weight Redistribution, Distance Imputation, and Linkage Expansion. Weight Redistribution removes fields with missing data from the set of quasi-identifiers and redistributes the weight from the missing attribute based on relative proportions across the remaining available linkage fields. Distance Imputation imputes the distance between the missing data fields rather than imputing the missing data value. Linkage Expansion adds previously considered non-linkage fields to the linkage field set to compensate for the missing information in a linkage field. We tested the linkage methods using simulated data sets with varying field value corruption rates. RESULTS: The methods developed had sensitivity ranging from .895 to .992 and positive predictive values (PPV) ranging from .865 to 1 in data sets with low corruption rates. Increased corruption rates lead to decreased sensitivity for all methods. CONCLUSIONS: These new record linkage algorithms show promise in terms of accuracy and efficiency and may be valuable for combining large data sets at the patient level to support biomedical and clinical research.
Toan Ong, Michael V. Mannino, Lisa M. Schilling, Michael G. Kahn
J. Biomed. Informatics4
2013 How Fit is Electronic Health Data for its Intended Uses? Exploring Data Quality across Clinical, Public Health, and Research Use Cases
Shaun J. Grannis, Brian E. Dixon, Siaw-Teng Liaw, Michael G. Kahn, Hamish S. F. Fraser
AMIA4
2013 CloudDRN: A Lightweight, End-to-End System for Sharing Distributed Research Data in the Cloud
abstract
The cloud has proven itself as a scalable platform for Web-based applications. However, scientists and medical researchers are still searching for a simple cloud-based architecture that enables secure collaboration and sharing of distributed datasets. To date, attempts at using the cloud for this purpose generally view the cloud as simply a pool of servers upon which to run their legacy software. This approach fails to leverage the unique platform capabilities of the cloud. In this paper, we describe our Cloud Distributed Research Network (CloudDRN). We leverage the cloud for availability, reliability, scalability, and improved security as compared to legacy distributed systems while still supporting site autonomy. Our philosophy is to adapt commercial software tooling that was originally designed for business use-cases, thereby benefiting from the large built-in user community. We describe our general architecture and show an example of our system created to share distributed clinical research data. We evaluate our system in Amazon Web Services (AWS) and in Microsoft Windows Azure and find that while each cloud achieves similar financial cost, representative queries are 3.5x slower on average in Windows Azure.
Marty Humphrey, Jacob Steele, In Kee Kim, Michael G. Kahn, Jessica Bondy, Michael Ames
e-Science4
2012 Clinical research informatics: a conceptual perspective
abstract
Clinical research informatics is the rapidly evolving sub-discipline within biomedical informatics that focuses on developing new informatics theories, tools, and solutions to accelerate the full translational continuum: basic research to clinical trials (T1), clinical trials to academic health center practice (T2), diffusion and implementation to community practice (T3), and 'real world' outcomes (T4). We present a conceptual model based on an informatics-enabled clinical research workflow, integration across heterogeneous data sources, and core informatics tools and platforms. We use this conceptual model to highlight 18 new articles in the JAMIA special issue on clinical research informatics.
Michael G. Kahn, Chunhua Weng
J. Am. Medical Informatics Assoc.1
2011 Selected Papers from the 2011 Summit on Clinical Research Informatics
Philip R. O. Payne, Peter J. Embí, Michael G. Kahn
J. Biomed. Informatics3
2010 The impact of electronic medical records data sources on an adverse drug event quality measure
abstract
OBJECTIVE: To examine the impact of billing and clinical data extracted from an electronic medical record system on the calculation of an adverse drug event (ADE) quality measure approved for use in The Joint Commission's ORYX program, a mandatory national hospital quality reporting system. DESIGN: The Child Health Corporation of America's "Use of Rescue Agents-ADE Trigger" quality measure uses medication billing data contained in the Pediatric Health Information Systems (PHIS) data warehouse to create The Joint Commission-approved quality measure. Using a similar query, we calculated the quality measure using PHIS plus four data sources extracted from our electronic medical record (EMR) system: medications charged, medication orders placed, medication orders with associated charges (orders charged), and medications administered. MEASUREMENTS: Inclusion and exclusion criteria were identical for all queries. Denominators and numerators were calculated using the five data sets. The reported quality measure is the ADE rate (numerator/denominator). RESULTS: Significant differences in denominators, numerators, and rates were calculated from different data sources within a single institution's EMR. Differences were due to both common clinical practices that may be similar across institutions and unique workflow practices not likely to be present at any other institution. The magnitude of the differences would significantly alter the national comparative ranking of our institution compared to other PHIS institutions. CONCLUSIONS: More detailed clinical information may result in quality measures that are not comparable across institutions due institution-specific workflow, differences that are exposed using EMR-derived data.
Michael G. Kahn, Daksha Ranade
J. Am. Medical Informatics Assoc.1
2005 Clinical Terminology Support for a National Ambulatory Practice Outcomes Research Network
Thomas N. Ricciardi, Michael I. Lieberman, Michael G. Kahn, Fred E. Masarie Jr.
AMIA3
2002 Protocol Design Patterns: Domain-oriented Abstractions to Support the Authoring of Computer-executable Clinical Trials
John H. Nguyen, Michael G. Kahn, Carol A. Broverman, Mark A. Musen
AMIA2
2002 Temporal knowledge representation for scheduling tasks in clinical trial protocols
Chunhua Weng, Michael G. Kahn, John H. Gennari
AMIA2
2001 The Expanding Informatics Community: Blessing or Curse?
abstract
In their introductory comments to the 2001 ACMI Symposium published in this issue, Friedman, Ozbolt, and Masys 1 note that the current fast-moving and turbulent times seem to be both a blessing and curse for biomedical informatics. Never before have so many modifiers appeared next to the term “informatics,” as collaborations spread out into an ever-widening array of professional and consumer medical/clinical/health fields. Their example of “public health informatics” is only one of dozens of examples that they could have chosen, in which an additional level of informatics specialization has occurred in the health sciences. Adding to the sense of expanding roles and opportunities, informatics teams play an essential role in strategic market needs analysis and in product design, development, and deployment within the clinical and biomedical commercial sectors. None of these features were prevalent in the informatics landscape less than 5 years ago. From the published report, it appears that the initial concern of the conference organizers was how to prevent the field of informatics from fractionating into such small subsegments that nothing remains as a core shared culture that binds and unifies all informatics practitioners and investigators. From there, the conference seems to have both embraced the widening scope of informatics and created an overall four-part superstructure in which to place the rapidly expanding disparate pieces together. Thus, although the primary data look scattered, the ACMI conference attendees seemed to find four “clusters” in which to organize the overarching themes for widely disparate biomedical informatics activities. Before commenting on some of the organizing themes that form the major contribution from the conference, I first would like to add a personal perspective on the underlying uneasy sense of “culture lost” or perhaps more appropriate “culture diluted” that motivated the conference initially. The strength of any interdisciplinary activity is also its biggest weakness. Informatics professionals and practitioners bring to the table a highly unique blend of clinical, technical, and methodological skills and assumptions. A blending of the computer scientist, the engineer, the health policy maker, the clinical setting, and the evaluation scientist all went into making the unique beast of the informatician of earlier years. Today we need to add a dash of the molecular biologist, a pinch of the geneticist, a smidgen of the image scientist, and a sprinkle of the cognitive psychologist, and the list keeps growing. I frequently describe my role as analogous to a “United Nations translator,” facilitating communication and understanding between clinician-speak and techno-speak. I often would see first-hand that when such a role was missing, serious misunderstandings and lost collaboration ensued. I believe the proliferation of additional informatics groupings validates the growing recognition of the success and contribution from such an integrative approach. However, I also have experienced directly the less settling side of being “jack of all trades, master of none” in both academic and commercial settings. Being neither the subspecialty domain expert and nor the professional software developer places the informatics professional in the position of being most able to see what needs to be done but least able to actually make it happen without significant contributions from others. Not one of the many projects to which I contributed significantly could have been accomplished without the extensive collaboration of others. It is still difficult to articulate precisely what component in a large project I specifically “owned” rather than “facilitated,” and yet I am certain that my direct participation enabled a set of activities to occur successfully that otherwise would not have happened. Thus, in both academic and commercial settings, it is often difficult to claim the science or claim the engineering (depending on which of the two you are seeking to claim). But none of these issues of professional angst should be unfamiliar or uncomfortable to any professional who seeks to live in the middle of an interdisciplinary field. It is all part of the informatics blessing and curse—the “culture” of informatics that Friedman et al. are concerned may be “dissolved into the cultures of these expanding work settings.” To turn to some of the organizing themes proposed by the conference attendees, here is a brief set of comments: Genome-enabled science and health care. Friedman states: “With a small number of exceptions, the community of scientists traditionally associated with medical informatics has not played a significant role in the work of structural genomics.” I would have constructed this statement differently: “With an extremely small number of exceptions, the community of scientists traditionally associated with structural genomics have not viewed informatics collaborators as more than providing IT infrastructure support so that the ‘real scientists' such as computational biologists can perform the ‘real science’ more effectively.” The Symposium urges “the group historically identified with AMIA and ACMI…to take aggressive steps to build collaborations with the communities of biologists who are extending their own expertise to include computational methods.” I wonder what it would take for both communities to meet halfway and still feel that they each bring unique skills to the collaboration. System-mindedness and error reduction. The Symposium participants call for additional multi-institutional evaluations of the effects of “real-time decision support and long-term learning from clinical data for quality improvement.” Although I do not disagree with the need for additional, well-designed, generalizable studies, it is interesting to wonder aloud why other technologies that have had far less prior research and have shown far less impact have nonetheless been widely adopted into clinical practice. This lack of adoption, despite strong preliminary evidence, also applies to the poor penetration of NIH consensus panel guidelines and protocols into routine clinical practice. Why are technologies with far less supporting data adopted rapidly while these technologies languish? My personal opinion is that systems-oriented technologies result in indirect cost avoidance rather than in direct cost savings or, better yet, new revenue streams. It is impossible to prove “what would have happened but didn't” (prevention/cost avoidance); it is much easier to convince others of “what was happening but isn't anymore” (intervention/cost reduction). Thus, the technologies of decision-support and error reduction are very tough sells to clinical thought leaders, health system board of directors, and corporate funders. Unfortunately, most of what we do is of the former ilk. Moving beyond the guild mentality. Of the four agendas, this one struck me as the boldest and the most naïve simultaneously. This agenda seeks to use informatics-enabled information access to break down the existing “guild mentality that views health professional practitioners as an estate externalized from general society and responsible more to its own norms and values than to society as a whole.” (For a moment, I thought Friedman et al. switched to describing lawyers and politicians.) Breaking down professional identities or “guilds” is a very different agenda from creating and supporting an informed, participatory patient/decision maker. Any guild, including the health professionals, is an extremely effective “barrier to entry” for excluding non-guild members, and is thus is a powerful means for creating a controlled shortage of skills that are “owned” and “certified” by the guild. Reducing these barriers effectively dilutes the exclusivity that guild members have over non-guild members. This potential reduction of the barrier to entry by non-guild members represents a significant economic threat if guild members are valued for the membership in an exclusive “club.” It is extremely difficult to imagine a scenario in which a professional organization would willingly surrender its “right” to determine membership criteria and hence guild exclusivity without stiff (as in, to-the-last-breathing-member) resistance. To do so would effectively surrender members' hard-earned special privileges that come from “paying the price” of becoming a guild member. Can the consumerism and personal activism implied by this agenda overcome this resistance in the health professions? One statistic—that health-related Web sites consistently are the most-often visited URLs—suggests perhaps so. The experience of medscape.com, drkoop.com, and a large number of less-visible defunct or moribund consumer-health related electronic resources suggests otherwise—at least not in a commercially sustainable model. The topic of the 2001 ACMI Symposium is timely. The concerns about “a field…that could…differentiate itself into oblivion” and the need to “reverse the trend to hyper-differentiation and replace it with a renewed integration” are real and are important. By proposing four action agendas, the Symposium participants emphasize some opportunities at the expense of others. This side effect is necessary for achieving focus. All successful enterprises must decide what is “core” and what is “not core” so as to preserve what is essential. What seems to be most missing from the Symposium is a common understanding what makes informatics types so “special” compared with others. What do we as a group bring to the table irrespective of the adjective placed before the word “informatics?” In what way do we “think differently” from our collaborators and how can we make that difference appreciated, acknowledged, and valued quite apart from the domain knowledge in which we chose to apply our expertise? These are questions that both AMIA and ACMI have examined and discussed in the past. It seems that the rapidly growing depth and breadth of our field requires another look at those common themes that define our unique culture and that will endure no matter what changes occur within the field. I congratulate the 2001 ACMI Symposium organizers for attacking this difficult issue “head-on.” Hearing directly from the readers about their thoughts on “What makes informatics researchers and investigators unique?” would be a fascinating exercise. I invite readers to send their responses directly to the JAMIA editor.
Michael G. Kahn
J. Am. Medical Informatics Assoc.1
1997 Using Logical Observation Identifier Names and Codes (LOINC) to exchange laboratory data among three academic hospitals
David M. Baorto, James J. Cimino, Curtis A. Parvin, Michael G. Kahn
AMIA4
1997 Iterative usability testing: ensuring a usable clinical workstation
Janette M. Coble, John Karat, Matthew J. Orland, Michael G. Kahn
AMIA4
1997 Maintaining a Focus on User Requirements Throughout the Development of Clinical Workstation Software
abstract
Article Free Access Share on Maintaining a focus on user requirements throughout the development of clinical workstation software Authors: Janette M. Coble Washington Univ. School of Medicine, Box 8005, 660 S. Euclid, St. Louis, MO Washington Univ. School of Medicine, Box 8005, 660 S. Euclid, St. Louis, MOView Profile , John Karat IBM T. J. Watson Research, 30 Saw Mill River Road, Hawthorne, NY IBM T. J. Watson Research, 30 Saw Mill River Road, Hawthorne, NYView Profile , Michael G. Kahn Washington Univ. School of Medicine, Box 8005, 660 S. Euclid, St. Louis, MO Washington Univ. School of Medicine, Box 8005, 660 S. Euclid, St. Louis, MOView Profile Authors Info & Claims CHI '97: Proceedings of the ACM SIGCHI Conference on Human factors in computing systemsMarch 1997 Pages 170–177https://doi.org/10.1145/258549.258698Published:27 March 1997Publication History 27citation1,256DownloadsMetricsTotal Citations27Total Downloads1,256Last 12 Months44Last 6 weeks9 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Janette M. Coble, John Karat, Michael G. Kahn
CHI3
1996 Review: Statistical Process Control Methods for Expert System Performance Monitoring
abstract
The literature on the performance evaluation of medical expert system is extensive, yet most of the techniques used in the early stages of system development are inappropriate for deployed expert systems. Because extensive clinical and informatics expertise and resources are required to perform evaluations, efficient yet effective methods of monitoring performance during the long-term maintenance phase of the expert system life cycle must be devised. Statistical process control techniques provide a well-established methodology that can be used to define policies and procedures for continuous, concurrent performance evaluation. Although the field of statistical process control has been developed for monitoring industrial processes, its tools, techniques, and theory are easily transferred to the evaluation of expert systems. Statistical process tools provide convenient visual methods and heuristic guidelines for detecting meaningful changes in expert system performance. The underlying statistical theory provides estimates of the detection capabilities of alternative evaluation strategies. This paper describes a set of statistical process control tools that can be used to monitor the performance of a number of deployed medical expert systems. It describes how p-charts are used in practice to monitor the GermWatcher expert system. The case volume and error rate of GermWatcher are then used to demonstrate how different inspection strategies would perform.
Michael G. Kahn, Thomas C. Bailey, Sherry A. Steib, Victoria J. Fraser, W. Claiborne Dunagan
J. Am. Medical Informatics Assoc.1
1996 Research Paper: Monitoring Expert System Performance Using Continuous User Feedback
abstract
OBJECTIVE: To evaluate the applicability of metrics collected during routine use to monitor the performance of a deployed expert system. METHODS: Two extensive formal evaluations of the GermWatcher (Washington University School of Medicine) expert system were performed approximately six months apart. Deficiencies noted during the first evaluation were corrected via a series of interim changes to the expert system rules, even though the expert system was in routine use. As part of their daily work routine, infection control nurses reviewed expert system output and changed the output results with which they disagreed. The rate of nurse disagreement with expert system output was used as an indirect or surrogate metric of expert system performance between formal evaluations. The results of the second evaluation were used to validate the disagreement rate as an indirect performance measure. Based on continued monitoring of user feedback, expert system changes incorporated after the second formal evaluation have resulted in additional improvements in performance. RESULTS: The rate of nurse disagreement with GermWatcher output decreased consistently after each change to the program. The second formal evaluation confirmed a marked improvement in the program's performance, justifying the use of the nurses' disagreement rate as an indirect performance metric. CONCLUSIONS: Metrics collected during the routine use of the GermWatcher expert system can be used to monitor the performance of the expert system. The impact of improvements to the program can be followed using continuous user feedback without requiring extensive formal evaluations after each modification. When possible, the design of an expert system should incorporate measures of system performance that can be collected and monitored during the routine use of the system.
Michael G. Kahn, Sherry A. Steib, W. Claiborne Dunagan, Victoria J. Fraser
J. Am. Medical Informatics Assoc.1
1992 Artificial intelligence in medicine workshop: AAAI 1992 Spring Symposium Series Stanford University March 25-27, 1992 Workshop Summary
Michael G. Kahn
Artif. Intell. Medicine1
1991 The visual display of temporal information
Steve B. Cousins, Michael G. Kahn
Artif. Intell. Medicine2