EDBT 2026 Demo / reviewers in the wild / expert
Albert M. Lai
dblp:39/2229 · also Albert Max Lai
· DBLP profile ↗
38ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0002-9241-2656ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 3Systems, architecture and hardware · 2 · 2 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semi-automated pipeline to accelerate multi-site flowsheet alignment and concept mapping in electronic health recordsabstractOBJECTIVES: Health-care institutions customize electronic health record (EHR) configurations to reflect their unique workflows and patient care priorities. Ensuring EHR alignment across sites facilitates seamless information exchange. We developed a pipeline for EHR flowsheet alignment between health-care organizations. The pipeline is augmented by mapping flowsheet data fields to concepts in the Clinical Care Classification (CCC) nursing terminology. MATERIALS AND METHODS: Flowsheet templates and measures from 2 study sites were transformed into template-measure (T-M) pairs. They were aligned through exact, lexical, or semantic matching. Lexical matches were assessed using Jaccard similarity and fuzzy matching methods. Semantic alignment was determined using cosine similarity between large language model-generated embeddings of T-M pairs and CCC concepts to rank and recommend the top n concepts in CCC. Concept mappings were evaluated based on whether concepts were mapped consistently within the CCC hierarchy. RESULTS: We totally aligned 31 255 unique T-M pairs in acute care units and 27 012 T-M pairs in intensive care units from 2 study sites. When restricted to the top-ranked CCC concept (n = 1), we achieved a 63% flowsheet alignment rate with a 53% concept mapping rate. Expanding to the top 3 concepts (n = 3) improved alignment to 96.5% and concept mapping to 96%. DISCUSSION AND CONCLUSION: Electronic health record data field alignment with concept mapping offers opportunities to standardize data elements presented in flowsheets across health-care sites. We demonstrated the feasibility of leveraging a semi-automated pipeline to streamline the EHR flowsheet alignment and accelerate the manual concept mapping process. Sarah Collins Rossetti, Jennifer Thate, Rosemary Mugoya, Albert M. Lai, Po-Yin Yen |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Comparison of rule- and large language model-based phenotype extraction from clinical notes for neurofibromatosis type 1abstractINTRODUCTION: Neurofibromatosis type 1 (NF1) is a rare genetic disorder affecting multiple organ systems with significant clinical heterogeneity. Managing individuals with NF1 is challenging due to variability in disease progression and outcomes and limited early risk assessment tools. OBJECTIVE: This study aims to develop an effective, generalizable, user-friendly clinical entity extraction pipeline for identifying NF1-related phenotypes from unstructured clinical notes to enhance research and risk-modeling efforts. We compare the benefits of rule-based natural language processing (NLP) vs large language models (LLMs) for this purpose. MATERIALS AND METHODS: Four phenotype extraction pipelines (3 LLM-based vs 1 rule-based) were developed to automatically extract selected NF1-relevant phenotypes. Subject matter experts manually reviewed clinical notes, generating a gold-standard annotation dataset for evaluation. In Phase 1, notes authored by a single NF1 physician were used to guide pipeline development and refinement. In Phase 2, notes from a second NF1 physician were used to assess pipeline generalizability, followed by further refinement to accommodate differences in physician terminology. RESULTS: With refinement, the rule-based model had higher distributions of F1 scores than the LLMs in both Phase 1 and Phase 2. However, the LLMs demonstrated better generalizability between physicians without refinement, showing lesser performance decreases (4.4%-5.1%) when transitioning from Phase 1 to Phase 2 without refinement, compared to an 8.8% decrease for the rule-based model. CONCLUSION: We highlight trade-offs between the effectiveness of rule-based NLP vs generalizability and ease of implementation of LLMs for clinical entity extraction, with implications for pipeline portability across providers and institutions. Levi Kaster, Ethan Hillis, Inez Y. Oh, Elizabeth C. Cordell, Randi E. Foraker, Albert M. Lai, Stephanie M. Morris, David H. Gutmann, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 6 |
| 2025 | CeRTS: certainty retrieval token search in large language model clinical information extractionabstractOBJECTIVE: Large language models (LLMs) must effectively communicate their uncertainty to be viable in clinical settings. As such, the need for reliable uncertainty estimation grows increasingly urgent with the expanding use of LLMs for information extraction from electronic health records. Previous token-level uncertainty estimators have only used token probabilities within a single output sequence. Here, by leveraging the constraints of JSON output structure, we instead consider all likely sequences and their respective probabilities to obtain a more robust measure of model confidence. We develop Certainty Retrieval Token Search (CeRTS), a new uncertainty estimator for structured information extraction. METHODS: We evaluated CeRTS against a previous gold-standard uncertainty estimator when extracting clinical features from lung cancer discharge summaries across eight open-source LLMs. Calibration (Brier score) and discrimination (AUROC) were used to quantify performance. RESULTS: CeRTS surpassed the previous gold-standard estimator in discriminatory power across every model and achieved better calibration in most cases. CeRTS had the strongest agreement between model confidence and accuracy with Qwen-2.5. CONCLUSION: CeRTS enhances LLM-based information extraction from unstructured clinical text by assigning well-calibrated confidence scores to each extracted item, providing medical researchers with a quantitative measure of reliability at minimal additional cost. Although its performance was generally robust, CeRTS struggled with DeepSeek-R1, which we attribute to the model's Chain-of-Thought reasoning steps. Our evaluation focused on clinical data, but CeRTS can be applied to any domain requiring reliable uncertainty estimation. Lars E. Schimmelpfennig, Kriti Bhattarai, Inez Y. Oh, Jake Lever, Obi L. Griffith, Malachi Griffith, Albert M. Lai, Zachary B. Abrams |
J. Biomed. Informatics | 7 |
| 2023 | Electronic health record data quality assessment and tools: a systematic reviewabstractOBJECTIVE: We extended a 2013 literature review on electronic health record (EHR) data quality assessment approaches and tools to determine recent improvements or changes in EHR data quality assessment methodologies. MATERIALS AND METHODS: We completed a systematic review of PubMed articles from 2013 to April 2023 that discussed the quality assessment of EHR data. We screened and reviewed papers for the dimensions and methods defined in the original 2013 manuscript. We categorized papers as data quality outcomes of interest, tools, or opinion pieces. We abstracted and defined additional themes and methods though an iterative review process. RESULTS: We included 103 papers in the review, of which 73 were data quality outcomes of interest papers, 22 were tools, and 8 were opinion pieces. The most common dimension of data quality assessed was completeness, followed by correctness, concordance, plausibility, and currency. We abstracted conformance and bias as 2 additional dimensions of data quality and structural agreement as an additional methodology. DISCUSSION: There has been an increase in EHR data quality assessment publications since the original 2013 review. Consistent dimensions of EHR data quality continue to be assessed across applications. Despite consistent patterns of assessment, there still does not exist a standard approach for assessing EHR data quality. CONCLUSION: Guidelines are needed for EHR data quality assessment to improve the efficiency, transparency, comparability, and interoperability of data quality assessment. These guidelines must be both scalable and flexible. Automation could be helpful in generalizing this process. Abigail E. Lewis, Nicole Gray Weiskopf, Zachary B. Abrams, Randi E. Foraker, Albert M. Lai, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Respiratory support status from EHR data for adult population: classification, heuristics, and usage in predictive modelingabstractOBJECTIVE: Respiratory support status is critical in understanding patient status, but electronic health record data are often scattered, incomplete, and contradictory. Further, there has been limited work on standardizing representations for respiratory support. The objective of this work was to (1) propose a practical terminology system for respiratory support methods; (2) develop (meta-)heuristics for constructing respiratory support episodes; and (3) evaluate the utility of respiratory support information for mortality prediction. MATERIALS AND METHODS: All analyses were performed using electronic health record data of COVID-19-tested, emergency department-admit, adult patients at a large, Midwestern healthcare system between March 1, 2020 and April 1, 2021. Logistic regression and XGBoost models were trained with and without respiratory support information, and performance metrics were compared. Importance of respiratory-support-based features was explored using absolute coefficient values for logistic regression and SHapley Additive exPlanations values for the XGBoost model. RESULTS: The proposed terminology system for respiratory support methods is as follows: Low-Flow Oxygen Therapy (LFOT), High-Flow Oxygen Therapy (HFOT), Non-Invasive Mechanical Ventilation (NIMV), Invasive Mechanical Ventilation (IMV), and ExtraCorporeal Membrane Oxygenation (ECMO). The addition of respiratory support information significantly improved mortality prediction (logistic regression area under receiver operating characteristic curve, median [IQR] from 0.855 [0.852-0.855] to 0.881 [0.876-0.884]; area under precision recall curve from 0.262 [0.245-0.268] to 0.319 [0.313-0.325], both P < 0.01). The proposed generalizable, interpretable, and episodic representation had commensurate performance compared to alternate representations despite loss of granularity. Respiratory support features were among the most important in both models. CONCLUSION: Respiratory support information is critical in understanding patient status and can facilitate downstream analyses. Sean C. Yu, Mackenzie R. Hofford, Albert M. Lai, Marin Kollef, Philip R. O. Payne, Andrew P. Michelson |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | When past is not a prologue: Adapting informatics practice during a pandemicabstractData and information technology are key to every aspect of our response to the current coronavirus disease 2019 (COVID-19) pandemic-including the diagnosis of patients and delivery of care, the development of predictive models of disease spread, and the management of personnel and equipment. The increasing engagement of informaticians at the forefront of these efforts has been a fundamental shift, from an academic to an operational role. However, the past history of informatics as a scientific domain and an area of applied practice provides little guidance or prologue for the incredible challenges that we are now tasked with performing. Building on our recent experiences, we present 4 critical lessons learned that have helped shape our scalable, data-driven response to COVID-19. We describe each of these lessons within the context of specific solutions and strategies we applied in addressing the challenges that we faced. Thomas George Kannampallil, Randi E. Foraker, Albert M. Lai, Keith F. Woeltje, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Automated classification of mobility activities in free text clinical narratives
Denis Newman-Griffis, Ayah Zirikly, Pei-Shu Ho, Jonathan Camacho, Maryanne Sacco, Alex Marr, Albert M. Lai, Eric Fosler-Lussier |
AMIA | 7 |
| 2017 | Novel Visualization of Clostridium Difficile Infection in Intensive Care Units
Sean C. Yu, Courtney Hebert, Albert M. Lai, Justin Smyer, Jennifer Flaherty, Julie Mangino, Susan D. Moffatt-Bruce |
AMIA | 3 |
| 2017 | Inductive identification of functional status information and establishing a gold standard corpus: A case study on the Mobility domainabstractThe importance of functional status information (FSI) has become increasingly evident in recent years [1, 2]. However, implementation, application, and normalization of FSI in health care and Electronic Health Records (EHRs) have been largely underexplored. The World Health Organization's International Classification of Functioning, Disability and Health (ICF) [3] is considered to be the international standard for describing and coding function and health states. Nevertheless, the ICF provides only a limited vocabulary for recognizing FSI descriptions, since its purpose is to organize concepts related to functioning rather than to provide a comprehensive terminology or a complete set of relations between concepts. While the free text portion of EHRs might provide a more complete picture of health status, treatment, and progress, current Natural Language Processing (NLP) methods largely focus on extracting medical conditions (e.g. diagnoses and symptoms, etc.). The absence of a standardized functional terminology and incompleteness of the ICF as a vocabulary source makes it challenging to build a NLP system to extract FSI from EHR free text. Our work takes the first step towards extraction of FSI from free text by systematically identifying the structure of FSI related to Mobility, a key domain of the ICF and an important domain in the determination of work disability. Our interdisciplinary research group inductively evaluated examples extracted from over 1,200 Physical Therapy (PT) notes from the Clinical Center of the National Institutes of Health (NIH). This extensive work resulted in a nested entity structure comprised of 2 entities, 3 sub-entities, 8 attributes, and 21 attribute values. Furthermore, we have manually curated the first gold standard corpus of 200 double-annotated and 50 triple-annotated PT notes. Our inter-annotator agreement (IAA) averages 97% F1-score on partial textual span matching and from 0.4 to 0.9 Siegel & Castellan's kappa on attribute value matching. Such a rich semantic corpus of Mobility FSI is valuable and a promising resource for future statistical learning. Our method is also adaptable to other domains of the ICF. Thanh Thieu, Jonathan Camacho, Pei-Shu Ho, Julia Porcino, Lisa Nelson, Elizabeth Rasch, Chunxiao Zhou, Leighton Chan, Diane Brandt, Denis Newman-Griffis, Ao Yuan, Albert M. Lai |
BIBM | 13 |
| 2016 | Parsing complex microbiology data for secondary use
Protiva Rahman, Courtney Hebert, Albert M. Lai |
AMIA | 3 |
| 2016 | Automatic data source identification for clinical trial eligibility criteria resolution
Chaitanya P. Shivade, Courtney Hebert, Kelly Regan-Fendt, Eric Fosler-Lussier, Albert M. Lai |
AMIA | 5 |
| 2016 | A Review of Mobile Phone-based Interventions and Applications for Medication Adherence
Po-Yin Yen, Jessica Garvey Smith, Michelle P. Zhou, Megan Chamberlain, Xiaonan Ji, Albert M. Lai |
AMIA | 6 |
| 2016 | The geographic distribution of cardiovascular health in the stroke prevention in healthcare delivery environments (SPHERE) study
Caryn Roth, Philip R. O. Payne, Rory C. Weier, Abigail B. Shoben, Erica N. Fletcher, Albert M. Lai, Marjorie M. Kelley, Jesse J. Plascak, Randi E. Foraker |
J. Biomed. Informatics | 6 |
| 2015 | RECRUIT: Roadmap to Enhance Clinical trial Recruitment Using Information Technology
Tasneem Motiwala, Chaitanya P. Shivade, Satyajeet Raje, Albert M. Lai, Philip R. O. Payne |
AMIA | 4 |
| 2015 | Corpus-based discovery of semantic intensity scalesabstractChaitanya Shivade, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Chaitanya P. Shivade, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai |
HLT-NAACL | 4 |
| 2015 | Textual inference for eligibility criteria resolution in clinical trialsabstractClinical trials are essential for determining whether new interventions are effective. In order to determine the eligibility of patients to enroll into these trials, clinical trial coordinators often perform a manual review of clinical notes in the electronic health record of patients. This is a very time-consuming and exhausting task. Efforts in this process can be expedited if these coordinators are directed toward specific parts of the text that are relevant for eligibility determination. In this study, we describe the creation of a dataset that can be used to evaluate automated methods capable of identifying sentences in a note that are relevant for screening a patient's eligibility in clinical trials. Using this dataset, we also present results for four simple methods in natural language processing that can be used to automate this task. We found that this is a challenging task (maximum F-score=26.25), but it is a promising direction for further research. Chaitanya P. Shivade, Courtney Hebert, Marcelo A. Lopetegui, Marie-Catherine de Marneffe, Eric Fosler-Lussier, Albert M. Lai |
J. Biomed. Informatics | 6 |
| 2015 | Comparison of UMLS terminologies to identify risk of heart disease using clinical notesabstractThe second track of the 2014 i2b2 challenge asked participants to automatically identify risk factors for heart disease among diabetic patients using natural language processing techniques for clinical notes. This paper describes a rule-based system developed using a combination of regular expressions, concepts from the Unified Medical Language System (UMLS), and freely-available resources from the community. With a performance (F1=90.7) that is significantly higher than the median (F1=87.20) and close to the top performing system (F1=92.8), it was the best rule-based system of all the submissions in the challenge. We also used this system to evaluate the utility of different terminologies in the UMLS towards the challenge task. Of the 155 terminologies in the UMLS, 129 (76.78%) have no representation in the corpus. The Consumer Health Vocabulary had very good coverage of relevant concepts and was the most useful terminology for the challenge task. While segmenting notes into sections and lists has a significant impact on the performance, identifying negations and experiencer of the medical event results in negligible gain. Chaitanya P. Shivade, Pranav Malewadkar, Eric Fosler-Lussier, Albert M. Lai |
J. Biomed. Informatics | 4 |
| 2014 | Cross-narrative Temporal Ordering of Medical EventsabstractCross-narrative temporal ordering of medical events is essential to the task of generating a comprehensive timeline over a patient's history.We address the problem of aligning multiple medical event sequences, corresponding to different clinical narratives, comparing the following approaches: (1) A novel weighted finite state transducer representation of medical event sequences that enables composition and search for decoding, and (2) Dynamic programming with iterative pairwise alignment of multiple sequences using global and local alignment algorithms.The cross-narrative coreference and temporal relation weights used in both these approaches are learned from a corpus of clinical narratives.We present results using both approaches and observe that the finite state transducer approach performs performs significantly better than the dynamic programming one by 6.8% for the problem of multiple-sequence alignment. Preethi Raghavan, Eric Fosler-Lussier, Noémie Elhadad, Albert M. Lai |
ACL (1) | 4 |
| 2014 | A Literature Review of Electronic Health Record Redesign for Optimization
Alissa A. Schultz, Albert M. Lai, Po-Yin Yen |
AMIA | 2 |
| 2014 | A review of approaches to identifying patient phenotype cohorts using electronic health recordsabstractOBJECTIVE: To summarize literature describing approaches aimed at automatically identifying patients with a common phenotype. MATERIALS AND METHODS: We performed a review of studies describing systems or reporting techniques developed for identifying cohorts of patients with specific phenotypes. Every full text article published in (1) Journal of American Medical Informatics Association, (2) Journal of Biomedical Informatics, (3) Proceedings of the Annual American Medical Informatics Association Symposium, and (4) Proceedings of Clinical Research Informatics Conference within the past 3 years was assessed for inclusion in the review. Only articles using automated techniques were included. RESULTS: Ninety-seven articles met our inclusion criteria. Forty-six used natural language processing (NLP)-based techniques, 24 described rule-based systems, 41 used statistical analyses, data mining, or machine learning techniques, while 22 described hybrid systems. Nine articles described the architecture of large-scale systems developed for determining cohort eligibility of patients. DISCUSSION: We observe that there is a rise in the number of studies associated with cohort identification using electronic medical records. Statistical analyses or machine learning, followed by NLP techniques, are gaining popularity over the years in comparison with rule-based systems. CONCLUSIONS: There are a variety of approaches for classifying patients into a particular phenotype. Different techniques and data sources are used, and good performance is reported on datasets at respective institutions. However, no system makes comprehensive use of electronic medical records addressing all of their known weaknesses. Chaitanya P. Shivade, Preethi Raghavan, Eric Fosler-Lussier, Peter J. Embí, Noémie Elhadad, Stephen B. Johnson, Albert M. Lai |
J. Am. Medical Informatics Assoc. | 7 |
| 2014 | Time motion studies in healthcare: What are we talking about?abstractTime motion studies were first described in the early 20th century in industrial engineering, referring to a quantitative data collection method where an external observer captured detailed data on the duration and movements required to accomplish a specific task, coupled with an analysis focused on improving efficiency. Since then, they have been broadly adopted by biomedical researchers and have become a focus of attention due to the current interest in clinical workflow related factors. However, attempts to aggregate results from these studies have been difficult, resulting from a significant variability in the implementation and reporting of methods. While efforts have been made to standardize the reporting of such data and findings, a lack of common understanding on what "time motion studies" are remains, which not only hinders reviews, but could also partially explain the methodological variability in the domain literature (duration of the observations, number of tasks, multitasking, training rigor and reliability assessments) caused by an attempt to cluster dissimilar sub-techniques. A crucial milestone towards the standardization and validation of time motion studies corresponds to a common understanding, accompanied by a proper recognition of the distinct techniques it encompasses. Towards this goal, we conducted a review of the literature aiming at identifying what is being referred to as "time motion studies". We provide a detailed description of the distinct methods used in articles referenced or classified as "time motion studies", and conclude that currently it is used not only to define the original technique, but also to describe a broad spectrum of studies whose only common factor is the capture and/or analysis of the duration of one or more events. To maintain alignment with the existing broad scope of the term, we propose a disambiguation approach by preserving the expanded conception, while recommending the use of a specific qualifier "continuous observation time motion studies" to refer to variations of the original method (the use of an external observer recording data continuously). In addition, we present a more granular naming for sub-techniques within continuous observation time motion studies, expecting to reduce the methodological variability within each sub-technique and facilitate future results aggregation. Marcelo A. Lopetegui, Po-Yin Yen, Albert M. Lai, Joseph Jeffries, Peter J. Embí, Philip R. O. Payne |
J. Biomed. Informatics | 3 |
| 2013 | Inter-Observer Reliability Assessments in Time Motion Studies: The Foundation for Meaningful Clinical Workflow Analysis
Marcelo A. Lopetegui, Shasha Bai, Po-Yin Yen, Albert M. Lai, Peter J. Embí, Philip R. O. Payne |
AMIA | 4 |
| 2012 | Time Capture Tool (TimeCaT): Development of a Comprehensive Application to Support Data Capture for Time Motion Studies
Marcelo A. Lopetegui, Po-Yin Yen, Albert M. Lai, Peter J. Embí, Philip R. O. Payne |
AMIA | 3 |
| 2012 | Inter-Annotator Reliability of Medical Events, Coreferences and Temporal Relations in Clinical Narratives by Annotators with Varying Levels of Clinical Expertise
Preethi Raghavan, Eric Fosler-Lussier, Albert M. Lai |
AMIA | 3 |
| 2012 | Exploring Semi-Supervised Coreference Resolution of Medical Concepts using Semantic and Temporal Features
Preethi Raghavan, Eric Fosler-Lussier, Albert M. Lai |
HLT-NAACL | 3 |
| 2012 | Applying knowledge-anchored hypothesis discovery methods to advance clinical and translational research: the OAMiner projectabstractThe conduct of clinical and translational research regularly involves the use of a variety of heterogeneous and large-scale data resources. Scalable methods for the integrative analysis of such resources, particularly when attempting to leverage computable domain knowledge in order to generate actionable hypotheses in a high-throughput manner, remain an open area of research. In this report, we describe both a generalizable design pattern for such integrative knowledge-anchored hypothesis discovery operations and our experience in applying that design pattern in the experimental context of a set of driving research questions related to the publicly available Osteoarthritis Initiative data repository. We believe that this 'test bed' project and the lessons learned during its execution are both generalizable and representative of common clinical and translational research paradigms. Philip R. O. Payne, Rebecca D. Jackson, Thomas M. Best, Tara Borlawsky, Albert M. Lai, Stephen L. James, Metin Nafi Gürcan |
J. Am. Medical Informatics Assoc. | 5 |
| 2011 | Utility-directed resource allocation in virtual desktop clouds
Prasad Calyam, Rohit Patali, Alex Berryman, Albert M. Lai, Rajiv Ramnath |
Comput. Networks | 4 |
| 2010 | VDBench: A Benchmarking Toolkit for Thin-Client Based Virtual Desktop EnvironmentsabstractThe recent advances in thin client devices and the push to transition users' desktop delivery to cloud environments will eventually transform how desktop computers are used today. The ability to measure and adapt the performance of virtual desktop environments is a major challenge for ''virtual desktop cloud'' service providers. In this paper, we present the ''VD Bench'' toolkit that uses a novel methodology and related metrics to benchmark thin-client based virtual desktop environments in terms of scalability and reliability. We also describe how we used a VD Bench instance to benchmark the performance of: (a) popular user applications (Spreadsheet Calculator, Internet Browser, Media Player, Interactive Visualization), (b) TCP/UDP based thin client protocols (RDP, RGS, PCoIP), and (c) remote user experience (interactive response times, perceived video quality), under a variety of system load and network health conditions. Our results can help service providers to mitigate over-provisioning in sizing virtual desktop resources, and guesswork in thin client protocol configurations, and thus obtain significant cost savings while simultaneously fostering satisfied customers. Alex Berryman, Prasad Calyam, Matthew Honigford, Albert M. Lai |
CloudCom | 4 |
| 2009 | Research Paper: Syndromic Surveillance Using Ambulatory Electronic Health RecordsabstractOBJECTIVE: To assess the performance of electronic health record data for syndromic surveillance and to assess the feasibility of broadly distributed surveillance. DESIGN: Two systems were developed to identify influenza-like illness and gastrointestinal infectious disease in ambulatory electronic health record data from a network of community health centers. The first system used queries on structured data and was designed for this specific electronic health record. The second used natural language processing of narrative data, but its queries were developed independently from this health record. Both were compared to influenza isolates and to a verified emergency department chief complaint surveillance system. MEASUREMENTS: Lagged cross-correlation and graphs of the three time series. RESULTS: For influenza-like illness, both the structured and narrative data correlated well with the influenza isolates and with the emergency department data, achieving cross-correlations of 0.89 (structured) and 0.84 (narrative) for isolates and 0.93 and 0.89 for emergency department data, and having similar peaks during influenza season. For gastrointestinal infectious disease, the structured data correlated fairly well with the emergency department data (0.81) with a similar peak, but the narrative data correlated less well (0.47). CONCLUSIONS: It is feasible to use electronic health records for syndromic surveillance. The structured data performed best but required knowledge engineering to match the health record data to the queries. The narrative data illustrated the potential performance of a broadly disseminated system and achieved mixed results. George Hripcsak, Nicholas D. Soulakis, Li Li 0062, Frances P. Morrison, Albert M. Lai, Carol Friedman, Neil S. Calman, Farzad Mostashari |
J. Am. Medical Informatics Assoc. | 5 |
| 2009 | Viewpoint Paper: Repurposing the Clinical Record: Can an Existing Natural Language Processing System De-identify Clinical Notes?abstractElectronic clinical documentation can be useful for activities such as public health surveillance, quality improvement, and research, but existing methods of de-identification may not provide sufficient protection of patient data. The general-purpose natural language processor MedLEE retains medical concepts while excluding the remaining text so, in addition to processing text into structured data, it may be able provide a secondary benefit of de-identification. Without modifying the system, the authors tested the ability of MedLEE to remove protected health information (PHI) by comparing 100 outpatient clinical notes with the corresponding XML-tagged output. Of 809 instances of PHI, 26 (3.2%) were detected in output as a result of processing and identification errors. However, PHI in the output was highly transformed, much appearing as normalized terms for medical concepts, potentially making re-identification more difficult. The MedLEE processor may be a good enhancement to other de-identification systems, both removing PHI and providing coded data from clinical text. Frances P. Morrison, Li Li 0062, Albert M. Lai, George Hripcsak |
J. Am. Medical Informatics Assoc. | 3 |
| 2009 | Research Paper: A Randomized Trial Comparing Telemedicine Case Management with Usual Care in Older, Ethnically Diverse, Medically Underserved Patients with Diabetes Mellitus: 5 Year Results of the IDEATel StudyabstractCONTEXT Telemedicine is a promising but largely unproven technology for providing case management services to patients with chronic conditions and lower access to care. OBJECTIVES To examine the effectiveness of a telemedicine intervention to achieve clinical management goals in older, ethnically diverse, medically underserved patients with diabetes. DESIGN, Setting, and Patients A randomized controlled trial was conducted, comparing telemedicine case management to usual care, with blinded outcome evaluation, in 1,665 Medicare recipients with diabetes, aged >/= 55 years, residing in federally designated medically underserved areas of New York State. Interventions Home telemedicine unit with nurse case management versus usual care. Main Outcome Measures The primary endpoints assessed over 5 years of follow-up were hemoglobin A1c (HgbA1c), low density lipoprotein (LDL) cholesterol, and blood pressure levels. RESULTS Intention-to-treat mixed models showed that telemedicine achieved net overall reductions over five years of follow-up in the primary endpoints (HgbA1c, p = 0.001; LDL, p < 0.001; systolic and diastolic blood pressure, p = 0.024; p < 0.001). Estimated differences (95% CI) in year 5 were 0.29 (0.12, 0.46)% for HgbA1c, 3.84 (-0.08, 7.77) mg/dL for LDL cholesterol, and 4.32 (1.93, 6.72) mm Hg for systolic and 2.64 (1.53, 3.74) mm Hg for diastolic blood pressure. There were 176 deaths in the intervention group and 169 in the usual care group (hazard ratio 1.01 [0.82, 1.24]). CONCLUSIONS Telemedicine case management resulted in net improvements in HgbA1c, LDL-cholesterol and blood pressure levels over 5 years in medically underserved Medicare beneficiaries. Mortality was not different between the groups, although power was limited. Trial Registration http://clinicaltrials.gov Identifier: NCT00271739. Steven Shea, Ruth S. Weinstock, Jeanne A. Teresi, Walter Palmas, Justin Starren, James J. Cimino, Albert M. Lai, Lesley Field, Philip C. Morin, Robin Goland, Roberto E. Izquierdo, Susana Ebner, Stephanie Silver, Eva Petkova, Joseph P. Eimicke |
J. Am. Medical Informatics Assoc. | 7 |
| 2008 | Fuzzy Temporal Constraint Networks for Clinical Information
Albert M. Lai, Simon Parsons, George Hripcsak |
AMIA | 1 |
| 2006 | Training Digital Divide Seniors to use a Telehealth System: A Remote Training Approach
Albert M. Lai, David R. Kaufman, Justin Starren |
AMIA | 1 |
| 2006 | On the performance of wide-area thin-client computingabstractWhile many application service providers have proposed using thin-client computing to deliver computational services over the Internet, little work has been done to evaluate the effectiveness of thin-client computing in a wide-area network. To assess the potential of thin-client computing in the context of future commodity high-bandwidth Internet access, we have used a novel, noninvasive slow-motion benchmarking technique to evaluate the performance of several popular thin-client computing platforms in delivering computational services cross-country over Internet2. Our results show that using thin-client computing in a wide-area network environment can deliver acceptable performance over Internet2, even when client and server are located thousands of miles apart on opposite ends of the country. However, performance varies widely among thin-client platforms and not all platforms are suitable for this environment. While many thin-client systems are touted as being bandwidth efficient, we show that network latency is often the key factor in limiting wide-area thin-client performance. Furthermore, we show that the same techniques used to improve bandwidth efficiency often result in worse overall performance in wide-area networks. We characterize and analyze the different design choices in the various thin-client platforms and explain which of these choices should be selected for supporting wide-area computing services. Albert M. Lai, Jason Nieh |
ACM Trans. Comput. Syst. | 1 |
| 2005 | Architecture for Remote Training of Home Telemedicine Patients
Albert M. Lai, Justin Starren, Steven Shea |
AMIA | 1 |
| 2004 | Improving web browsing performance on wireless pdas using thin-client computingabstractWeb applications are becoming increasingly popular for mobile wireless PDAs. However, web browsing on these systems can be quite slow. An alternative approach is handheld thin-client computing, in which the web browser and associated application logic run on a server, which then sends simple screen updates to thePDA for display. To assess the viability of this thin-client approach, we compare the web browsing performance of thin clients against fat clients that run the web browser locally on a PDA. Our results show that thin clients can provide better web browsing performance compared to fat clients, both in terms of speed and ability to correctly display web content. Surprisingly, thin clients are faster even when having to send more data over the network. We characterize and analyze different design choices in various thin-client systems and explain why these approaches can yield superior web browsing performance on mobile wireless PDAs. Albert M. Lai, Jason Nieh, Bhagyashree Bohra, Vijayarka Nandikonda, Abhishek P. Surana, Suchita Varshneya |
WWW | 1 |
| 2003 | Thin Client Performance for Remote 3-D Image Display
Albert M. Lai, Jason Nieh, Andrew F. Laine, Justin Starren |
AMIA | 1 |
| 2002 | Limits of wide-area thin-client computingabstractWhile many application service providers have proposed using thin-client computing to deliver computational services over the Internet, little work has been done to evaluate the effectiveness of thin-client computing in a wide-area network. To assess the potential of thin-client computing in the context of future commodity high-bandwidth Internet access, we have used a novel, non-invasive slow-motion benchmarking technique to evaluate the performance of several popular thin-client computing platforms in delivering computational services cross-country over Internet2. Our results show that using thin-client computing in a wide-area network environment can deliver acceptable performance over Internet2, even when client and server are located thousands of miles apart on opposite ends of the country. However, performance varies widely among thin-client platforms and not all platforms are suitable for this environment. While many thin-client systems are touted as being bandwidth efficient, we show that network latency is often the key factor in limiting wide-area thin-client performance. Furthermore, we show that the same techniques used to improve bandwidth efficiency often result in worse overall performance in wide-area networks. We characterize and analyze the different design choices in the various thin-client platforms and explain which of these choices should be selected for supporting wide-area computing services. Albert M. Lai, Jason Nieh |
SIGMETRICS | 1 |