VLDB 2026 Research / reviewers in the wild / expert
Bradley A. Malin
dblp:m/BradleyMalin
· DBLP profile ↗
163ranked-venue papers
16as first author
46since 2021 · last 2026
0000-0003-3040-5175ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 113 · 11 first-author · 35 since 2021Databases, data management, data science and information retrieval · 25 · 3 first-author · 2 since 2021Security and privacy · 19 · 6 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Auditor models to suppress poor artificial intelligence predictions can improve human-artificial intelligence collaborative performanceabstractOBJECTIVE: Healthcare decisions are increasingly made with the assistance of machine learning (ML). ML has been known to have unfairness-inconsistent outcomes across subpopulations. Clinicians interacting with these systems can perpetuate such unfairness by overreliance. Recent work exploring ML suppression-silencing predictions based on auditing the ML-shows promise in mitigating performance issues originating from overreliance. This study aims to evaluate the impact of suppression on collaboration fairness and evaluate ML uncertainty as desiderata to audit the ML. MATERIALS AND METHODS: We used data from the Vanderbilt University Medical Center electronic health record (n = 58 817) and the MIMIC-IV-ED dataset (n = 363 145) to predict likelihood of death or intensive care unit transfer and likelihood of 30-day readmission using gradient-boosted trees and an artificially high-performing oracle model. We derived clinician decisions directly from the dataset and simulated clinician acceptance of ML predictions based on previous empirical work on acceptance of clinical decision support alerts. We measured performance as area under the receiver operating characteristic curve and algorithmic fairness using absolute averaged odds difference. RESULTS: When the ML outperforms humans, suppression outperforms the human alone (P < 8.2 × 10-6) and at least does not degrade fairness. When the human outperforms the ML, the human is either fairer than suppression (P < 8.2 × 10-4) or there is no statistically significant difference in fairness. Incorporating uncertainty quantification into suppression approaches can improve performance. CONCLUSION: Suppression of poor-quality ML predictions through an auditor model shows promise in improving collaborative human-AI performance and fairness. Katherine E. Brown, Jesse O. Wrenn, Nicholas J. Jackson, Michael R. Cauley, Benjamin X. Collins, Laurie L. Novak, Bradley A. Malin, Jessica S. Ancker |
J. Am. Medical Informatics Assoc. | 7 |
| 2026 | Re-identification risk for common privacy preserving patient matching strategies when shared with de-identified demographicsabstractOBJECTIVE: Privacy preserving record linkage (PPRL) refers to techniques used to identify which records refer to the same person across disparate datasets while safeguarding their identities. PPRL is increasingly relied upon to facilitate biomedical research. A common strategy encodes personally identifying information for comparison without disclosing underlying identifiers. As the scale of research datasets expands, it becomes crucial to reassess the privacy risks associated with these encodings. This paper highlights the potential re-identification risks of some of these encodings, demonstrating an attack that exploits encoding repetition across patients. MATERIALS AND METHODS: The attack leverages repeated PPRL encoding values combined with common demographics shared during PPRL in the clear (e.g., 3-digit ZIP code) to distinguish encodings from one another and ultimately link them to identities in a reference dataset. Using US Census statistics and voter registries, we empirically estimate encodings' re-identification risk against such an attack, while varying multiple factors that influence the risk. RESULTS: Re-identification risk for PPRL encodings increases with population size, number of distinct encodings per patient, and amount of demographic information available. Commonly used encodings typically grow from <1% re-identification rate for datasets under one million individuals to 10%-20% for 250 million individuals. DISCUSSION AND CONCLUSION: Re-identification risk often remains low in smaller populations, but increases significantly at the larger scales increasingly encountered today. These risks are common in many PPRL implementations, although, as our work shows, they are avoidable. Choosing better tokens or matching tokens through a third party without the underlying demographics effectively eliminates these risks. Austin Eliazar, J. Thomas Brown, Sara Cinamon, Murat Kantarcioglu, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2026 | A novel analysis methodology for assessment of re-identification risks for the National Cancer Institute cancer registry privacy preserving record linkage techniqueabstractOBJECTIVE: The National Cancer Institute (NCI), part of the National Institutes of Health (NIH) supports efforts to address critical challenges in advancing cancer research. As part of this effort, NCI sponsored the development of a privacy-preserving record linkage (PPRL) software that transforms identifying patient information into multiple tokens through a set of cryptographically secure keyed hash functions. This project aims to evaluate the PPRL software in the perspective of re-identification risks and propose effective strategies to sufficiently mitigate these risks. MATERIALS AND METHODS: To achieve the goals, we developed a novel re-identification risk assessment framework, based on token frequency analysis, to estimate the privacy impact of hashed tokens shared for record linkage. We assessed privacy risk through empirical analysis on a state-level voter registration database, a public dataset commonly used for re-identification, under various scenarios. These scenarios are defined based on several factors, including the size of the dataset used for linkage and a group size parameter that determines when an adversary can claim that a record has been re-identified. RESULTS: We found that the re-identification risk based on frequency analysis attack is approximately 0.0002 (ie, 2 patients out of 10 000 are potentially identifiable) under reasonable adversarial settings, with a group size parameter of k = 12 and a dataset size of 400 000 patients. Additionally, our analysis reveals a negative correlation between dataset size and re-identification risk. DISCUSSION: Re-identification risk is deemed low for the new NCI PPRL software. Token frequency analysis provides a reliable estimate of the re-identification risk in token-based PPRL tools. Murat Kantarcioglu, Will Howe, Benmei Liu, Valentina Petkov, Esmeralda Casas-Silva, Diana Velasquez-Kolnik, Bradley A. Malin, Lynne Penberthy |
J. Am. Medical Informatics Assoc. | 7 |
| 2026 | Community medical centers struggle to produce well-calibrated clinical prediction models: Data augmentation can helpabstractOBJECTIVE: Machine learning models (ML) often require localization to perform optimally in local populations. We hypothesize that smaller community healthcare centers may not have the necessary patient volume to facilitate localization based on statistical guidelines. This work investigates the ability for community medical centers to localize ML and performs a simulation study to evaluate synthetic data generation (SDG) to augment local data for recalibration. METHODS: We conducted an experiment using data from a real network of hospitals (two rural, one urban academic medical center) to predict 30-day unplanned hospital readmission and using data from a multi-site ICU dataset to simulate using synthetic data generation (SDG) in a network of hospitals of various sizes. We also performed a simulation study using data from a multi-site ICU dataset to evaluate the utility of SDG to augment local data volumes. RESULTS: In the real-world evaluation, the urban medical center met the guidelines for the number of samples for recalibration (Required: 14,224, Available: 42,303) and had the best calibrated model using local data (α=0.1,β=1.05; best: α=0,β=1). For the smaller sites, neither site had the samples required for recalibration (Site 1: Required: 16461, Available: 3187; Site 2: Required: 15299, Available: 905). In the simulation study, deep learning-based SDG was most effective at improving calibration performance. CONCLUSIONS: Connections to large medical centers are not enough to promote accurate ML at all sites within a healthcare system. Data augmentation and SDG may provide the necessary data volumes to enable local recalibration at smaller facilities. Katherine E. Brown, Bradley A. Malin, Sharon E. Davis |
J. Biomed. Informatics | 2 |
| 2025 | SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt OptimizationabstractWendi Cui, Jiaxin Zhang, Zhuohang Li, Hao Sun, Damien Lopez, Kamalika Das, Bradley A. Malin, Sricharan Kumar. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wendi Cui, Jiaxin Zhang 0005, Damien Lopez, Kamalika Das, Bradley A. Malin, Kumar Sricharan |
ACL (1) | 7 |
| 2025 | Using Large Language Model for Efficient Extraction of Treatment Discontinuation Information - A Study of Online Breast Cancer Community Posts
Qingyuan Song, Jessie Yang, Ndidiamaka Obi, Congning Ni, Jeremy L. Warner, Qingxia Chen, S. Trent Rosenbloom, Bradley A. Malin, Zhijun Yin |
AIME (2) | 9 |
| 2025 | From Voice to Diagnosis: A Hybrid Approach to Multi-Label Disease Classification with Uncertainty AwarenessabstractThe human voice encodes a wealth of acoustic biomarkers linked to various health conditions, including but not limited to neurological, mood, respiratory, and laryngeal disorders. Recent advancements in artificial intelligence (AI) offer a promising solution for leveraging voice to perform noninvasive, cost-effective, and scalable screening of health-related conditions. However, classifying comorbid clinical conditions presents a significant multi-label classification challenge, which is further complicated by class imbalance, feature noise, and the “black-box” nature of modern deep learning models, a critical barrier to model interpretability, error analysis, and potential clinical translation. To address these challenges, we introduce and evaluate a two-stage framework for voice-based, multi-label disease classification on a newly curated dataset from the NIH Bridge2AI initiative. This hybrid approach fuses theory-driven, handcrafted acoustic features with data-driven deep features from a pre-trained ResNet-18. To identify the most effective model configuration, we enhance a Feed-Forward Neural Network (FFNN) with an attention mechanism and Focal Loss, systematically comparing it against traditional classifiers and end-to-end fine-tuning benchmarks. Finally, we incorporate an uncertainty quantification (UQ) layer using Monte Carlo (MC) Dropout and Deep Ensembles to assess prediction reliability. Our results show that the proposed two-stage FFNN with Focal Loss and Attention emerges as the best-performing model in our evaluation, delivering the best balance of high discriminative power (Macro AUC of$\mathbf{0. 8 1 0}$) and robust classification accuracy (Macro F1 of 0.610). While an end-to-end model achieved the highest Macro F1 score (0.638), our two-stage approach proved to be more reliable and computationally efficient. Furthermore, our UQ analysis identified MC Dropout with Predictive Entropy as a practical and effective method, revealing a statistically significant correlation between model uncertainty and prediction error (overall$\mathbf{r} \boldsymbol{=} \mathbf{0. 2 8 0}$,$\mathbf{p}<\mathbf{0. 0 0 1}$), with the strongest link for Voice Disorders ($\mathbf{r} \boldsymbol{=} \mathbf{0. 3 8 0}$). In summary, this study presents a robust technical blueprint for a voice-based AI system, establishing a strong performance benchmark to guide future model development using the Bridge2AI voice dataset. Weixin Liu 0001, Bowen Qu, Matthew E. Pontell, Maria E. Powell, Bradley A. Malin, Zhijun Yin |
BIBM | 5 |
| 2025 | Towards Statistical Factuality Guarantee for Large Vision-Language ModelsabstractAdvancements in Large Vision-Language Models (LVLMs) have demonstrated impressive performance in image-conditioned text generation; however, hallucinated outputs-text that misaligns with the visual input-pose a major barrier to their use in safety-critical applications.We introduce CONFLVLM, a conformal-prediction-based framework that achieves finite-sample distribution-free statistical guarantees to the factuality of LVLM output.Taking each generated detail as a hypothesis, CONFLVLM statistically tests factuality via efficient heuristic uncertainty measures to filter out unreliable claims.We conduct extensive experiments covering three representative application domains: general scene understanding, medical radiology report generation, and document understanding.Remarkably, CON-FLVLM reduces the error rate of claims generated by LLaVa-1.5 for scene descriptions from 87.8% to 10.0% by filtering out erroneous claims with a 95.3% true positive rate.Our results further show that CONFLVLM is highly flexible, and can be applied to any black-box LVLMs paired with any uncertainty measure for any image-conditioned free-form text generation task while providing a rigorous guarantee on controlling hallucination risk. Chao Yan 0004, Nicholas J. Jackson, Wendi Cui, Bo Li 0026, Jiaxin Zhang 0005, Bradley A. Malin |
EMNLP | 7 |
| 2025 | Differential Confounding Privacy and Inverse CompositionabstractDifferential privacy (DP) has become the gold standard for privacy-preserving data analysis, but its applicability can be limited in scenarios involving complex dependencies between sensitive information and datasets. To address this, we introduce differential confounding privacy (DCP), a specialized form of the Pufferfish privacy (PP) framework that generalizes DP by accounting for broader relationships between sensitive information and datasets. DCP adopts the$(\epsilon, \delta)$-indistinguishability framework to quantify privacy loss. We show that while DCP mechanisms retain privacy guarantees under composition, they lack the graceful compositional properties of DP. To overcome this, we propose an Inverse Composition (IC) framework, where a leader-follower model optimally designs a privacy strategy to achieve target guarantees without relying on worst-case privacy proofs, such as sensitivity calculation. Experimental results validate IC's effectiveness in managing privacy budgets and ensuring rigorous privacy guarantees under composition. Tao Zhang 0011, Bradley A. Malin, Netanel Raviv, Yevgeniy Vorobeychik |
ISIT | 2 |
| 2025 | What Really is a Member? Discrediting Membership Inference via PoisoningabstractMembership inference tests aim to determine whether a particular data point was included in a language model's training set. However, recent works have shown that such tests often fail under the strict definition of membership based on exact matching, and have suggested relaxing this definition to include semantic neighbors as members as well. In this work, we show that membership inference tests are still *unreliable* under this relaxation - it is possible to poison the training dataset in a way that causes the test to produce incorrect predictions for a target point. We theoretically reveal a trade-off between a test’s accuracy and its robustness to poisoning. We also present a concrete instantiation of this poisoning attack and empirically validate its effectiveness. Our results show that it can degrade the performance of existing tests to well below random. Neal Mangaokar, Ashish Hooda, Bradley A. Malin, Kassem Fawaz, Somesh Jha, Atul Prakash 0001, Amrita Roy Chowdhury 0001 |
NeurIPS | 4 |
| 2025 | Catalysts of Conversation: Examining Interaction Dynamics Between Topic Initiators and Commentors in Alzheimer's Disease Online CommunitiesabstractInformal caregivers (e.g., family members or friends) of people living with Alzheimer's Disease and Related Dementias (ADRD) face substantial challenges and often seek support through online communities. Understanding the factors driving engagement within these platforms is crucial, as it can enhance communities' long-term value to meet their needs effectively. This study investigated the user interaction dynamics within two large, popular ADRD communities, TalkingPoint and ALZConnected, focusing on topic initiator engagement, initial post content, and the linguistic patterns of comments at the thread level. Using analytical methods such as propensity score matching, topic modeling, and predictive modeling, we found that active topic initiator engagement drives a higher comment volume, and reciprocal replies from topic initiators encourage further commentor engagement at the community level. Practical caregiving topics prompt more re-engagement of topic initiators, while emotional support topics attract more comments from commentors. Additionally, the linguistic complexity and emotional tone of a comment are associated with its likelihood of receiving replies from topic initiators. These findings highlight the importance of fostering active and reciprocal engagement and providing effective strategies to enhance sustainability in ADRD caregiving and broader health-related online communities. Congning Ni, Qingxia Chen, Patricia Commiskey, Qingyuan Song, Bradley A. Malin, Zhijun Yin |
WWW | 6 |
| 2025 | Large language models are less effective at clinical prediction tasks than locally trained machine learning modelsabstractOBJECTIVES: To determine the extent to which current large language models (LLMs) can serve as substitutes for traditional machine learning (ML) as clinical predictors using data from electronic health records (EHRs), we investigated various factors that can impact their adoption, including overall performance, calibration, fairness, and resilience to privacy protections that reduce data fidelity. MATERIALS AND METHODS: We evaluated GPT-3.5, GPT-4, and traditional ML (as gradient-boosting trees) on clinical prediction tasks in EHR data from Vanderbilt University Medical Center (VUMC) and MIMIC IV. We measured predictive performance with area under the receiver operating characteristic (AUROC) and model calibration using Brier Score. To evaluate the impact of data privacy protections, we assessed AUROC when demographic variables are generalized. We evaluated algorithmic fairness using equalized odds and statistical parity across race, sex, and age of patients. We also considered the impact of using in-context learning by incorporating labeled examples within the prompt. RESULTS: Traditional ML [AUROC: 0.847, 0.894 (VUMC, MIMIC)] substantially outperformed GPT-3.5 (AUROC: 0.537, 0.517) and GPT-4 (AUROC: 0.629, 0.602) (with and without in-context learning) in predictive performance and output probability calibration [Brier Score (ML vs GPT-3.5 vs GPT-4): 0.134 vs 0.384 vs 0.251, 0.042 vs 0.06 vs 0.219)]. DISCUSSION: Traditional ML is more robust than GPT-3.5 and GPT-4 in generalizing demographic information to protect privacy. GPT-4 is the fairest model according to our selected metrics but at the cost of poor model performance. CONCLUSION: These findings suggest that non-fine-tuned LLMs are less effective and robust than locally trained ML for clinical prediction tasks, but they are improving across releases. Katherine E. Brown, Chao Yan 0004, Xinmeng Zhang, Benjamin X. Collins, You Chen 0001, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 10 |
| 2025 | Beyond Phecodes: leveraging PheMAP to identify patients lacking diagnosis codes in electronic health recordsabstractOBJECTIVE: Diagnosis codes documented in electronic health records (EHR) are often relied upon to clinically phenotype patients for biomedical research. However, these diagnoses can be incomplete and inaccurate, leading to false negatives when searching for patients with phenotypes of interest. This study aims to determine whether PheMAP, a comprehensive knowledgebase integrating multiple clinical terminologies beyond diagnosis to capture phenotypes, can effectively identify patients lacking relevant EHR diagnosis codes. MATERIALS AND METHODS: We investigated a collection of 3.5 million patient records from Vanderbilt University Medical Center's EHR and focused on 4 well-studied phenotypes: (1) type 2 diabetes mellitus (T2DM), (2) dementia, (3) prostate cancer, and (4) sensorineural hearing loss. We applied PheMAP to match structured concepts in patient records and calculated a phenotype risk score (PheScore) to indicate patient-phenotype similarity. Patients meeting predefined PheScore criteria but lacking diagnosis codes were identified. Clinically knowledgeable experts adjudicated randomly selected patients per phenotype as Positive, Possibly Positive, or Negative. RESULTS: Our approach indicated that 5.3% of patients lacked a diagnosis for T2DM, 4.5% for dementia, 2.2% for prostate cancer, and 0.2% for sensorineural hearing loss. The expert review indicated 100% precision (for Possibly Positive or Positive cases) for dementia and sensorineural hearing loss, and 90.0% and 85.0% precision for T2DM and prostate cancer, respectively. Excluding Possibly Positive cases, the precision for T2DM and prostate cancer was 88.9% and 81.3%, respectively. CONCLUSIONS: Leveraging clinical terminologies incorporated by PheMAP can effectively identify patients with phenotypes who lack EHR diagnosis codes, thereby enhancing phenotyping quality and related research reliability. Chao Yan 0004, Monika E. Grabowska, Rut Thakkar, Alyson L. Dickson, Peter J. Embí, QiPing Feng, Joshua C. Denny, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Analyzing Inference Privacy Risks Through Gradients In Machine LearningabstractIn distributed learning settings, models are iteratively updated with shared gradients computed from potentially sensitive user data. While previous work has studied various privacy risks of sharing gradients, our paper aims to provide a systematic approach to analyze private information leakage from gradients. We present a unified game-based framework that encompasses a broad range of attacks including attribute, property, distributional, and user disclosures. We investigate how different uncertainties of the adversary affect their inferential power via extensive experiments on five datasets across various data modalities. Our results demonstrate the inefficacy of solely relying on data aggregation to achieve privacy against inference attacks in distributed learning. We further evaluate five types of defenses, namely, gradient pruning, signed gradient descent, adversarial perturbations, variational information bottleneck, and differential privacy, under both static and adaptive adversary settings. We provide an information-theoretic view for analyzing the effectiveness of these defenses against inference from gradients. Finally, we introduce a method for auditing attribute inference privacy, improving the empirical estimation of worst-case privacy through crafting adversarial canary records. Andrew Lowy, Jing Liu 0009, Toshiaki Koike-Akino, Kieran Parsons, Bradley A. Malin, Ye Wang 0001 |
CCS | 6 |
| 2024 | Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented GenerationabstractLanguage models (LMs) are known to suffer from hallucinations and misinformation.Retrieval augmented generation (RAG) that retrieves verifiable information from an external knowledge corpus to complement the parametric knowledge in LMs provides a tangible solution to these problems.However, the generation quality of RAG is highly dependent on the relevance between a user's query and the retrieved documents.Inaccurate responses may be generated when the query is outside of the scope of knowledge represented in the external knowledge corpus or if the information in the corpus is out-of-date.In this work, we establish a statistical framework that assesses how well a query can be answered by an RAG system by capturing the relevance of knowledge.We introduce an online testing procedure that employs goodness-of-fit (GoF) tests to inspect the relevance of each user query to detect out-of-knowledge queries with low knowledge relevance.Additionally, we develop an offline testing framework that examines a collection of user queries, aiming to detect significant shifts in the query distribution which indicates the knowledge corpus is no longer sufficiently capable of supporting the interests of the users.We demonstrate the capabilities of these strategies through a systematic evaluation on eight question-answering (QA) datasets, the results of which indicate that the new testing framework is an efficient solution to enhance the reliability of existing RAG systems. Jiaxin Zhang 0005, Chao Yan 0004, Kamalika Das, Kumar Sricharan, Murat Kantarcioglu, Bradley A. Malin |
EMNLP | 7 |
| 2024 | A Flexible Generative Model for Heterogeneous Tabular EHR with Missing ModalityabstractRealistic synthetic electronic health records (EHRs) can be leveraged to acceler- ate methodological developments for research purposes while mitigating privacy concerns associated with data sharing. However, the training of Generative Ad- versarial Networks remains challenging, often resulting in issues like mode col- lapse. While diffusion models have demonstrated progress in generating qual- ity synthetic samples for tabular EHRs given ample denoising steps, their perfor- mance wanes when confronted with missing modalities in heterogeneous tabular EHRs data. For example, some EHRs contain solely static measurements, and some contain only contain temporal measurements, or a blend of both data types. To bridge this gap, we introduce FLEXGEN-EHR– a versatile diffusion model tai- lored for heterogeneous tabular EHRs, equipped with the capability of handling missing modalities in an integrative learning framework. We define an optimal transport module to align and accentuate the common feature space of hetero- geneity of EHRs. We empirically show that our model consistently outperforms existing state-of-the-art synthetic EHR generation methods both in fidelity by up to 3.10% and utility by up to 7.16%. Additionally, we show that our method can be successfully used in privacy-sensitive settings, where the original patient-level data cannot be shared. William Hao, Yuanzhe Xi, Yong Chen 0016, Bradley A. Malin, Joyce C. Ho |
ICLR | 5 |
| 2024 | Robin Hood: A De-identification Method to Preserve Minority Representation for Disparities Research
J. Thomas Brown, Ellen Wright Clayton, Michael E. Matheny, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
PSD | 6 |
| 2024 | Evaluating site-of-care-related racial disparities in kidney graft failure using a novel federated learning frameworkabstractOBJECTIVES: Racial disparities in kidney transplant access and posttransplant outcomes exist between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in the United States, with the site of care being a key contributor. Using multi-site data to examine the effect of site of care on racial disparities, the key challenge is the dilemma in sharing patient-level data due to regulations for protecting patients' privacy. MATERIALS AND METHODS: We developed a federated learning framework, named dGEM-disparity (decentralized algorithm for Generalized linear mixed Effect Model for disparity quantification). Consisting of 2 modules, dGEM-disparity first provides accurately estimated common effects and calibrated hospital-specific effects by requiring only aggregated data from each center and then adopts a counterfactual modeling approach to assess whether the graft failure rates differ if NHB patients had been admitted at transplant centers in the same distribution as NHW patients were admitted. RESULTS: Utilizing United States Renal Data System data from 39 043 adult patients across 73 transplant centers over 10 years, we found that if NHB patients had followed the distribution of NHW patients in admissions, there would be 38 fewer deaths or graft failures per 10 000 NHB patients (95% CI, 35-40) within 1 year of receiving a kidney transplant on average. DISCUSSION: The proposed framework facilitates efficient collaborations in clinical research networks. Additionally, the framework, by using counterfactual modeling to calculate the event rate, allows us to investigate contributions to racial disparities that may occur at the level of site of care. CONCLUSIONS: Our framework is broadly applicable to other decentralized datasets and disparities research related to differential access to care. Ultimately, our proposed framework will advance equity in human health by identifying and addressing hospital-level racial disparities. Jiayi Tong, Yishan Shen, Alice Xu, Xing He 0003, Chongliang Luo, Mackenzie J. Edmondson, Dazheng Zhang, Chao Yan 0004, Ruowang Li, Lianne Siegel, Lichao Sun 0001, Elizabeth Shenkman, Sally C. Morton, Bradley A. Malin, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 15 |
| 2024 | Large language models facilitate the generation of electronic health record phenotyping algorithmsabstractOBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms. Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 12 |
| 2024 | Leveraging generative AI for clinical evidence synthesis needs to ensure trustworthiness
Qiao Jin 0001, Denis Jered McInerney, Yong Chen 0016, Fei Wang 0001, Curtis L. Cole, Qian Yang 0004, Yanshan Wang, Bradley A. Malin, Mor Peleg, Byron C. Wallace, Zhiyong Lu, Chunhua Weng, Yifan Peng 0002 |
J. Biomed. Informatics | 9 |
| 2023 | Managing re-identification risks while providing access to the All of Us research programabstractOBJECTIVE: The All of Us Research Program makes individual-level data available to researchers while protecting the participants' privacy. This article describes the protections embedded in the multistep access process, with a particular focus on how the data was transformed to meet generally accepted re-identification risk levels. METHODS: At the time of the study, the resource consisted of 329 084 participants. Systematic amendments were applied to the data to mitigate re-identification risk (eg, generalization of geographic regions, suppression of public events, and randomization of dates). We computed the re-identification risk for each participant using a state-of-the-art adversarial model specifically assuming that it is known that someone is a participant in the program. We confirmed the expected risk is no greater than 0.09, a threshold that is consistent with guidelines from various US state and federal agencies. We further investigated how risk varied as a function of participant demographics. RESULTS: The results indicated that 95th percentile of the re-identification risk of all the participants is below current thresholds. At the same time, we observed that risk levels were higher for certain race, ethnic, and genders. CONCLUSIONS: While the re-identification risk was sufficiently low, this does not imply that the system is devoid of risk. Rather, All of Us uses a multipronged data protection strategy that includes strong authentication practices, active monitoring of data misuse, and penalization mechanisms for users who violate terms of service. Weiyi Xia, Melissa A. Basford, Robert J. Carroll, Ellen Wright Clayton, Paul A. Harris, Murat Kantarcioglu, Yongtai Liu, Steve Nyemba, Yevgeniy Vorobeychik, Zhiyu Wan, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | Defending Against Membership Inference Attacks on Beacon ServicesabstractLarge genomic datasets are created through numerous activities, including recreational genealogical investigations, biomedical research, and clinical care. At the same time, genomic data has become valuable for reuse beyond their initial point of collection, but privacy concerns often hinder access. Beacon services have emerged to broaden accessibility to such data. These services enable users to query for the presence of a particular minor allele in a dataset, and information helps care providers determine if genomic variation is spurious or has some known clinical indication. However, various studies have shown that this process can leak information regarding if individuals are members of the underlying dataset. There are various approaches to mitigate this vulnerability, but they are limited in that they (1) typically rely on heuristics to add noise to the Beacon responses; (2) offer probabilistic privacy guarantees only, neglecting data utility; and (3) assume a batch setting where all queries arrive at once. In this article, we present a novel algorithmic framework to ensure privacy in a Beacon service setting with a minimal number of query response flips. We represent this problem as one of combinatorial optimization in both the batch setting and the online setting (where queries arrive sequentially). We introduce principled algorithms with both privacy and, in some cases, worst-case utility guarantees. Moreover, through extensive experiments, we show that the proposed approaches significantly outperform the state of the art in terms of privacy and utility, using a dataset consisting of 800 individuals and 1.3 million single nucleotide variants. Rajagopal Venkatesaramani, Zhiyu Wan, Bradley A. Malin, Yevgeniy Vorobeychik |
ACM Trans. Priv. Secur. | 3 |
| 2022 | Assessing Machine Learning Based Generators for Synthetic Electronic Health Records: A Benchmarking
Chao Yan 0004, Ziqi Zhang 0005, Zhiyu Wan, Justin Guinney, Sean D. Mooney, Bradley A. Malin |
AMIA | 7 |
| 2022 | A Representativeness-informed Model for Research Record Selection from Electronic Medical Record Systems
Victor A. Borza, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
AMIA | 5 |
| 2022 | Supporting COVID-19 Disparity Investigations with Dynamically Adjusting Case Reporting Policies
J. Thomas Brown, Zhiyu Wan, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin |
AMIA | 5 |
| 2022 | How to Achieve Privacy in Large Diverse Health Systems
Gamze Gürsoy, Bradley A. Malin, Erman Ayday, Ellen Wright Clayton |
AMIA | 2 |
| 2022 | A Scalable Tool for Realistic Health Data Re-identification Risk Assessment
Weiyi Xia, Yongtai Liu, Zhiyu Wan, Yevgeniy Vorobeychik, Murat Kantarcioglu, Ellen Wright Clayton, Bradley A. Malin |
AMIA | 7 |
| 2022 | Privacy-Preserving Publishing of Individual-Level Pandemic Data Based on a Game Theoretic ModelabstractSharing individual-level pandemic data is essential for accelerating the understanding of a disease. For example, COVID-19 data have been widely collected to support public health surveillance and research. In the United States, these data need to be de-identified before being released to the public due to privacy concerns. However, current data publishing approaches for individual-level pandemic data, such as those adopted by the U.S. Centers for Disease Control and Prevention (CDC), have not flexed over time to account for the dynamic nature of infection rates. Thus, the policies generated by these strategies may either raise privacy risks or impair the data utility (or usability). To optimize the tradeoff between privacy risk and data utility, we introduce a game theoretic model that adaptively generates policies to publish individual-level COVID-19 data according to infection dynamics. We model the data publishing process as a two-player Stackelberg game between a data publisher and a data recipient and then search for the best strategy for the publisher. In this game, we consider 1) the average accuracy of predicting future case counts for all demographic groups, and 2) the mutual information between the original data and the released data. We use COVID-19 case data from Vanderbilt University Medical Center from March 2020 to December 2021 to demonstrate our model and evaluate its effectiveness. The experimental results show that our game theoretic model outperforms all baseline approaches, including those adopted by CDC, while maintaining low privacy risk. Abinitha Gourabathina, Zhiyu Wan, J. Thomas Brown, Chao Yan 0004, Bradley A. Malin |
BIBM | 5 |
| 2022 | GINN: Fast GPU-TEE Based Integrity for Neural Network TrainingabstractMachine learning models based on Deep Neural Networks (DNNs) are increasingly deployed in a wide variety of applications, ranging from self-driving cars to COVID-19 diagnosis. To support the computational power necessary to train a DNN, cloud environments with dedicated Graphical Processing Unit (GPU) hardware support have emerged as critical infrastructure. However, there are many integrity challenges associated with outsourcing the computation to use GPU power, due to its inherent lack of safeguards to ensure computational integrity. Various approaches have been developed to address these challenges, building on trusted execution environments (TEE). Yet, no existing approach scales up to support realistic integrity-preserving DNN model training for heavy workloads (e.g., deep architectures and millions of training examples) without sustaining a significant performance hit. To mitigate the running time difference between pure TEE (i.e., full integrity) and pure GPU (i.e., no integrity) , we combine random verification of selected computation steps with systematic adjustments of DNN hyperparameters (e.g., a narrow gradient clipping range), which limits the attacker's ability to shift the model parameters arbitrarily. Experimental analysis shows that the new approach can achieve a 2X to 20X performance improvement over a pure TEE-based solution while guaranteeing an extremely high probability of integrity (e.g., 0.999) with respect to state-of-the-art DNN backdoor attacks. Aref Asvadishirehjini, Murat Kantarcioglu, Bradley A. Malin |
CODASPY | 3 |
| 2022 | "Rough Day ... Need a Hug": Learning Challenges and Experiences of the Alzheimer's Disease and Related Dementia Caregivers on Reddit
Congning Ni, Bradley A. Malin, Angela L. Jefferson, Patricia Commiskey, Zhijun Yin |
ICWSM | 2 |
| 2022 | How Adversarial Assumptions Influence Re-identification Risk Measures: A COVID-19 Case Study
Xinmeng Zhang, Zhiyu Wan, Chao Yan 0004, J. Thomas Brown, Weiyi Xia, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin |
PSD | 8 |
| 2022 | Dynamically adjusting case reporting policy to maximize privacy and public health utility in the face of a pandemicabstractOBJECTIVE: Supporting public health research and the public's situational awareness during a pandemic requires continuous dissemination of infectious disease surveillance data. Legislation, such as the Health Insurance Portability and Accountability Act of 1996 and recent state-level regulations, permits sharing deidentified person-level data; however, current deidentification approaches are limited. Namely, they are inefficient, relying on retrospective disclosure risk assessments, and do not flex with changes in infection rates or population demographics over time. In this paper, we introduce a framework to dynamically adapt deidentification for near-real time sharing of person-level surveillance data. MATERIALS AND METHODS: The framework leverages a simulation mechanism, capable of application at any geographic level, to forecast the reidentification risk of sharing the data under a wide range of generalization policies. The estimates inform weekly, prospective policy selection to maintain the proportion of records corresponding to a group size less than 11 (PK11) at or below 0.1. Fixing the policy at the start of each week facilitates timely dataset updates and supports sharing granular date information. We use August 2020 through October 2021 case data from Johns Hopkins University and the Centers for Disease Control and Prevention to demonstrate the framework's effectiveness in maintaining the PK11 threshold of 0.01. RESULTS: When sharing COVID-19 county-level case data across all US counties, the framework's approach meets the threshold for 96.2% of daily data releases, while a policy based on current deidentification techniques meets the threshold for 32.3%. CONCLUSION: Periodically adapting the data publication policies preserves privacy while enhancing public health utility through timely updates and sharing epidemiologically critical features. J. Thomas Brown, Chao Yan 0004, Weiyi Xia, Zhijun Yin, Zhiyu Wan, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 8 |
| 2022 | Dobbs and the future of health data privacy for patients and healthcare organizationsabstractThe Supreme Court recently overturned settled case law that affirmed a pregnant individual's Constitutional right to an abortion. While many states will commit to protect this right, a large number of others have enacted laws that limit or outright ban abortion within their borders. Additional efforts are underway to prevent pregnant individuals from seeking care outside their home state. These changes have significant implications for delivery of healthcare as well as for patient-provider confidentiality. In particular, these laws will influence how information is documented in and accessed via electronic health records and how personal health applications are utilized in the consumer domain. We discuss how these changes may lead to confusion and conflict regarding use of health information, both within and across state lines, why current health information security practices may need to be reconsidered, and what policy options may be possible to protect individuals' health information. Ellen Wright Clayton, Peter J. Embí, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Keeping synthetic patients on track: feedback mechanisms to mitigate performance drift in longitudinal health data simulationabstractOBJECTIVE: Synthetic data are increasingly relied upon to share electronic health record (EHR) data while maintaining patient privacy. Current simulation methods can generate longitudinal data, but the results are unreliable for several reasons. First, the synthetic data drifts from the real data distribution over time. Second, the typical approach to quality assessment, which is based on the extent to which real records can be distinguished from synthetic records using a critic model, often fails to recognize poor simulation results. In this article, we introduce a longitudinal simulation framework, called LS-EHR, which addresses these issues. MATERIALS AND METHODS: LS-EHR enhances simulation through conditional fuzzing and regularization, rejection sampling, and prior knowledge embedding. We compare LS-EHR to the state-of-the-art using data from 60 000 EHRs from Vanderbilt University Medical Center (VUMC) and the All of Us Research Program. We assess discrimination between real and synthetic data over time. We evaluate the generation process and critic model using the area under the receiver operating characteristic curve (AUROC). For the critic, a higher value indicates a more robust model for quality assessment. For the generation process, a lower value indicates better synthetic data quality. RESULTS: The LS-EHR critic improves discrimination AUROC from 0.655 to 0.909 and 0.692 to 0.918 for VUMC and All of Us data, respectively. By using the new critic, the LS-EHR generation model reduces the AUROC from 0.909 to 0.758 and 0.918 to 0.806. CONCLUSION: LS-EHR can substantially improve the usability of simulated longitudinal EHR data. Ziqi Zhang 0005, Chao Yan 0004, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Forecasting the future clinical events of a patient through contrastive learningabstractOBJECTIVE: Deep learning models for clinical event forecasting (CEF) based on a patient's medical history have improved significantly over the past decade. However, their transition into practice has been limited, particularly for diseases with very low prevalence. In this paper, we introduce CEF-CL, a novel method based on contrastive learning to forecast in the face of a limited number of positive training instances. MATERIALS AND METHODS: CEF-CL consists of two primary components: (1) unsupervised contrastive learning for patient representation and (2) supervised transfer learning over the derived representation. We evaluate the new method along with state-of-the-art model architectures trained in a supervised manner with electronic health records data from Vanderbilt University Medical Center and the All of Us Research Program, covering 48 000 and 16 000 patients, respectively. We assess forecasting for over 100 diagnosis codes with respect to their area under the receiver operator characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). We investigate the correlation between forecasting performance improvement and code prevalence via a Wald Test. RESULTS: CEF-CL achieved an average AUROC and AUPRC performance improvement over the state-of-the-art of 8.0%-9.3% and 11.7%-32.0%, respectively. The improvement in AUROC was negatively correlated with the number of positive training instances (P < .001). CONCLUSION: This investigation indicates that clinical event forecasting can be improved significantly through contrastive representation learning, especially when the number of positive training instances is small. Ziqi Zhang 0005, Chao Yan 0004, Xinmeng Zhang, Steve Nyemba, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Membership inference attacks against synthetic health data
Ziqi Zhang 0005, Chao Yan 0004, Bradley A. Malin |
J. Biomed. Informatics | 3 |
| 2021 | Synthetic Data to Support Engineering and Demonstrations in the All of Us Research Program
Chao Yan 0004, Steve Nyemba, Kelsey R. Mayo, Ziqi Zhang 0005, Francis Ratsimbazafy, Bradley A. Malin |
AMIA | 6 |
| 2021 | Telehealth Uptake and Continuing Usage During the COVID-19 Pandemic
Bradley A. Malin, You Chen 0001 |
AMIA | 2 |
| 2021 | De-identifying Socioeconomic Data at the Census Tract Level for Medical Research Through Constraint-based Clustering
Yongtai Liu, Douglas Conway, Zhiyu Wan, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
AMIA | 6 |
| 2021 | Predicting Next-Day Discharge via Electronic Health Record Audit Logs
Xinmeng Zhang, Chao Yan 0004, Mayur B. Patel, Bradley A. Malin, You Chen 0001 |
AMIA | 4 |
| 2021 | CCF-CL: Forecasting the Clinical Status of a Patient Through Contrastive Learning
Ziqi Zhang 0005, Chao Yan 0004, Xinmeng Zhang, Steve Nyemba, Bradley A. Malin |
AMIA | 5 |
| 2021 | Mining tasks and task characteristics from electronic health record audit logs with unsupervised machine learningabstractOBJECTIVE: The characteristics of clinician activities while interacting with electronic health record (EHR) systems can influence the time spent in EHRs and workload. This study aims to characterize EHR activities as tasks and define novel, data-driven metrics. MATERIALS AND METHODS: We leveraged unsupervised learning approaches to learn tasks from sequences of events in EHR audit logs. We developed metrics characterizing the prevalence of unique events and event repetition and applied them to categorize tasks into 4 complexity profiles. Between these profiles, Mann-Whitney U tests were applied to measure the differences in performance time, event type, and clinician prevalence, or the number of unique clinicians who were observed performing these tasks. In addition, we apply process mining frameworks paired with clinical annotations to support the validity of a sample of our identified tasks. We apply our approaches to learn tasks performed by nurses in the Vanderbilt University Medical Center neonatal intensive care unit. RESULTS: We examined EHR audit logs generated by 33 neonatal intensive care unit nurses resulting in 57 234 sessions and 81 tasks. Our results indicated significant differences in performance time for each observed task complexity profile. There were no significant differences in clinician prevalence or in the frequency of viewing and modifying event types between tasks of different complexities. We presented a sample of expert-reviewed, annotated task workflows supporting the interpretation of their clinical meaningfulness. CONCLUSIONS: The use of the audit log provides an opportunity to assist hospitals in further investigating clinician activities to optimize EHR workflows. Bob Chen 0001, Mhd Wael Alrifai, Barrett Jones, Laurie L. Novak, Nancy M. Lorenzi, Daniel J. France, Bradley A. Malin, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | Predicting brain function status changes in critically ill patients via Machine learningabstractOBJECTIVE: In intensive care units (ICUs), a patient's brain function status can shift from a state of acute brain dysfunction (ABD) to one that is ABD-free and vice versa, which is challenging to forecast and, in turn, hampers the allocation of hospital resources. We aim to develop a machine learning model to predict next-day brain function status changes. MATERIALS AND METHODS: Using multicenter prospective adult cohorts involving medical and surgical ICU patients from 2 civilian and 3 Veteran Affairs hospitals, we trained and externally validated a light gradient boosting machine to predict brain function status changes. We compared the performances of the boosting model against state-of-the-art models-an ABD predictive model and its variants. We applied Shapley additive explanations to identify influential factors to develop a compact model. RESULTS: There were 1026 critically ill patients without evidence of prior major dementia, or structural brain diseases, from whom 12 295 daily transitions (ABD: 5847 days; ABD-free: 6448 days) were observed. The boosting model achieved an area under the receiver-operating characteristic curve (AUROC) of 0.824 (95% confidence interval [CI], 0.821-0.827), compared with the state-of-the-art models of 0.697 (95% CI, 0.693-0.701) with P < .001. Using 13 identified top influential factors, the compact model achieved 99.4% of the boosting model on AUROC. The boosting and the compact models demonstrated high generalizability in external validation by achieving an AUROC of 0.812 (95% CI, 0.812-0.813). CONCLUSION: The inputs of the compact model are based on several simple questions that clinicians can quickly answer in practice, which demonstrates the model has direct prospective deployment potential into clinical practice, aiding in critical hospital resource allocation. Chao Yan 0004, Ziqi Zhang 0005, Wencong Chen, Bradley A. Malin, Eugene Wesley Ely, Mayur B. Patel, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | SynTEG: a framework for temporal structured electronic health data simulationabstractOBJECTIVE: Simulating electronic health record data offers an opportunity to resolve the tension between data sharing and patient privacy. Recent techniques based on generative adversarial networks have shown promise but neglect the temporal aspect of healthcare. We introduce a generative framework for simulating the trajectory of patients' diagnoses and measures to evaluate utility and privacy. MATERIALS AND METHODS: The framework simulates date-stamped diagnosis sequences based on a 2-stage process that 1) sequentially extracts temporal patterns from clinical visits and 2) generates synthetic data conditioned on the learned patterns. We designed 3 utility measures to characterize the extent to which the framework maintains feature correlations and temporal patterns in clinical events. We evaluated the framework with billing codes, represented as phenome-wide association study codes (phecodes), from over 500 000 Vanderbilt University Medical Center electronic health records. We further assessed the privacy risks based on membership inference and attribute disclosure attacks. RESULTS: The simulated temporal sequences exhibited similar characteristics to real sequences on the utility measures. Notably, diagnosis prediction models based on real versus synthetic temporal data exhibited an average relative difference in area under the ROC curve of 1.6% with standard deviation of 3.8% for 1276 phecodes. Additionally, the relative difference in the mean occurrence age and time between visits were 4.9% and 4.2%, respectively. The privacy risks in synthetic data, with respect to the membership and attribute inference were negligible. CONCLUSION: This investigation indicates that temporal diagnosis code sequences can be simulated in a manner that provides utility and respects privacy. Ziqi Zhang 0005, Chao Yan 0004, Thomas A. Lasko, Jimeng Sun 0001, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Predicting next-day discharge via electronic health record access logsabstractOBJECTIVE: Hospital capacity management depends on accurate real-time estimates of hospital-wide discharges. Estimation by a clinician requires an excessively large amount of effort and, even when attempted, accuracy in forecasting next-day patient-level discharge is poor. This study aims to support next-day discharge predictions with machine learning by incorporating electronic health record (EHR) audit log data, a resource that captures EHR users' granular interactions with patients' records by communicating various semantics and has been neglected in outcome predictions. MATERIALS AND METHODS: This study focused on the EHR data for all adults admitted to Vanderbilt University Medical Center in 2019. We learned multiple advanced models to assess the value that EHR audit log data adds to the daily prediction of discharge likelihood within 24 h and to compare different representation strategies. We applied Shapley additive explanations to identify the most influential types of user-EHR interactions for discharge prediction. RESULTS: The data include 26 283 inpatient stays, 133 398 patient-day observations, and 819 types of user-EHR interactions. The model using the count of each type of interaction in the recent 24 h and other commonly used features, including demographics and admission diagnoses, achieved the highest area under the receiver operating characteristics (AUROC) curve of 0.921 (95% CI: 0.919-0.923). By contrast, the model lacking user-EHR interactions achieved a worse AUROC of 0.862 (0.860-0.865). In addition, 10 of the 20 (50%) most influential factors were user-EHR interaction features. CONCLUSION: EHR audit log data contain rich information such that it can improve hospital-wide discharge predictions. Xinmeng Zhang, Chao Yan 0004, Bradley A. Malin, Mayur B. Patel, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Lucene-P$^2$2: A Distributed Platform for Privacy-Preserving Text-Based SearchabstractInformation retrieval (IR) plays an essential role in daily life. However, currently deployed IR technologies, e.g., Apache Lucene – open-source search software, are insufficient when the information is protected or deemed to be private. For example, submitting a query to a publicly available search engine (e.g., Bing or Google) requires disclosing potentially delicate facts (e.g., thoughts about abortion), as well as the websites the user considers interesting. Similarly, when a private database contains sensitive information needed by the user, it cannot be searched freely. Over the past decade, various approaches, generally referred to as private information retrieval, have been proposed to obfuscate queries and responses, but they are limited in that the retrieved information is inadequate to compute relevancy. To address such limitations, this article introduces the necessary techniques to build Lucene-P$^2$that allows one party to discover whether a second party harbors any relevant textual information without either party disclosing any information. Nitish M. Uplavikar, Bradley A. Malin, Wei Jiang 0026 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | Generating Electronic Health Records with Multiple Data Types and Constraints
Chao Yan 0004, Ziqi Zhang 0005, Steve Nyemba, Bradley A. Malin |
AMIA | 4 |
| 2020 | Learning Tasks of Pediatric Providers from Electronic Health Record Audit Logs
Barrett Jones, Xinmeng Zhang, Bradley A. Malin, You Chen 0001 |
AMIA | 3 |
| 2020 | De-Identification of Clinical Text: Stakeholders' Perspectives and Acceptance of Automatic De-Identification
Stéphane M. Meystre, Jonathan C. Silverstein, Guergana K. Savova, Valentina Petkov, Bradley A. Malin |
AMIA | 5 |
| 2020 | To Warn or Not to Warn: Online Signaling in Audit GamesabstractRoutine operational use of sensitive data is often governed by law and regulation. For instance, in the medical domain, there are various statues at the state and federal level that dictate who is permitted to work with patients' records and under what conditions. To screen for potential privacy breaches, logging systems are usually deployed to trigger alerts whenever a suspicious access is detected. However, such mechanisms are often inefficient because 1) the vast majority of triggered alerts are false positives, 2) small budgets make it unlikely that a real attack will be detected, and 3) attackers can behave strategically, such that traditional auditing mechanisms cannot easily catch them. To improve efficiency, information systems may invoke signaling, so that whenever a suspicious access request occurs, the system can, in real time, warn the user that the access may be audited. Then, at the close of a finite period, a selected subset of suspicious accesses are audited. This gives rise to an online problem in which one needs to determine 1) whether a warning should be triggered and 2) the likelihood that the data request event will be audited. In this paper, we formalize this auditing problem as a Signaling Audit Game (SAG), in which we model the interactions between an auditor and an attacker in the context of signaling and the usability cost is represented as a factor of the auditor's payoff. We study the properties of its Stackelberg equilibria and develop a scalable approach to compute its solution. We show that a strategic presentation of warnings adds value in that SAGs realize significantly higher utility for the auditor than systems without signaling. We perform a series of experiments with 10 million real access events, containing over 26K alerts, from a large academic medical center to illustrate the value of the proposed auditing model and the consistency of its advantages over existing baseline methods. Chao Yan 0004, Yevgeniy Vorobeychik, Bo Li 0026, Daniel Fabbri, Bradley A. Malin |
ICDE | 6 |
| 2020 | Resilience of clinical text de-identified with "hiding in plain sight" to hostile reidentification attacks by human readersabstractOBJECTIVE: Effective, scalable de-identification of personally identifying information (PII) for information-rich clinical text is critical to support secondary use, but no method is 100% effective. The hiding-in-plain-sight (HIPS) approach attempts to solve this "residual PII problem." HIPS replaces PII tagged by a de-identification system with realistic but fictitious (resynthesized) content, making it harder to detect remaining unredacted PII. MATERIALS AND METHODS: Using 2000 representative clinical documents from 2 healthcare settings (4000 total), we used a novel method to generate 2 de-identified 100-document corpora (200 documents total) in which PII tagged by a typical automated machine-learned tagger was replaced by HIPS-resynthesized content. Four readers conducted aggressive reidentification attacks to isolate leaked PII: 2 readers from within the originating institution and 2 external readers. RESULTS: Overall, mean recall of leaked PII was 26.8% and mean precision was 37.2%. Mean recall was 9% (mean precision = 37%) for patient ages, 32% (mean precision = 26%) for dates, 25% (mean precision = 37%) for doctor names, 45% (mean precision = 55%) for organization names, and 23% (mean precision = 57%) for patient names. Recall was 32% (precision = 40%) for internal and 22% (precision =33%) for external readers. DISCUSSION AND CONCLUSIONS: Approximately 70% of leaked PII "hiding" in a corpus de-identified with HIPS resynthesis is resilient to detection by human readers in a realistic, aggressive reidentification attack scenario-more than double the rate reported in previous studies but less than the rate reported for an attack assisted by machine learning methods. David Carrell, Bradley A. Malin, David J. Cronkite, John S. Aberdeen, Cheryl Clark, Muqun Li, Dikshya Bastakoty, Steve Nyemba, Lynette Hirschman |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | SCOR: A secure international informatics infrastructure to investigate COVID-19abstractGlobal pandemics call for large and diverse healthcare data to study various risk factors, treatment options, and disease progression patterns. Despite the enormous efforts of many large data consortium initiatives, scientific community still lacks a secure and privacy-preserving infrastructure to support auditable data sharing and facilitate automated and legally compliant federated analysis on an international scale. Existing health informatics systems do not incorporate the latest progress in modern security and federated machine learning algorithms, which are poised to offer solutions. An international group of passionate researchers came together with a joint mission to solve the problem with our finest models and tools. The SCOR Consortium has developed a ready-to-deploy secure infrastructure using world-class privacy and security technologies to reconcile the privacy/utility conflicts. We hope our effort will make a change and accelerate research in future pandemics with broad and diverse samples on an international scale. Jean Louis Raisaro, Juan Ramón Troncoso-Pastoriza, Raphaelle Beau-Lejdstrom, Riccardo Bellazzi, Robert Murphy, Elmer V. Bernstam, Henry Wang, Mauro Bucalo, Yong Chen 0016, Assaf Gottlieb, Arif Ozgun Harmanci, Miran Kim, Yejin Kim 0001, Jeffrey G. Klann, Catherine Klersy, Bradley A. Malin, Marie Méan, Fabian Prasser, Luigia Scudeller, Ali Torkamani, Julien Vaucher, Mamta Puppala, Stephen T. C. Wong, Milana Frenkel-Morgenstern, Hua Xu 0001, Baba Maiyaki Musa, Abdulrazaq G. Habib, Trevor Cohen, Adam B. Wilcox, Hamisu M. Salihu, Heidi Sofia, Xiaoqian Jiang, Jean-Pierre Hubaux |
J. Am. Medical Informatics Assoc. | 17 |
| 2020 | Ensuring electronic medical record simulation through better training, modeling, and evaluationabstractOBJECTIVE: Electronic medical records (EMRs) can support medical research and discovery, but privacy risks limit the sharing of such data on a wide scale. Various approaches have been developed to mitigate risk, including record simulation via generative adversarial networks (GANs). While showing promise in certain application domains, GANs lack a principled approach for EMR data that induces subpar simulation. In this article, we improve EMR simulation through a novel pipeline that (1) enhances the learning model, (2) incorporates evaluation criteria for data utility that informs learning, and (3) refines the training process. MATERIALS AND METHODS: We propose a new electronic health record generator using a GAN with a Wasserstein divergence and layer normalization techniques. We designed 2 utility measures to characterize similarity in the structural properties of real and simulated EMRs in the original and latent space, respectively. We applied a filtering strategy to enhance GAN training for low-prevalence clinical concepts. We evaluated the new and existing GANs with utility and privacy measures (membership and disclosure attacks) using billing codes from over 1 million EMRs at Vanderbilt University Medical Center. RESULTS: The proposed model outperformed the state-of-the-art approaches with significant improvement in retaining the nature of real records, including prediction performance and structural properties, without sacrificing privacy. Additionally, the filtering strategy achieved higher utility when the EMR training dataset was small. CONCLUSIONS: These findings illustrate that EMR simulation through GANs can be substantially improved through more appropriate training, modeling, and evaluation criteria. Ziqi Zhang 0005, Chao Yan 0004, Diego Mesa, Jimeng Sun 0001, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Corpus Size Influences Clinical Concept Embeddings
Chao Yan 0004, Bradley A. Malin, You Chen 0001 |
AMIA | 3 |
| 2019 | Biomedical Research Cohort Membership Disclosure on Social Media
Yongtai Liu, Chao Yan 0004, Zhijun Yin, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Bradley A. Malin |
AMIA | 9 |
| 2019 | Public Attitudes Toward Direct to Consumer Genetic Testing
Grayson Ruhl, James Hazel, Ellen Wright Clayton, Bradley A. Malin |
AMIA | 4 |
| 2019 | Why Patient Portal Messages Indicate Risk of Readmission for Patients with Ischemic Heart Disease
Lina M. Sulieman, Zhijun Yin, Bradley A. Malin |
AMIA | 3 |
| 2019 | Patient Messaging Content Associated with Initiating Hormonal Therapy after a Breast Cancer Diagnosis
Zhijun Yin, Jeremy L. Warner, Qingxia Chen, Bradley A. Malin |
AMIA | 4 |
| 2019 | Determining the Impact of Missing Values on Blocking in Record Linkage
Imrul Chowdhury Anindya, Murat Kantarcioglu, Bradley A. Malin |
PAKDD (3) | 3 |
| 2019 | CP Tensor Decomposition with Cannot-Link Intermode ConstraintsabstractTensor factorization is a methodology that is applied in a variety of fields, ranging from climate modeling to medical informatics. A tensor is an n-way array that captures the relationship between n objects. These multiway arrays can be factored to study the underlying bases present in the data. Two challenges arising in tensor factorization are 1) the resulting factors can be noisy and highly overlapping with one another and 2) they may not map to insights within a domain. However, incorporating supervision to increase the number of insightful factors can be costly in terms of the time and domain expertise necessary for gathering labels or domain-specific constraints. To meet these challenges, we introduce CANDECOMP/PARAFAC (CP) tensor factorization with Cannot-Link Intermode Constraints (CP-CLIC), a framework that achieves succinct, diverse, interpretable factors. This is accomplished by gradually learning constraints that are verified with auxiliary information during the decomposition process. We demonstrate CP-CLIC's potential to extract sparse, diverse, and interpretable factors through experiments on simulated data and a real-world application in medical informatics. Jette Henderson, Bradley A. Malin, Joshua C. Denny, Abel N. Kho, Jimeng Sun 0001, Joydeep Ghosh, Joyce C. Ho |
SDM | 2 |
| 2019 | The machine giveth and the machine taketh away: a parrot attack on clinical text deidentified with hiding in plain sightabstractOBJECTIVE: Clinical corpora can be deidentified using a combination of machine-learned automated taggers and hiding in plain sight (HIPS) resynthesis. The latter replaces detected personally identifiable information (PII) with random surrogates, allowing leaked PII to blend in or "hide in plain sight." We evaluated the extent to which a malicious attacker could expose leaked PII in such a corpus. MATERIALS AND METHODS: We modeled a scenario where an institution (the defender) externally shared an 800-note corpus of actual outpatient clinical encounter notes from a large, integrated health care delivery system in Washington State. These notes were deidentified by a machine-learned PII tagger and HIPS resynthesis. A malicious attacker obtained and performed a parrot attack intending to expose leaked PII in this corpus. Specifically, the attacker mimicked the defender's process by manually annotating all PII-like content in half of the released corpus, training a PII tagger on these data, and using the trained model to tag the remaining encounter notes. The attacker hypothesized that untagged identifiers would be leaked PII, discoverable by manual review. We evaluated the attacker's success using measures of leak-detection rate and accuracy. RESULTS: The attacker correctly hypothesized that 211 (68%) of 310 actual PII leaks in the corpus were leaks, and wrongly hypothesized that 191 resynthesized PII instances were also leaks. One-third of actual leaks remained undetected. DISCUSSION AND CONCLUSION: A malicious parrot attack to reveal leaked PII in clinical text deidentified by machine-learned HIPS resynthesis can attenuate but not eliminate the protective effect of HIPS deidentification. David Carrell, David J. Cronkite, Muqun Li, Steve Nyemba, Bradley A. Malin, John S. Aberdeen, Lynette Hirschman |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | A systematic literature review of machine learning in online personal health dataabstractOBJECTIVE: User-generated content (UGC) in online environments provides opportunities to learn an individual's health status outside of clinical settings. However, the nature of UGC brings challenges in both data collecting and processing. The purpose of this study is to systematically review the effectiveness of applying machine learning (ML) methodologies to UGC for personal health investigations. MATERIALS AND METHODS: We searched PubMed, Web of Science, IEEE Library, ACM library, AAAI library, and the ACL anthology. We focused on research articles that were published in English and in peer-reviewed journals or conference proceedings between 2010 and 2018. Publications that applied ML to UGC with a focus on personal health were identified for further systematic review. RESULTS: We identified 103 eligible studies which we summarized with respect to 5 research categories, 3 data collection strategies, 3 gold standard dataset creation methods, and 4 types of features applied in ML models. Popular off-the-shelf ML models were logistic regression (n = 22), support vector machines (n = 18), naive Bayes (n = 17), ensemble learning (n = 12), and deep learning (n = 11). The most investigated problems were mental health (n = 39) and cancer (n = 15). Common health-related aspects extracted from UGC were treatment experience, sentiments and emotions, coping strategies, and social support. CONCLUSIONS: The systematic review indicated that ML can be effectively applied to UGC in facilitating the description and inference of personal health. Future research needs to focus on mitigating bias introduced when building study cohorts, creating features from free text, improving clinical creditability of UGC, and model interpretability. Zhijun Yin, Lina M. Sulieman, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Deep learning predicts extreme preterm birth from electronic health records
Sarah Osmundson, Digna Velez Edwards, Gretchen Purcell Jackson, Bradley A. Malin, You Chen 0001 |
J. Biomed. Informatics | 5 |
| 2019 | Systematizing Genome Privacy Research: A Privacy-Enhancing Technologies PerspectiveabstractAbstract Rapid advances in human genomics are enabling researchers to gain a better understanding of the role of the genome in our health and well-being, stimulating hope for more effective and cost efficient healthcare. However, this also prompts a number of security and privacy concerns stemming from the distinctive characteristics of genomic data. To address them, a new research community has emerged and produced a large number of publications and initiatives. In this paper, we rely on a structured methodology to contextualize and provide a critical analysis of the current knowledge on privacy-enhancing technologies used for testing, storing, and sharing genomic data, using a representative sample of the work published in the past decade. We identify and discuss limitations, technical challenges, and issues faced by the community, focusing in particular on those that are inherently tied to the nature of the problem and are harder for the community alone to address. Finally, we report on the importance and difficulty of the identified challenges based on an online survey of genome data privacy experts. Alexandros Mittos, Bradley A. Malin, Emiliano De Cristofaro |
Proc. Priv. Enhancing Technol. | 2 |
| 2019 | Database Audit Workload Prioritization via Game TheoryabstractThe quantity of personal data that is collected, stored, and subsequently processed continues to grow rapidly. Given its sensitivity, ensuring privacy protections has become a necessary component of database management. To enhance protection, a number of mechanisms have been developed, such as audit logging and alert triggers, which notify administrators about suspicious activities. However, this approach is limited. First, the volume of alerts is often substantially greater than the auditing capabilities of organizations. Second, strategic attackers can attempt to disguise their actions or carefully choose targets, thus hide illicit activities. In this article, we introduce an auditing approach that accounts for adversarial behavior by (1) prioritizing the order in which types of alerts are investigated and (2) providing an upper bound on how much resource to allocate for each type. Specifically, we model the interaction between a database auditor and attackers as a Stackelberg game. We show that even a highly constrained version of such problem is NP-Hard. Then, we introduce a method that combines linear programming, column generation, and heuristic searching to derive an auditing policy. On the synthetic data, we perform an extensive evaluation on the approximation degree of our solution with the optimal one. The two real datasets, (1) 1.5 months of audit logs from Vanderbilt University Medical Center and (2) a publicly available credit card application dataset, are used to test the policy-searching performance. The findings demonstrate the effectiveness of the proposed methods for searching the audit strategies, and our general approach significantly outperforms non-game-theoretic baselines. Chao Yan 0004, Bo Li 0026, Yevgeniy Vorobeychik, Aron Laszka, Daniel Fabbri, Bradley A. Malin |
ACM Trans. Priv. Secur. | 6 |
| 2018 | Crowdsourcing Clinical Chart Reviews
Joseph R. Coco, Cheng Ye 0001, Chen Hajaj, Yevgeniy Vorobeychik, Joshua C. Denny, Laurie L. Novak, Bradley A. Malin, Thomas A. Lasko, Daniel Fabbri |
AMIA | 7 |
| 2018 | Phenotyping through Semi-Supervised Tensor Factorization (PSST)
Jette Henderson, Bradley A. Malin, Joshua C. Denny, Abel N. Kho, Joydeep Ghosh, Joyce C. Ho |
AMIA | 3 |
| 2018 | Detecting the Presence of an Individual in Phenotypic Summary Data
Yongtai Liu, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Abel N. Kho, David Carrell, Bradley A. Malin |
AMIA | 9 |
| 2018 | Learning When Communications Between Healthcare Providers Indicate Hormonal Therapy Medication Discontinuation
Zhijun Yin, Jeremy L. Warner, Bradley A. Malin |
AMIA | 3 |
| 2018 | Get Your Workload in Order: Game Theoretic Prioritization of Database AuditingabstractA wide variety of mechanisms, such as alert triggers and auditing routines, have been developed to notify administrators about types of suspicious activities in the daily use of large databases of personal and sensitive information. However, such mechanisms are limited in that: 1) the volume of such alerts is often substantially greater than the auditing capabilities of budget-constrained organizations and 2) strategic attackers may disguise their actions or carefully choose which records they touch, thus evading auditing routines. To address these problems, we introduce a novel approach to database auditing that explicitly accounts for adversarial behavior by 1) prioritizing the order in which types of alerts are investigated and 2) providing an upper bound on how much budget to allocate for auditing each alert type. We model the interaction between a database auditor and potential attackers as a Stackelberg game in which the auditor chooses an auditing policy and attackers choose which records in a database to target. We further introduce an efficient approach that combines linear programming, column generation, and heuristic search to derive an auditing policy, in the form of a mixed strategy. We assess the performance of the policy selection method using a publicly available credit card application dataset, the results of which indicate that our method produces high-quality database audit policies, significantly outperforming baselines that are not based in a game theoretic framing. Chao Yan 0004, Bo Li 0026, Yevgeniy Vorobeychik, Aron Laszka, Daniel Fabbri, Bradley A. Malin |
ICDE | 6 |
| 2018 | Interaction patterns of trauma providers are associated with length of stayabstractBackground: Trauma-related hospitalizations drive a high percentage of health care expenditure and inpatient resource consumption, which is directly related to length of stay (LOS). Robust and reliable interactions among health care employees can reduce LOS. However, there is little known about whether certain patterns of interactions exist and how they relate to LOS and its variability. The objective of this study is to learn interaction patterns and quantify the relationship to LOS within a mature trauma system and long-standing electronic medical record (EMR). Methods: We adapted a spectral co-clustering methodology to infer the interaction patterns of health care employees based on the EMR of 5588 hospitalized adult trauma survivors. The relationship between interaction patterns and LOS was assessed via a negative binomial regression model. We further assessed the influence of potential confounders by age, number of health care encounters to date, number of access action types care providers committed to patient EMRs, month of admission, phenome-wide association study codes, procedure codes, and insurance status. Results: Three types of interaction patterns were discovered. The first pattern exhibited the most collaboration between employees and was associated with the shortest LOS. Compared to this pattern, LOS for the second and third patterns was 0.61 days (P = 0.014) and 0.43 days (P = 0.037) longer, respectively. Although the 3 interaction patterns dealt with different numbers of patients in each admission month, our results suggest that care was provided for similar patients. Discussion: The results of this study indicate there is an association between LOS and the extent to which health care employees interact in the care of an injured patient. The findings further suggest that there is merit in ascertaining the content of these interactions and the factors that induce these differences in interaction patterns within a trauma system. You Chen 0001, Mayur B. Patel, Candace D. McNaughton, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | It's all in the timing: calibrating temporal penalties for biomedical data sharingabstractObjective: Biomedical science is driven by datasets that are being accumulated at an unprecedented rate, with ever-growing volume and richness. There are various initiatives to make these datasets more widely available to recipients who sign Data Use Certificate agreements, whereby penalties are levied for violations. A particularly popular penalty is the temporary revocation, often for several months, of the recipient's data usage rights. This policy is based on the assumption that the value of biomedical research data depreciates significantly over time; however, no studies have been performed to substantiate this belief. This study investigates whether this assumption holds true and the data science policy implications. Methods: This study tests the hypothesis that the value of data for scientific investigators, in terms of the impact of the publications based on the data, decreases over time. The hypothesis is tested formally through a mixed linear effects model using approximately 1200 publications between 2007 and 2013 that used datasets from the Database of Genotypes and Phenotypes, a data-sharing initiative of the National Institutes of Health. Results: The analysis shows that the impact factors for publications based on Database of Genotypes and Phenotypes datasets depreciate in a statistically significant manner. However, we further discover that the depreciation rate is slow, only ∼10% per year, on average. Conclusion: The enduring value of data for subsequent studies implies that revoking usage for short periods of time may not sufficiently deter those who would violate Data Use Certificate agreements and that alternative penalty mechanisms may need to be invoked. Weiyi Xia, Zhiyu Wan, Zhijun Yin, James Gaupp, Yongtai Liu, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 9 |
| 2018 | The therapy is making me sick: how online portal communications between breast cancer patients and physicians indicate medication discontinuationabstractObjective: Online platforms have created a variety of opportunities for breast patients to discuss their hormonal therapy, a long-term adjuvant treatment to reduce the chance of breast cancer occurrence and mortality. The goal of this investigation is to ascertain the extent to which the messages breast cancer patients communicated through an online portal can indicate their potential for discontinuing hormonal therapy. Materials and Methods: We studied the de-identified electronic medical records of 1106 breast cancer patients who were prescribed hormonal therapy at Vanderbilt University Medical Center over a 12-year period. We designed a data-driven approach to investigate patients' patterns of messaging with healthcare providers, the topics they communicated, and the extent to which these messaging behaviors associate with the likelihood that a patient will discontinue a prescribed 5-year regimen of therapy. Results: The results indicates that messaging rate over time [hazard ratio (HR) = 1.373, P = 0.002], mentions of side effects (HR = 1.214, P = 0.006), and surgery-related topics (HR = 1.170, P = 0.034) were associated with increased risk of early medication discontinuation. In contrast, seeking professional suggestions (HR = 0.766, P = 0.002), expressing gratitude to healthcare providers (HR = 0.872, P = 0.044), and mentions of drugs used to treat side effects (HR = 0.807, P = 0.013) were associated with decreased risk of medication discontinuation. Discussion and Conclusion: This investigation suggests that patient-generated content can inform the study of health-related behaviors. Given that approximately 50% of breast cancer patients do not complete a course of hormonal therapy as described, the identification of factors associated with medication discontinuation can facilitate real-time interventions to prevent early discontinuation. Zhijun Yin, Morgan Harrell, Jeremy L. Warner, Qingxia Chen, Daniel Fabbri, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Learning bundled care opportunities from electronic medical records
You Chen 0001, Abel N. Kho, David M. Liebovitz, Catherine Ivory, Sarah Osmundson, Jiang Bian 0001, Bradley A. Malin |
J. Biomed. Informatics | 7 |
| 2018 | Development of an automated phenotyping algorithm for hepatorenal syndrome
Jejo Koola, Sharon E. Davis, Omar Al-Nimri, Sharidan K. Parr, Daniel Fabbri, Bradley A. Malin, Samuel B. Ho, Michael E. Matheny |
J. Biomed. Informatics | 6 |
| 2018 | GenoPri'16: International Workshop on Genome Privacy and SecurityabstractThe three papers included this special section were presented at the 3rd International Workshop on Genome Privacy and Security (GenoPri) in 2016. GenoPri’16 was collocated with the American Medical Informatics Association Annual Fall Symposium (AMIA), a premier medical informatics venue. Erman Ayday, Xiaoqian Jiang, Bradley A. Malin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Evaluating the Effectiveness of Auditing Rules for Electronic Health Record Systems
Monica Hedda, Bradley A. Malin, Chao Yan 0004, Daniel Fabbri |
AMIA | 2 |
| 2017 | An Open Source Tool for Game Theoretic Health Data De-Identification
Fabian Prasser, James Gaupp, Zhiyu Wan, Weiyi Xia, Yevgeniy Vorobeychik, Murat Kantarcioglu, Klaus A. Kuhn, Bradley A. Malin |
AMIA | 8 |
| 2017 | Computational Phenotyping on Diverse Data Sources
Jimeng Sun 0001, Bradley A. Malin, Abel N. Kho, Mark W. Craven, Joydeep Ghosh |
AMIA | 2 |
| 2017 | Talking About My Care: Detecting Mentions of Hormonal Therapy Adherence Behavior in an Online Breast Cancer Community
Zhijun Yin, Wei Xie 0002, Bradley A. Malin |
AMIA | 3 |
| 2017 | Reciprocity and its Association with Treatment Adherence in an Online Breast Cancer ForumabstractOnline health communities (OHCs) are increasingly relied upon by individuals exchanging social support for diagnoses and treatment regimens. It has been shown that social support from trusted relationships (e.g., family and friends) positively influence treatment adherence in offline environments, but much less is known about the online setting. In this study, we focus on how relationships established in an online breast cancer discussion board induce reciprocity (specifically in the form of reciprocal exchange of support) and its impact on adherence to a five-year hormonal therapy, a highly prevalent long-term treatment for breast cancers, with varying completion rates. We measure reciprocity as responses to forum posts to analyze interactions of over 6,000 patients and 100,000 responses. In doing so, we assess how reciprocity is related to time active in the OHC and the tones communicated by authors in their posts (e.g., emotions, writing styles and social tendencies). We further assess if such reciprocity is associated with treatment adherence. We find the volume of the reciprocity is positively associated with completing the five-year protocol, rather than the rate of the reciprocity or the fraction of the posts that received replies. Zhijun Yin, Bradley A. Malin |
CBMS | 3 |
| 2017 | Building a Dossier on the Cheap: Integrating Distributed Personal Data Resources Under Cost ConstraintsabstractA wide variety of personal data is routinely collected by numerous organizations that, in turn, share and sell their collections for analytic investigations (e.g., market research). To preserve privacy, certain identifiers are often redacted, perturbed or even removed. A substantial number of attacks have shown that, if care is not taken, such data can be linked to external resources to determine the explicit identifiers (e.g., personal names) or infer sensitive attributes (e.g., income) for the individuals from whom the data was collected. As such, organizations increasingly rely upon record linkage methods to assess the risk such attacks pose and adopt countermeasures accordingly. Traditional linkage methods assume only two datasets would be linked (e.g., linking de-identified hospital discharge to identified voter registration lists), but with the advent of a multi-billion dollar data broker industry, modern adversaries have access to a massive data stash of multiple datasets that can be leveraged. Still, realistic adversaries have budget constraints that prevent them from obtaining and integrating all relevant datasets. Thus, in this work, we investigate a novel privacy risk assessment framework, based on adversaries who plan an integration of datasets for the most accurate estimate of targeted sensitive attributes under a certain budget. To solve this problem, we introduce a graph-based formulation of the problem and predictive modeling methods to prioritize data resources for linkage. We perform an empirical analysis using real world voter registration data from two different U.S. states and show that the methods can be used efficiently to accurately estimate potentially sensitive information disclosure risks even under a non-trivial amount of noise. Imrul Chowdhury Anindya, Harichandan Roy, Murat Kantarcioglu, Bradley A. Malin |
CIKM | 4 |
| 2017 | The Power of the Patient Voice: Learning Indicators of Treatment Adherence From An Online Breast Cancer Forum
Zhijun Yin, Bradley A. Malin, Jeremy L. Warner, Pei-Yun Sabrina Hsueh, Ching-Hua Chen |
ICWSM | 2 |
| 2017 | Identifying collaborative care teams through electronic medical record utilization patternsabstractOBJECTIVE: The goal of this investigation was to determine whether automated approaches can learn patient-oriented care teams via utilization of an electronic medical record (EMR) system. MATERIALS AND METHODS: To perform this investigation, we designed a data-mining framework that relies on a combination of latent topic modeling and network analysis to infer patterns of collaborative teams. We applied the framework to the EMR utilization records of over 10 000 employees and 17 000 inpatients at a large academic medical center during a 4-month window in 2010. Next, we conducted an extrinsic evaluation of the patterns to determine the plausibility of the inferred care teams via surveys with knowledgeable experts. Finally, we conducted an intrinsic evaluation to contextualize each team in terms of collaboration strength (via a cluster coefficient) and clinical credibility (via associations between teams and patient comorbidities). RESULTS: The framework discovered 34 collaborative care teams, 27 (79.4%) of which were confirmed as administratively plausible. Of those, 26 teams depicted strong collaborations, with a cluster coefficient > 0.5. There were 119 diagnostic conditions associated with 34 care teams. Additionally, to provide clarity on how the survey respondents arrived at their determinations, we worked with several oncologists to develop an illustrative example of how a certain team functions in cancer care. DISCUSSION: Inferred collaborative teams are plausible; translating such patterns into optimized collaborative care will require administrative review and integration with management practices. CONCLUSIONS: EMR utilization records can be mined for collaborative care patterns in large complex medical centers. You Chen 0001, Nancy M. Lorenzi, Warren S. Sandberg, Kelly Wolgast, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Towards a privacy preserving cohort discovery framework for clinical research networks
Bradley A. Malin, François Modave, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
J. Biomed. Informatics | 2 |
| 2017 | Scalable Iterative Classification for Sanitizing Large-Scale DatasetsabstractCheap ubiquitous computing enables the collection of massive amounts of personal data in a wide variety of domains. Many organizations aim to share such data while obscuring features that could disclose personally identifiable information. Much of this data exhibits weak structure (e.g., text), such that machine learning approaches have been developed to detect and remove identifiers from it. While learning is never perfect, and relying on such approaches to sanitize data can leak sensitive information, a small risk is often acceptable. Our goal is to balance the value of published data and the risk of an adversary discovering leaked identifiers. We model data sanitization as a game between 1) a publisher who chooses a set of classifiers to apply to data and publishes only instances predicted as non-sensitive and 2) an attacker who combines machine learning and manual inspection to uncover leaked identifying information. We introduce a fast iterative greedy algorithm for the publisher that ensures a low utility for a resource-limited adversary. Moreover, using five text data sets we illustrate that our algorithm leaves virtually no automatically identifiable sensitive instances for a state-of-the-art learning algorithm, while sharing over 93% of the original data, and completes after at most 5 iterations. Bo Li 0026, Yevgeniy Vorobeychik, Muqun Li, Bradley A. Malin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2016 | Big Data for Healthcare and Life Sciences: Learning Useful Insights from Imperfect Data
Jianying Hu, Nigam H. Shah, Bradley A. Malin, Patrick B. Ryan |
AMIA | 3 |
| 2016 | Predicting Negative Events: Using Post-discharge Data to Detect High-Risk Patients
Lina M. Sulieman, Daniel Fabbri, Fei Wang 0001, Jianying Hu, Bradley A. Malin |
AMIA | 5 |
| 2016 | Learning Clinical Workflows to Identify Subgroups of Heart Failure Patients
Chao Yan 0004, You Chen 0001, Bo Li 0026, David M. Liebovitz, Bradley A. Malin |
AMIA | 5 |
| 2016 | CheapSMC: A Framework to Minimize Secure Multiparty Computation Cost in the Cloud
Erman Pattuk, Murat Kantarcioglu, Huseyin Ulusoy, Bradley A. Malin |
DBSec | 4 |
| 2016 | Optimizing secure classification performance with privacy-aware feature selectionabstractRecent advances in personalized medicine point towards a future where clinical decision making will be dependent upon the individual characteristics of the patient, e.g., their age, race, genomic variation, and lifestyle. Already, there are numerous commercial entities working towards the provision of software to support such decisions as cloud-based services. However, deployment of such services in such settings raises important challenges for privacy. A recent attack shows that disclosing personalized drug dosage recommendations, combined with several pieces of demographic knowledge, can be leveraged to infer single nucleotide polymorphism variants of a patient. One manner to prevent such inference is to apply secure multi-party computation (SMC) techniques that hide all patient data, so that no information, including the clinical recommendation, is disclosed during the decision making process. Yet, SMC is a computationally cumbersome process and disclosing some information may be necessary for various compliance purposes. Additionally, certain information (e.g., demographic information) may already be publicly available. In this work, we provide a novel approach to selectively disclose certain information before the SMC process to significantly improve personalized decision making performance while preserving desired levels of privacy. To achieve this goal, we introduce mechanisms to quickly compute the loss in privacy due to information disclosure while considering its performance impact on SMC execution phase. Our empirical analysis show that we can achieve up to three orders of magnitude improvement compared to pure SMC solutions with only a slight increase in privacy risks. Erman Pattuk, Murat Kantarcioglu, Huseyin Ulusoy, Bradley A. Malin |
ICDE | 4 |
| 2016 | #PrayForDad: Learning the Semantics Behind Why Social Media Users Disclose Health Information
Zhijun Yin, You Chen 0001, Daniel Fabbri, Jimeng Sun 0001, Bradley A. Malin |
ICWSM | 5 |
| 2016 | A multi-institution evaluation of clinical profile anonymizationabstractBACKGROUND AND OBJECTIVE: There is an increasing desire to share de-identified electronic health records (EHRs) for secondary uses, but there are concerns that clinical terms can be exploited to compromise patient identities. Anonymization algorithms mitigate such threats while enabling novel discoveries, but their evaluation has been limited to single institutions. Here, we study how an existing clinical profile anonymization fares at multiple medical centers. METHODS: We apply a state-of-the-artk-anonymization algorithm, withkset to the standard value 5, to the International Classification of Disease, ninth edition codes for patients in a hypothyroidism association study at three medical centers: Marshfield Clinic, Northwestern University, and Vanderbilt University. We assess utility when anonymizing at three population levels: all patients in 1) the EHR system; 2) the biorepository; and 3) a hypothyroidism study. We evaluate utility using 1) changes to the number included in the dataset, 2) number of codes included, and 3) regions generalization and suppression were required. RESULTS: Our findings yield several notable results. First, we show that anonymizing in the context of the entire EHR yields a significantly greater quantity of data by reducing the amount of generalized regions from ∼15% to ∼0.5%. Second, ∼70% of codes that needed generalization only generalized two or three codes in the largest anonymization. CONCLUSIONS: Sharing large volumes of clinical data in support of phenome-wide association studies is possible while safeguarding privacy to the underlying individuals. Raymond Heatherly, Luke V. Rasmussen, Peggy L. Peissig, Jennifer A. Pacheco, Paul A. Harris, Joshua C. Denny, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 7 |
| 2016 | Preserving temporal relations in clinical data while maintaining privacyabstractOBJECTIVE: Maintaining patient privacy is a challenge in large-scale observational research. To assist in reducing the risk of identifying study subjects through publicly available data, we introduce a method for obscuring date information for clinical events and patient characteristics. METHODS: The method, which we call Shift and Truncate (SANT), obscures date information to any desired granularity. Shift and Truncate first assigns each patient a random shift value, such that all dates in that patient's record are shifted by that amount. Data are then truncated from the beginning and end of the data set. RESULTS: The data set can be proven to not disclose temporal information finer than the chosen granularity. Unlike previous strategies such as a simple shift, it remains robust to frequent - even daily - updates and robust to inferring dates at the beginning and end of date-shifted data sets. Time-of-day may be retained or obscured, depending on the goal and anticipated knowledge of the data recipient. CONCLUSIONS: The method can be useful as a scientific approach for reducing re-identification risk under the Privacy Rule of the Health Insurance Portability and Accountability Act and may contribute to qualification for the Safe Harbor implementation. George Hripcsak, Parsa Mirhaji, Alexander F. H. Low, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | Optimizing annotation resources for natural language de-identification via a game theoretic framework
Muqun Li, David Carrell, John S. Aberdeen, Lynette Hirschman, Jacqueline Kirby, Bo Li 0026, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Biomed. Informatics | 8 |
| 2015 | Inferring Clinical Workflow Efficiency via Electronic Medical Record Utilization
You Chen 0001, Wei Xie 0002, Carl A. Gunter, David M. Liebovitz, Sanjay Mehrotra, Bradley A. Malin |
AMIA | 7 |
| 2015 | Mining Twitter as a First Step toward Assessing the Adequacy of Gender Identification Terms on Intake Forms
Amanda Hicks, William R. Hogan, Michael W. Rutherford, Bradley A. Malin, Mengjun Xie, Christiane Fellbaum, Zhijun Yin, Daniel Fabbri, Josh Hanna, Jiang Bian 0001 |
AMIA | 4 |
| 2015 | Patient privacy and "de-identified" health records in the genomic era
Jessica D. Tenenbaum, Greg Biggers, Bradley A. Malin, Lucila Ohno-Machado, Leslie Wolf |
AMIA | 3 |
| 2015 | Process-Driven Data PrivacyabstractThe quantity of personal data gathered by service providers via our daily activities continues to grow at a rapid pace. The sharing, and the subsequent analysis of, such data can support a wide range of activities, but concerns around privacy often prompt an organization to transform the data to meet certain protection models (e.g., k-anonymity or ε-differential privacy). These models, however, are based on simplistic adversarial frameworks, which can lead to both under- and over-protection. For instance, such models often assume that an adversary attacks a protected record exactly once. We introduce a principled approach to explicitly model the attack process as a series of steps. Specifically, we engineer a factored Markov decision process (FMDP) to optimally plan an attack from the adversary's perspective and assess the privacy risk accordingly. The FMDP captures the uncertainty in the adversary's belief (e.g., the number of identified individuals that match the de-identified data) and enables the analysis of various real world deterrence mechanisms beyond a traditional protection model, such as a penalty for committing an attack. We present an algorithm to solve the FMDP and illustrate its efficiency by simulating an attack on publicly accessible U.S. census records against a real identified resource of over 500,000 individuals in a voter registry. Our results demonstrate that while traditional privacy models commonly expect an adversary to attack exactly once per record, an optimal attack in our model may involve exploiting none, one, or more individuals in the pool of candidates, depending on context. Weiyi Xia, Murat Kantarcioglu, Zhiyu Wan, Raymond Heatherly, Yevgeniy Vorobeychik, Bradley A. Malin |
CIKM | 6 |
| 2015 | Privacy-aware dynamic feature selectionabstractBig data will enable the development of novel services that enhance a company's market advantage, competition, or productivity. At the same time, the utilization of such a service could disclose sensitive data in the process, which raises significant privacy concerns. To protect individuals, various policies, such as the Code of Fair Information Practices, as well as recent laws require organizations to capture only the minimal amount of data necessary to support a service. While this is a notable goal, choosing the minimal data is a non-trivial process, especially while considering privacy and utility constraints. In this paper, we introduce a technique to minimize sensitive data disclosure by focusing on privacy-aware feature selection. During model deployment, the service provider requests only a subset of the available features from the client, such that it can produce results with maximal confidence, while minimizing its ability to violate a client's privacy. We propose an iterative approach, where the server requests information one feature at a time until the client-specified privacy budget is exhausted. The overall process is dynamic, such that the feature selected at each step depends on the previously selected features and their corresponding values. We demonstrate our technique with three popular classification algorithms and perform an empirical analysis over three real world datasets to illustrate that, in almost all cases, classifiers that select features using our strategy have the same error-rate as state-of-the art static feature selection methods that fail to preserve privacy. Erman Pattuk, Murat Kantarcioglu, Huseyin Ulusoy, Bradley A. Malin |
ICDE | 4 |
| 2015 | Iterative Classification for Sanitizing Large-Scale DatasetsabstractCheap ubiquitous computing enables the collection of massive amounts of personal data in a wide variety of domains. Many organizations aimto share such data while obscuring features that could discloseidentities or other sensitive information. Much of the data now collected exhibits weak structure (e.g., natural language text) and machine learning approaches have been developed to identify andremove sensitive entities in such data. Learning-based approaches are never perfect and relying upon them tosanitize datacan leak sensitive information as a consequence. However, a small amount of risk is permissible in practice, and, thus, our goal is to balance the value of datapublished and the risk of an adversary discovering leaked sensitiveinformation. We model data sanitization as a game between1) a publisher who chooses a set of classifiers to apply to data andpublishes only instances predicted to be non-sensitive and 2) an attackerwho combines machine learning and manual inspection to uncover leakedsensitive entities (e.g., personal names). We introduce aniterative greedy algorithm for the publisher that provablyexecutes no more than a linear number of iterations, and ensures a lowutility for a resource-limited adversary. Moreover, using several real world natural language corpora, weillustrate that our greedy algorithm leaves virtually no automaticallyidentifiable sensitive instances for a state-of-the-art learningalgorithm, while sharing over 93% of the original data, and completesafter at most 5 iterations. Bo Li 0026, Yevgeniy Vorobeychik, Muqun Li, Bradley A. Malin |
ICDM | 4 |
| 2015 | Rubik: Knowledge Guided Tensor Factorization and Completion for Health Data AnalyticsabstractComputational phenotyping is the process of converting heterogeneous electronic health records (EHRs) into meaningful clinical concepts. Unsupervised phenotyping methods have the potential to leverage a vast amount of labeled EHR data for phenotype discovery. However, existing unsupervised phenotyping methods do not incorporate current medical knowledge and cannot directly handle missing, or noisy data. We propose Rubik, a constrained non-negative tensor factorization and completion method for phenotyping. Rubik incorporates 1) guidance constraints to align with existing medical knowledge, and 2) pairwise constraints for obtaining distinct, non-overlapping phenotypes. Rubik also has built-in tensor completion that can significantly alleviate the impact of noisy and missing data. We utilize the Alternating Direction Method of Multipliers (ADMM) framework to tensor factorization and completion, which can be easily scaled through parallel computing. We evaluate Rubik on two EHR datasets, one of which contains 647,118 records for 7,744 patients from an outpatient clinic, the other of which is a public dataset containing 1,018,614 CMS claims records for 472,645 patients. Our results show that Rubik can discover more meaningful and distinct phenotypes than the baselines. In particular, by using knowledge guidance constraints, Rubik can also discover sub-phenotypes for several major diseases. Rubik also runs around seven times faster than current state-of-the-art tensor methods. Finally, Rubik is scalable to large datasets containing millions of EHR records. Yichen Wang 0001, Robert Chen 0001, Joydeep Ghosh, Joshua C. Denny, Abel N. Kho, You Chen 0001, Bradley A. Malin, Jimeng Sun 0001 |
KDD | 7 |
| 2015 | Design and implementation of a privacy preserving electronic health record linkage tool in ChicagoabstractOBJECTIVE: To design and implement a tool that creates a secure, privacy preserving linkage of electronic health record (EHR) data across multiple sites in a large metropolitan area in the United States (Chicago, IL), for use in clinical research. METHODS: The authors developed and distributed a software application that performs standardized data cleaning, preprocessing, and hashing of patient identifiers to remove all protected health information. The application creates seeded hash code combinations of patient identifiers using a Health Insurance Portability and Accountability Act compliant SHA-512 algorithm that minimizes re-identification risk. The authors subsequently linked individual records using a central honest broker with an algorithm that assigns weights to hash combinations in order to generate high specificity matches. RESULTS: The software application successfully linked and de-duplicated 7 million records across 6 institutions, resulting in a cohort of 5 million unique records. Using a manually reconciled set of 11 292 patients as a gold standard, the software achieved a sensitivity of 96% and a specificity of 100%, with a majority of the missed matches accounted for by patients with both a missing social security number and last name change. Using 3 disease examples, it is demonstrated that the software can reduce duplication of patient records across sites by as much as 28%. CONCLUSIONS: Software that standardizes the assignment of a unique seeded hash identifier merged through an agreed upon third-party honest broker can enable large-scale secure linkage of EHR data for epidemiologic and public health research. The software algorithm can improve future epidemiologic research by providing more comprehensive data given that patients may make use of multiple healthcare systems. Abel N. Kho, John P. Cashy, Kathryn L. Jackson, Adam R. Pah, Satyender Goel, Jörn Boehnke, John Eric Humphries, Scott Duke Kominers, Bala Hota, Shannon A. Sims, Bradley A. Malin, Dustin D. French, Theresa Walunas, David O. Meltzer, Erin O. Kaleba, Roderick C. Jones, William L. Galanter |
J. Am. Medical Informatics Assoc. | 11 |
| 2015 | R-U policy frontiers for health data de-identificationabstractOBJECTIVE: The Health Insurance Portability and Accountability Act Privacy Rule enables healthcare organizations to share de-identified data via two routes. They can either 1) show re-identification risk is small (e.g., via a formal model, such as k-anonymity) with respect to an anticipated recipient or 2) apply a rule-based policy (i.e., Safe Harbor) that enumerates attributes to be altered (e.g., dates to years). The latter is often invoked because it is interpretable, but it fails to tailor protections to the capabilities of the recipient. The paper shows rule-based policies can be mapped to a utility (U) and re-identification risk (R) space, which can be searched for a collection, or frontier, of policies that systematically trade off between these goals. METHODS: We extend an algorithm to efficiently compose an R-U frontier using a lattice of policy options. Risk is proportional to the number of patients to which a record corresponds, while utility is proportional to similarity of the original and de-identified distribution. We allow our method to search 20 000 rule-based policies (out of 2(700)) and compare the resulting frontier with k-anonymous solutions and Safe Harbor using the demographics of 10 U.S. states. RESULTS: The results demonstrate the rule-based frontier 1) consists, on average, of 5000 policies, 2% of which enable better utility with less risk than Safe Harbor and 2) the policies cover a broader spectrum of utility and risk than k-anonymity frontiers. CONCLUSIONS: R-U frontiers of de-identification policies can be discovered efficiently, allowing healthcare organizations to tailor protections to anticipated needs and trustworthiness of recipients. Weiyi Xia, Raymond Heatherly, Xiaofeng Ding 0001, Jiuyong Li, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Building bridges across electronic health record systems through inferred phenotypic topics
You Chen 0001, Joydeep Ghosh, Cosmin Adrian Bejan, Carl A. Gunter, Siddharth Gupta 0005, Abel N. Kho, David M. Liebovitz, Jimeng Sun 0001, Joshua C. Denny, Bradley A. Malin |
J. Biomed. Informatics | 10 |
| 2014 | SOEMPI: A Secure Open Enterprise Master Patient Index Software Toolkit for Private Record Linkage
Csaba Tóth, Elizabeth Durham, Murat Kantarcioglu, Yuan Xue 0001, Bradley A. Malin |
AMIA | 5 |
| 2014 | Decide Now or Decide Later?: Quantifying the Tradeoff between Prospective and Retrospective Access DecisionsabstractOne of the greatest challenges an organization faces is determining when an employee is permitted to utilize a certain resource in a system. This "insider threat" can be addressed through two strategies: i) prospective methods, such as access control, that make a decision at the time of a request, and ii) retrospective methods, such as post hoc auditing, that make a decision in the light of the knowledge gathered afterwards. While it is recognized that each strategy has a distinct set of benefits and drawbacks, there has been little investigation into how to provide system administrators with practical guidance on when one or the other should be applied. To address this problem, we introduce a framework to compare these strategies on a common quantitative scale. In doing so, we translate these strategies into classification problems using a context-based feature space that assesses the likelihood that an access request is legitimate. We then introduce a technique called bispective analysis to compare the performance of the classification models under the situation of non-equivalent costs for false positive and negative instances, a significant extension on traditional cost analysis techniques, such as analysis of the receiver operator characteristic (ROC) curve. Using domain-specific cost estimates and access logs of several months from a large Electronic Medical Record (EMR) system, we demonstrate how bispective analysis can support meaningful decisions about the relative merits of prospective and retrospective decision making for specific types of hospital personnel. You Chen 0001, Thaddeus Cybulski, Daniel Fabbri, Carl A. Gunter, Patrick N. Lawlor, David M. Liebovitz, Bradley A. Malin |
CCS | 8 |
| 2014 | SecureMA: protecting participant privacy in genetic association meta-analysisabstractMOTIVATION: Sharing genomic data is crucial to support scientific investigation such as genome-wide association studies. However, recent investigations suggest the privacy of the individual participants in these studies can be compromised, leading to serious concerns and consequences, such as overly restricted access to data. RESULTS: We introduce a novel cryptographic strategy to securely perform meta-analysis for genetic association studies in large consortia. Our methodology is useful for supporting joint studies among disparate data sites, where privacy or confidentiality is of concern. We validate our method using three multisite association studies. Our research shows that genetic associations can be analyzed efficiently and accurately across substudy sites, without leaking information on individual participants and site-level association summaries. AVAILABILITY AND IMPLEMENTATION: Our software for secure meta-analysis of genetic association studies, SecureMA, is publicly available at http://github.com/XieConnect/SecureMA. Our customized secure computation framework is also publicly available at http://github.com/XieConnect/CircuitService. Wei Xie 0002, Murat Kantarcioglu, William S. Bush, Dana C. Crawford, Joshua C. Denny, Raymond Heatherly, Bradley A. Malin |
Bioinform. | 7 |
| 2014 | Predicting changes in hypertension control using electronic health records from a chronic disease management programabstractOBJECTIVE: Common chronic diseases such as hypertension are costly and difficult to manage. Our ultimate goal is to use data from electronic health records to predict the risk and timing of deterioration in hypertension control. Towards this goal, this work predicts the transition points at which hypertension is brought into, as well as pushed out of, control. METHOD: In a cohort of 1294 patients with hypertension enrolled in a chronic disease management program at the Vanderbilt University Medical Center, patients are modeled as an array of features derived from the clinical domain over time, which are distilled into a core set using an information gain criteria regarding their predictive performance. A model for transition point prediction was then computed using a random forest classifier. RESULTS: The most predictive features for transitions in hypertension control status included hypertension assessment patterns, comorbid diagnoses, procedures and medication history. The final random forest model achieved a c-statistic of 0.836 (95% CI 0.830 to 0.842) and an accuracy of 0.773 (95% CI 0.766 to 0.780). CONCLUSIONS: This study achieved accurate prediction of transition points of hypertension control status, an important first step in the long-term goal of developing personalized hypertension management plans. Jimeng Sun 0001, Candace D. McNaughton, Ping Zhang 0016, Adam Perer, Aris Gkoulalas-Divanis, Joshua C. Denny, Jacqueline Kirby, Thomas A. Lasko, Alexander Saip, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 10 |
| 2014 | Size matters: How population size influences genotype-phenotype association studies in anonymized data
Raymond Heatherly, Joshua C. Denny, Jonathan L. Haines, Dan M. Roden, Bradley A. Malin |
J. Biomed. Informatics | 5 |
| 2014 | Limestone: High-throughput candidate phenotype generation via tensor factorization
Joyce C. Ho, Joydeep Ghosh, Steven R. Steinhubl, Walter F. Stewart, Joshua C. Denny, Bradley A. Malin, Jimeng Sun 0001 |
J. Biomed. Informatics | 6 |
| 2014 | PARAMO: A PARAllel predictive MOdeling platform for healthcare analytic research using electronic health records
Kenney Ng, Amol Ghoting, Steven R. Steinhubl, Walter F. Stewart, Bradley A. Malin, Jimeng Sun 0001 |
J. Biomed. Informatics | 5 |
| 2014 | A probabilistic approach to mitigate composition attacks on privacy in non-coordinated environments
A. H. M. Sarowar Sattar, Jiuyong Li, Jixue Liu, Raymond Heatherly, Bradley A. Malin |
Knowl. Based Syst. | 5 |
| 2014 | Composite Bloom Filters for Secure Record Linkageabstract), however, when databases are maintained by disparate organizations, the disclosure of such information can breach the privacy of the corresponding individuals. Various private record linkage (PRL) methods have been developed to obscure such identifiers, but they vary widely in their ability to balance competing goals of accuracy, efficiency and security. The tokenization and hashing of field values into Bloom filters (BF) enables greater linkage accuracy and efficiency than other PRL methods, but the encodings may be compromised through frequency-based cryptanalysis. Our objective is to adapt a BF encoding technique to mitigate such attacks with minimal sacrifices in accuracy and efficiency. To accomplish these goals, we introduce a statistically-informed method to generate BF encodings that integrate bits from multiple fields, the frequencies of which are provably associated with a minimum number of fields. Our method enables a user-specified tradeoff between security and accuracy. We compare our encoding method with other techniques using a public dataset of voter registration records and demonstrate that the increases in security come with only minor losses to accuracy. Elizabeth Durham, Murat Kantarcioglu, Yuan Xue 0001, Csaba Tóth, Mehmet Kuzu, Bradley A. Malin |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2013 | Location Bias of Identifiers in Clinical Narratives
David A. Hanauer, Qiaozhu Mei, Bradley A. Malin, Kai Zheng 0002 |
AMIA | 3 |
| 2013 | Scalable and robust key group size estimation for reducer load balancing in MapReduceabstractModern parallel computing systems, such as MapReduce, often assume data values are uniformly distributed. However, in the real world, data is often highly skewed, which may cause workload imbalance among parallel running tasks. In this paper, we study the reduce-phase skew problem in MapReduce, where reduce tasks are often assgined imbalance load (in terms of key groups). We introduce a sketch-based data structure for capturing MapReduce key group size statistics and present an optimal packing algorithm which assigns the key groups to the reducers in a load balancing manner. We perform an empirical evaluation with several real and synthetic datasets over two distinct types of applications. The results show that our load balancing algorithm can strongly mitigate the reduce-phase skew. It can decrease the overall job completion time by 45.5% of the default settings in Hadoop and by 38.3% in comparison to the state-of-the-art solution. Wei Yan 0025, Yuan Xue 0001, Bradley A. Malin |
IEEE BigData | 3 |
| 2013 | Efficient discovery of de-identification policy options through a risk-utility frontierabstractModern information technologies enable organizations to capture large quantities of person-specific data while providing routine services. Many organizations hope, or are legally required, to share such data for secondary purposes (e.g., validation of research findings) in a de-identified manner. In previous work, it was shown de-identification policy alternatives could be modeled on a lattice, which could be searched for policies that met a prespecified risk threshold (e.g., likelihood of re-identification). However, the search was limited in several ways. First, its definition of utility was syntactic - based on the level of the lattice - and not semantic - based on the actual changes induced in the resulting data. Second, the threshold may not be known in advance. The goal of this work is to build the optimal set of policies that trade-off between privacy risk (R) and utility (U), which we refer to as a R-U frontier. To model this problem, we introduce a semantic definition of utility, based on information theory, that is compatible with the lattice representation of policies. To solve the problem, we initially build a set of policies that define a frontier. We then use a probability-guided heuristic to search the lattice for policies likely to update the frontier. To demonstrate the effectiveness of our approach, we perform an empirical analysis with the Adult dataset of the UCI Machine Learning Repository. We show that our approach can construct a frontier closer to optimal than competitive approaches by searching a smaller number of policies. In addition, we show that a frequently followed de-identification policy (i.e., the Safe Harbor standard of the HIPAA Privacy Rule) is suboptimal in comparison to the frontier discovered by our approach. Weiyi Xia, Raymond Heatherly, Xiaofeng Ding 0001, Jiuyong Li, Bradley A. Malin |
CODASPY | 5 |
| 2013 | Efficient privacy-aware record integrationabstractThe integration of information dispersed among multiple repositories is a crucial step for accurate data analysis in various domains. In support of this goal, it is critical to devise procedures for identifying similar records across distinct data sources. At the same time, to adhere to privacy regulations and policies, such procedures should protect the confidentiality of the individuals to whom the information corresponds. Various private record linkage (PRL) protocols have been proposed to achieve this goal, involving secure multi-party computation (SMC) and similarity preserving data transformation techniques. SMC methods provide secure and accurate solutions to the PRL problem, but are prohibitively expensive in practice, mainly due to excessive computational requirements. Data transformation techniques offer more practical solutions, but incur the cost of information leakage and false matches. In this paper, we introduce a novel model for practical PRL, which 1) affords controlled and limited information leakage, 2) avoids false matches resulting from data transformation. Initially, we partition the data sources into blocks to eliminate comparisons for records that are unlikely to match. Then, to identify matches, we apply an efficient SMC technique between the candidate record pairs. To enable efficiency and privacy, our model leaks a controlled amount of obfuscated data prior to the secure computations. Applied obfuscation relies on differential privacy which provides strong privacy guarantees against adversaries with arbitrary background knowledge. In addition, we illustrate the practical nature of our approach through an empirical analysis with data derived from public voter records. Mehmet Kuzu, Murat Kantarcioglu, Ali Inan, Elisa Bertino, Elizabeth Durham, Bradley A. Malin |
EDBT | 6 |
| 2013 | Scalable load balancing for mapreduce-based record linkageabstractRecent research has introduced load balancing schemes that are aware of the input data distribution (i.e., data profile) to mitigate data skew and fully exploit the parallel capability of the MapReduce framework to support record linkage. However, existing solutions face a significant scalability issue when applied to massive data sets with millions or billions of blocks (a basic unit in record linkage) because their data profiles can not be maintained precisely in an efficient manner. The goal of this paper is to introduce a profiling method based on the notion of a sketch, which allows for a compact scalable solution for maintaining block size statistics. In addition, we propose two load balancing algorithms to work over sketch-based profiles while solving the data skew problem associated with record linkage. We provide an analytical analysis and extensive experiments (using Hadoop), with real and controlled synthetic data sets, to illustrate the effectiveness of our solution. The experimental results show that our load balancing algorithms can decrease the overall job completion time by 71.56% and 70.73% of the default settings in Hadoop using a set of DBLP data sets, which have 2.5 to 50.4 million records. Wei Yan 0025, Yuan Xue 0001, Bradley A. Malin |
IPCCC | 3 |
| 2013 | Modeling and detecting anomalous topic accessabstractThere has been considerable success in developing strategies to detect insider threats in information systems based on what one might call the random object access model or ROA. This approach models illegitimate users as ones who randomly access records. The goal is to use statistics, machine learning, knowledge of workflows and other techniques to support an anomaly detection framework that finds such users. In this paper we introduce and study a random topic access model or RTA aimed at users whose access may be illegitimate but is not fully random because it is focused on common semantic themes. We argue that this model is appropriate for a meaningful range of attacks and develop a system based on topic summarization that is able to formalize the model and provide anomalous user detection effectively for it. To this end, we use healthcare as an example and propose a framework for evaluating the ability to recognize various types of random users called random topic access detection or RTAD. Specifically, we utilize a combination of Latent Dirichlet Allocation (LDA), for feature extraction, a k-nearest neighbor (k-NN) algorithm for outlier detection and evaluate the ability to identify different adversarial types. We validate the technique in the context of hospital audit logs where we show varying degrees of success based on user roles and the anticipated characteristics of attackers. In particular, it was found that RTAD exhibits strong performance for roles are described by a few topics, but weaker performance when users are more topic-agnostic. Siddharth Gupta 0005, Casey Hanson, Carl A. Gunter, Mario Frank 0001, David M. Liebovitz, Bradley A. Malin |
ISI | 6 |
| 2013 | Evolving role definitions through permission invocation patternsabstractIn role-based access control (RBAC), roles are traditionally defined as sets of permissions. Roles specified by administrators may be inaccurate, however, such that data mining methods have been proposed to learn roles from actual permission utilization. These methods minimize variation from an information theoretic perspective, but they neglect the expert knowledge of administrators. In this paper, we propose a strategy to enable a controlled evolution of RBAC based on utilization. To accomplish this goal, we extend a subset enumeration framework to search candidate roles for an RBAC model that addresses an objective function which balances administrator beliefs and permission utilization. The rate of role evolution is controlled by an administrator-specified parameter. To assess effectiveness, we perform an empirical analysis using simulations, as well as a real world dataset from an electronic medical record system (EMR) in use at a large academic medical center (over 8000 users, 140 roles, and 140 permissions). We compare the results with several state-of-the-art role mining algorithms using 1) an outlier detection method on the new roles to evaluate the homogeneity of their behavior and 2)a set-based similarity measure between the original and new roles. The results illustrate our method is comparable to the state-of-the-art, but allows for a range of RBAC models which tradeoff user behavior and administrator expectations. For instance, in the EMR dataset, we find the resulting RBAC model contains 22% outliers and a distance of 0.02 to the original RBAC model when the system is biased toward administrator belief, and 13% outliers and a distance of 0.26 to the original RBAC model when biased toward permission utilization. You Chen 0001, Carl A. Gunter, David M. Liebovitz, Bradley A. Malin |
SACMAT | 5 |
| 2013 | Reducing patient re-identification risk for laboratory results within research datasetsabstractOBJECTIVE: To try to lower patient re-identification risks for biomedical research databases containing laboratory test results while also minimizing changes in clinical data interpretation. MATERIALS AND METHODS: In our threat model, an attacker obtains 5-7 laboratory results from one patient and uses them as a search key to discover the corresponding record in a de-identified biomedical research database. To test our models, the existing Vanderbilt TIME database of 8.5 million Safe Harbor de-identified laboratory results from 61 280 patients was used. The uniqueness of unaltered laboratory results in the dataset was examined, and then two data perturbation models were applied-simple random offsets and an expert-derived clinical meaning-preserving model. A rank-based re-identification algorithm to mimic an attack was used. The re-identification risk and the retention of clinical meaning for each model's perturbed laboratory results were assessed. RESULTS: Differences in re-identification rates between the algorithms were small despite substantial divergence in altered clinical meaning. The expert algorithm maintained the clinical meaning of laboratory results better (affecting up to 4% of test results) than simple perturbation (affecting up to 26%). DISCUSSION AND CONCLUSION: With growing impetus for sharing clinical data for research, and in view of healthcare-related federal privacy regulation, methods to mitigate risks of re-identification are important. A practical, expert-derived perturbation algorithm that demonstrated potential utility was developed. Similar approaches might enable administrators to select data protection scheme parameters that meet their preferences in the trade-off between the protection of privacy and the retention of clinical meaning of shared data. Ravi V. Atreya, Joshua C. Smith, Allison B. McCoy, Bradley A. Malin, Randolph A. Miller |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical textabstractOBJECTIVE: Secondary use of clinical text is impeded by a lack of highly effective, low-cost de-identification methods. Both, manual and automated methods for removing protected health information, are known to leave behind residual identifiers. The authors propose a novel approach for addressing the residual identifier problem based on the theory of Hiding In Plain Sight (HIPS). MATERIALS AND METHODS: HIPS relies on obfuscation to conceal residual identifiers. According to this theory, replacing the detected identifiers with realistic but synthetic surrogates should collectively render the few 'leaked' identifiers difficult to distinguish from the synthetic surrogates. The authors conducted a pilot study to test this theory on clinical narrative, de-identified by an automated system. Test corpora included 31 oncology and 50 family practice progress notes read by two trained chart abstractors and an informaticist. RESULTS: Experimental results suggest approximately 90% of residual identifiers can be effectively concealed by the HIPS approach in text containing average and high densities of personal identifying information. DISCUSSION: This pilot test suggests HIPS is feasible, but requires further evaluation. The results need to be replicated on larger corpora of diverse origin under a range of detection scenarios. Error analyses also suggest areas where surrogate generation techniques can be refined to improve efficacy. CONCLUSIONS: If these results generalize to existing high-performing de-identification systems with recall rates of 94-98%, HIPS could increase the effective de-identification rates of these systems to levels above 99% without further advancements in system recall. Additional and more rigorous assessment of the HIPS approach is warranted. David Carrell, Bradley A. Malin, John S. Aberdeen, Samuel Bayer, Cheryl Clark, Ben Wellner, Lynette Hirschman |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | A practical approach to achieve private medical record linkage in light of public resourcesabstractOBJECTIVE: Integration of patients' records across resources enhances analytics. To address privacy concerns, emerging strategies such as Bloom filter encodings (BFEs), enable integration while obscuring identifiers. However, recent investigations demonstrate BFEs are, in theory, vulnerable to cryptanalysis when encoded identifiers are randomly selected from a public resource. This study investigates the extent to which cryptanalysis conditions hold for (1) real patient records and (2) a countermeasure that obscures the frequencies of the identifying values in encoded datasets. DESIGN: First, to investigate the strength of cryptanalysis for real patient records, we build BFEs from identifiers in an electronic medical record system and apply cryptanalysis using identifiers in a publicly available voter registry. Second, to investigate the countermeasure under ideal cryptanalysis conditions, we compose BFEs from the identifiers that are randomly selected from a public voter registry. MEASUREMENT: We utilize precision (ie, rate of correct re-identified encodings) and computation efficiency (ie, time to complete cryptanalysis) to assess the performance of cryptanalysis in BFEs before and after application of the countermeasure. RESULTS: Cryptanalysis can achieve high precision when the encoded identifiers are composed of a random sample of a public resource (ie, a voter registry). However, we also find that the attack is less efficient and may not be practical for more realistic scenarios. By contrast, the proposed countermeasure made cryptanalysis impractical in terms of precision and efficiency. CONCLUSIONS: Performance of cryptanalysis against BFEs based on patient data is significantly lower than theoretical estimates. The proposed countermeasure makes BFEs resistant to known practical attacks. Mehmet Kuzu, Murat Kantarcioglu, Elizabeth Durham, Csaba Tóth, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2013 | Biomedical data privacy: problems, perspectives, and recent advancesabstractThe notion of privacy in the healthcare domain is at least as old as the ancient Greeks. Several decades ago, as electronic medical record (EMR) systems began to take hold, the necessity of patient privacy was recognized as a core principle, or even a right, that must be upheld.1,2 This belief was re-enforced as computers and EMRs became more common in clinical environments.3–5 However, the arrival of ultra-cheap data collection and processing technologies is fundamentally changing the face of healthcare. The traditional boundaries of primary and tertiary care environments are breaking down and health information is increasingly collected through mobile devices,6 in personal domains (eg, in one's home7), and from sensors attached on or in the human body (eg, body area networks8–10). At the same time, the detail and diversity of information collected in the context of healthcare and biomedical research is increasing at an unprecedented rate, with clinical and administrative health data being complemented with a range of *omics data, where genomics11 and proteomics12 are currently leading the charge, with other types of molecular data on the horizon.13 Healthcare organizations (HCOs) are adopting and adapting information technologies to support an expanding array of activities designed to derive value from these growing data archives, in terms of enhanced health outcomes.14 The ready availability of such large volumes of detailed data has also been accompanied by privacy invasions. Recent breach notification laws at the US federal and state levels have brought to the public's attention the scope and frequency of these invasions. For example, there are cases of healthcare provider snooping on the medical records of famous people, family, and friends, use of personal information for identity fraud, and millions of records disclosed through lost and stolen unencrypted mobile devices.15 The danger is that such publicized incidents will erode patient trust over time, and lead to privacy protective behaviors. For example, between 15% and 17% of US adults have changed their behavior to protect the privacy of their health information, doing things such as: going to another doctor, paying out-of-pocket when insured to avoid disclosure, not seeking care to avoid disclosure to an employer, giving inaccurate or incomplete information on medical history, self-treating or self-medicating rather than seeing a provider, or asking a doctor not to write down the health problem or record a less serious or embarrassing condition.16–18 A survey of service members who had been on active duty found that respondents were concerned that if they received treatment for their mental health problems, it would not be kept confidential and would have a negative impact on future job assignments and career advancement.19 Specific vulnerable populations have reported similar privacy protective behaviors, such as adolescents, people with HIV or at high risk for HIV, women undergoing genetic testing, mental health patients, and victims of domestic violence.20–26 A survey of Californian residents found that discussing depression with their primary care physician was a barrier to 15% of the respondents because of privacy concerns.27 On the other hand, some legal scholars are questioning the survival of conventional privacy expectations.28 Privacy has conventionally been defined as an individual's ability to control the disclosure of personal facts.29,30 However, privacy is also a multi-dimensional concept31,32 and any shifts in privacy expectations are not homogeneous in direction and intensity across all of these dimensions. Furthermore, advances in informatics that may be eroding individuals' control over their information are being countered by advances in privacy enhancing technologies, as well as regulatory and policy changes that give individuals back control over their information. This special issue was established to solicit current research in privacy as it is currently understood and is being redefined for emerging biomedical systems. The selected articles consider the different dimensions of privacy, and describe some novel privacy enhancing technologies and their applications, as well as the governance, regulatory, and policy mechanisms that are being used to manage privacy risks. Privacy is a major patient, provider, regulator, and legislator concern today. There is therefore a need to address these concerns in a practical way that can be deployed in the short term. Deployment must be preceded by a convincing evidence base demonstrating the rationale, costs, and benefits of an intervention. At the same time, new theoretical models and novel approaches that still need to be evaluated and tested in the field, are also necessary to ensure that the field keeps evolving. In putting together this special issue we attempted to balance these two perspectives, with articles presenting results of immediate relevance and applicability, and material covering theoretical work that remains to be proven in practical settings. There were 53 papers submitted for consideration in this issue, of which 13 were accepted for publication, for an acceptance rate of 25%. All papers were subject to a rigorous review by at least two referees and oversight by one of the guest editors. The review process for papers authored by guest editors, as well as the editor-in-chief, was handled by an unaffiliated associate editor of the journal. In addition to peer-reviewed manuscripts, two invited papers were solicited for the special issue to address the topics of privacy policy and technical data protection mechanisms. Privacy is an overloaded and complex term.33 The concept of privacy often subsumes various constructs, such as anonymity (ie, the ability to hide one's identity), confidentiality (ie, the ability to share information with a second party without the information being publicly revealed), and solitude (ie, the right to be left alone). Even when the particular construct is unambiguous, it remains difficult to have discussions around the topic of privacy because it is highly contextual, such that the expectations of privacy are often specialized to the situation.34 For instance, a patient's expectation of privacy changes when disclosing information to a care provider versus a random person on the street. The expectation is further modified by the perceived sensitivity of the health information in question. And, the extent to which health information (eg, a positive assertion of an HIV diagnosis) is deemed to be sensitive varies from patient to patient. This special issue is organized to trace the lifecycle of biomedical information, which we coarsely partition into three zones that follow the general data lifecycle as follows. The first zone corresponds to the point at which health information is collected from patients. The collection may occur while an individual is physically located at a healthcare provider or beyond (eg, such as through a website on the internet or an application running on a mobile device). In this zone, privacy tends to be concerned with who can collect health information, how much information should be collected, at what time, and for what purposes. The specific notions of privacy addressed in this zone tend to be associated with anonymity (eg, Is the recipient of the data permitted to know the identity of the information from which it is being collected?), limiting content (eg, What is the minimal amount of information that will satisfy the purpose?), and consent (eg, Did the patient agree to the terms of the data collection?). The second zone corresponds to the context in which the data have left the control of the patient and are housed in a system controlled, or accessed by, those who provide a primary service (eg, provision of care, study of biomedical data explicitly solicited for a specific research project). In this zone, privacy tends to be realized through confidentiality (eg, Who is permitted to access or use the data and for what purposes?) and security (eg, How can we ensure that the data are protected from misuse or abuse while at rest in a database or in transit between authorized entities?). The third zone corresponds to the scenario in which biomedical data are utilized for purposes which are different from their primary use. The data may be used by the organization that initially collected the data (eg, repurposing of clinical data for research) or disseminated to external entities (eg, publication of public use datasets) for the performance of certain tasks (eg, evaluation of health policies). In this zone, the privacy issues that tend to arise are anonymity and consent for individuals and groups (eg, Can data collected from a particular ethnic group be reused to study a specific phenotype?). Each of these zones can be partitioned into two interacting, although conceptually distinct, tracks. In the first track, biomedical privacy is defined and regulated via socio-legal mechanisms. This is the arena where the public, ethicists, and policy and law makers come together to define what privacy rights and responsibilities exist. In the second track, technical controls are specified and realized in working information technologies to maintain societal expectations of privacy or requirements specified in policy and law. It is critical to integrate these tracks to ensure that privacy expectations are appropriately represented in technical controls and that policies are designed to realistically account for state-of-the-art technical capabilities. It is further important that societal expectations of privacy and privacy enhancing technologies are current and cognizant of shifts in expectation of technical sophistication. The socio-legal track of this special issue commences by delving into the desires and expectations society harbors for privacy. As mentioned earlier, privacy is a societal phenomenon, such that the extent to which it is realized is dependent on how society chooses to codify the concept in policy and law. This process often begins with field studies and sessions that engage stakeholders in their preferences.35 In this vein, Caine and Hanania report on a study that asks patients who should control access to health information and what granularity of control is desirable.36 Often, the expectations of privacy are dependent on the domain in which information is collected. As the traditional boundaries of the healthcare domain expand, it is important to determine how individuals' perspectives on privacy relate to new technologies. To begin to address this issue, van der Velden and El Emam focus on the use of social media by teenage patients, and how they perceive their health information privacy when interacting online.37 Insights gained here should inform the more general health data context. The next set of papers in the socio-legal track move beyond primary uses for biomedical data and into secondary settings. In this environment, it is assumed that policies and laws have been codified. However, policy and law is dependent on the locale, such that it is critical to understand how it guides data management practices. The first paper in this group, by Pencarrick Hertzman, Meagher, and McGrail, presents a case study about how the ‘Privacy by Design’ framework was applied in British Columbia (BC), Canada, to facilitate access to health information in Population Data BC.38 To date, Population Data BC has facilitated over 350 research studies. This work is followed by the first invited paper by McGraw, which takes a look at the de-identification strategy of the US Health Insurance Portability and Accountability Act (HIPAA).39 This strategy enables HCOs to disclose information about patients in a manner that is no longer subject to oversight by the regulatory authorities because the risk that they would be individually identifiable is deemed very small. Clarifications to what de-identification means and how it can be achieved in accordance with HIPAA were recently published by the U.S. federal government.40 The paper by McGraw reports on a workshop on various stakeholders' support for the current HIPAA de-identification strategy held by the Center for Democracy and Technology and discusses policy proposals to address concerns and improve trust in the process. Peterson and colleagues then recount a recent case before the Supreme Court, Sorrell vs IMS Health, and illustrate the challenges associated with selling prescription records for various purposes, such as post-market effectiveness.41 They highlight some of the concerns associated with the dissemination of identifiable prescriber and de-identified patient information. While the previous papers focus on traditional health information, the last paper in this track, by Kosseim and colleagues, addresses privacy issues associated with the management of *omics data in particular, with a specific focus on genomics.42 This paper describes the legal and ethical principles and practices adopted in the Canadian provinces of Newfoundland and Labrador to enable research with genomic, phenomic, and genealogical data. While law and policy codify the rights and requirements for managing data privacy in the biomedical domain, information technology is necessary to uphold and ensure their realization in practice. In this regard, certain aspects of privacy can be achieved through information security. The Security Rule of HIPAA specifies various administrative, physical, and technical safeguards that covered entities must have in place (ie, required controls) or document why such protections are not prudent (ie, addressable controls). For instance, it is required that all covered entities ensure that appropriate authorization is provided before employees of an HCO access a patient's EMR. By contrast, the encryption of health information at rest within the HCO is addressable, but is not required. Along these lines, the technical track of this special issue begins with a paper by Kwon and Johnson that assesses the extent to which 250 HCOs in the USA have (or have not) adopted various security practices.43 Their analysis demonstrates patterns of leaders, followers, and laggers in their adoption, and provides recommendations for improving regulatory compliance. Although this work provides a high-level assessment of the adoption of security practices, it does not provide specific guidance on data management strategies. Thus, the next paper, by Fabbri and LeFevre, discusses a privacy threat encountered on a daily basis in primary care settings, specifically the insider threat.44 This threat is particularly important to study in the healthcare domain because traditional information security controls (eg, role-based access control) are difficult to realize in care settings due to the highly dynamic nature of healthcare teams. An increasing number of publications have proposed auditing strategies for EMRs,45,46 however, this line of work is unique in that it suggests EMR users can be ‘explained’ by the diagnoses that are assigned to the patient records they access. This work suggests that data-driven auditing strategies may help winnow the set of accesses to patient records to a manageable size for review by administrative officials (eg, privacy officers) of HCOs. The next set of papers in the technical section focus on various strategies that can be invoked to protect patients' privacy when data are shared for secondary use. However, before presenting specific protection methodologies, this section begins with an illustration of types of research studies that can be enabled through de-identified data. The paper by White and Horvitz integrates web search data from Bing and geocoded data from mobile devices to show how search for health information online correlates with an individual's physical presence at a healthcare providing facility.47 This research is performed on data that are stripped of user identifiers and location information prior to the analysis. However, there may be times when it is beneficial to link a patient's record across multiple healthcare institutions, or within a single institution. To support such efforts without revealing a patient's identity, there has been a flurry of research in private record linkage48–51 (or entity resolution). Such linkage is increasingly based on hashed versions of patient identifiers (eg, personal names) or quasi-identifiers (eg, demographics). The paper from Cassa, Miller, and Mandl suggests a protocol to derive a secure fingerprint from genomic data.52 They indicate how this approach may be applied to track a patient's record across the research enterprise in place of explicitly identifying information. The final set of papers focus on strategies for de-identifying various types of health information. Biomedical data can take a wide array of forms, ranging from free text (eg, natural language clinical notes) to structured information (eg, such as discharge databases) to high-dimensional information (eg, genome-wide scans of single nucleotide polymorphisms). The majority of health information is in free text form and so a significant amount of research53,54 over the past several years has investigated how to detect and redact a prespecified set of potential identifiers (such as the list of 18 features in the HIPAA Safe Harbor de-identification standard). The paper by Ferrandez and colleagues provides an example of how rule-based (eg, dictionaries, regular expressions, and rules) and machine (eg, random support and can be to construct a free text for over different types of Health clinical The paper by and colleagues how machine text de-identification has impact on clinical concept in the form of from over types from the Although the previous and in the illustrate how identifiers in clinical text can be information in the text may still or of the patient (or their As have to de-identification strategies that are more in their strategies are often applied to more data such as database The paper by and colleagues how a patient's of results can be unique and used as a to track a patient back to their To this they a which to the An analysis patients' records that such it highly that a patient's record be to a group of less than individuals while minimal on clinical Although some this of addition does not an In this regard, a significant number of publications have (ie, strategies be applied to ensure that record corresponds to at least patients (ie, the it has been that de-identification strategies based on the of a prespecified list of features or more strategies may not be an appropriate of privacy protection because they can about the patients from the data were While this general notion has been models have been In the second invited paper, and describe the notion of privacy from a theoretical In this of are permitted to of a which with a (eg, a of may be reported as This is such that it is that the determine a specific individual to the database within a certain and that the is within a certain of the This has a number of important but also a number of and practical to in healthcare The final paper of the special issue, by and colleagues, provides a more practical on how privacy be applied to of health They how this privacy protection be applied to a from but there are challenges to this approach to high-dimensional data. As this special issue the of data privacy in the biomedical domain is and It and technical boundaries and is specialized to the of data and process being it is not to review the field in this As we this to a we that topics (eg, access consent disclosure and policy to manage health information were not but are no less important than those reported on in this At the same time, we that new and technologies are new challenges to privacy that the biomedical will need to in the not technology that we to highlight is As and the amount of data by healthcare it is increasingly the case the health information is being in systems beyond the control and oversight of and in with different privacy laws and this issue demonstrates that appropriate protections can be defined for emerging systems and that there is a working on and are that research in this area will lead to new appropriate that balance privacy and data and system Bradley A. Malin, Khaled El Emam, Christine M. O'Keefe |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Auditing Medical Records Accesses via Healthcare Interaction Networks
You Chen 0001, Steve Nyemba, Bradley A. Malin |
AMIA | 3 |
| 2012 | The Chicago Health Atlas: A Public Resource to Visualize Health Conditions and Resources in Chicago
Abel N. Kho, John P. Cashy, Bala Hota, Shannon A. Sims, Bradley A. Malin, David O. Meltzer, Erin O. Kaleba, William L. Galanter |
AMIA | 5 |
| 2012 | Rethinking the "Honest Broker" in the Changing Face of Security and Privacy
Luke V. Rasmussen, Brian D. Athey, Andrew D. Boyd, Bradley A. Malin, Shawn N. Murphy |
AMIA | 4 |
| 2012 | Detecting Anomalous User Behaviors in Workflow-Driven Web ApplicationsabstractWeb applications are increasingly used as portals to interact with back-end database systems and support business processes. This type of data-centric workflow-driven web application is vulnerable to two types of security threats. The first is an request integrity attack, which stems from the vulnerabilities in the implementation of business logic within web applications. The second is guideline violation, which stems from privilege misuse in scenarios where business logic and policies are too complex to be accurately defined and enforced. Both threats can lead to sequences of web requests that deviate from typical user behaviors. The objective of this paper is to detect anomalous user behaviors based on the sequence of their requests within a web session. We first decompose web sessions into workflows based on their data objects. In doing so, the detection of anomalous sessions is reduced to detection of anomalous workflows. Next, we apply a hidden Markov model (HMM) to characterize workflows on a per-object basis. In this model, the implicit business logic involved in this object defines the unobserved states of the Markov process, where the web requests are observations. To derive more robust HMMs, we extend the object-specific approach to an object-cluster approach, where objects with similar workflows are clustered and HMM models are derived on a per-cluster basis. We evaluate our models using two real systems, including an open source web application and a large web-based electronic medical record system. The results show that our approach can detect anomalous web sessions and lend evidence to suggest that the clustering approach can achieve relatively low false positive rates while maintaining its detection accuracy. Xiaowei Li 0003, Yuan Xue 0001, Bradley A. Malin |
SRDS | 3 |
| 2012 | Detecting Anomalous Insiders in Collaborative Information SystemsabstractCollaborative information systems (CISs) are deployed within a diverse array of environments that manage sensitive information. Current security mechanisms detect insider threats, but they are ill-suited to monitor systems in which users function in dynamic teams. In this paper, we introduce the community anomaly detection system (CADS), an unsupervised learning framework to detect insider threats based on the access logs of collaborative environments. The framework is based on the observation that typical CIS users tend to form community structures based on the subjects accessed (e.g., patients' records viewed by healthcare providers). CADS consists of two components: 1) relational pattern extraction, which derives community structures and 2) anomaly prediction, which leverages a statistical model to determine when users have sufficiently deviated from communities. We further extend CADS into MetaCADS to account for the semantics of subjects (e.g., patients' diagnoses). To empirically evaluate the framework, we perform an assessment with three months of access logs from a real electronic health record (EHR) system in a large medical center. The results illustrate our models exhibit significant performance gains over state-of-the-art competitors. When the number of illicit users is low, MetaCADS is the best model, but as the number grows, commonly accessed semantics lead to hiding in a crowd, such that CADS is more prudent. You Chen 0001, Steve Nyemba, Bradley A. Malin |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2012 | Secure Management of Biomedical Data With Cryptographic HardwareabstractThe biomedical community is increasingly migrating toward research endeavors that are dependent on large quantities of genomic and clinical data. At the same time, various regulations require that such data be shared beyond the initial collecting organization (e.g., an academic medical center). It is of critical importance to ensure that when such data are shared, as well as managed, it is done so in a manner that upholds the privacy of the corresponding individuals and the overall security of the system. In general, organizations have attempted to achieve these goals through deidentification methods that remove explicitly, and potentially, identifying features (e.g., names, dates, and geocodes). However, a growing number of studies demonstrate that deidentified data can be reidentified to named individuals using simple automated methods. As an alternative, it was shown that biomedical data could be shared, managed, and analyzed through practical cryptographic protocols without revealing the contents of any particular record. Yet, such protocols required the inclusion of multiple third parties, which may not always be feasible in the context of trust or bandwidth constraints. Thus, in this paper, we introduce a framework that removes the need for multiple third parties by collocating services to store and to process sensitive biomedical data through the integration of cryptographic hardware. Within this framework, we define a secure protocol to process genomic data and perform a series of experiments to demonstrate that such an approach can be run in an efficient manner for typical biomedical investigations. Mustafa Canim, Murat Kantarcioglu, Bradley A. Malin |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2012 | Anonymization of Longitudinal Electronic Medical RecordsabstractElectronic medical record (EMR) systems have enabled healthcare providers to collect detailed patient information from the primary care domain. At the same time, longitudinal data from EMRs are increasingly combined with biorepositories to generate personalized clinical decision support protocols. Emerging policies encourage investigators to disseminate such data in a deidentified form for reuse and collaboration, but organizations are hesitant to do so because they fear such actions will jeopardize patient privacy. In particular, there are concerns that residual demographic and clinical features could be exploited for reidentification purposes. Various approaches have been developed to anonymize clinical data, but they neglect temporal information and are, thus, insufficient for emerging biomedical research paradigms. This paper proposes a novel approach to share patient-specific longitudinal data that offers robust privacy guarantees, while preserving data utility for many biomedical investigations. Our approach aggregates temporal and diagnostic information using heuristics inspired from sequence alignment and clustering methods. We demonstrate that the proposed approach can generate anonymized data that permit effective biomedical analysis using several patient cohorts derived from the EMR system of the Vanderbilt University Medical Center. Acar Tamersoy, Grigorios Loukides, Mehmet Ercan Nergiz, Yücel Saygin, Bradley A. Malin |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2011 | Detection of anomalous insiders in collaborative environments via relational analysis of access logsabstractCollaborative information systems (CIS) are deployed within a diverse array of environments, ranging from the Internet to intelligence agencies to healthcare. It is increasingly the case that such systems are applied to manage sensitive information, making them targets for malicious insiders. While sophisticated security mechanisms have been developed to detect insider threats in various file systems, they are neither designed to model nor to monitor collaborative environments in which users function in dynamic teams with complex behavior. In this paper, we introduce a community-based anomaly detection system (CADS), an unsupervised learning framework to detect insider threats based on information recorded in the access logs of collaborative environments. CADS is based on the observation that typical users tend to form community structures, such that users with low affinity to such communities are indicative of anomalous and potentially illicit behavior. The model consists of two primary components: relational pattern extraction and anomaly detection. For relational pattern extraction, CADS infers community structures from CIS access logs, and subsequently derives communities, which serve as the CADS pattern core. CADS then uses a formal statistical model to measure the deviation of users from the inferred communities to predict which users are anomalies. To empirically evaluate the threat detection model, we perform an analysis with six months of access logs from a real electronic health record system in a large medical center, as well as a publicly available dataset for replication purposes. The results illustrate that CADS can distinguish simulated anomalous users in the context of real user behavior with a high degree of certainty and with significant performance gains in comparison to several competing anomaly detection models. You Chen 0001, Bradley A. Malin |
CODASPY | 2 |
| 2011 | Leveraging social networks to detect anomalous insider actions in collaborative environmentsabstractCollaborative information systems (CIS) enable users to coordinate efficiently over shared tasks. T hey are often deployed in complex dynamic systems that provide users with broad access privileges, but also leave the system vulnerable to various attacks. Techniques to detect threats originating from beyond the system are relatively mature, but methods to detect insider threats are still evolving. A promising class of insider threat detection models for CIS focus on the communities that manifest between users based on the usage of common subjects in the system. However, current methods detect only when a user's aggregate behavior is intruding, not when specific actions have deviated from expectation. In this paper, we introduce a method called specialized network anomaly detection (SNAD) to detect such events. SNAD assembles the community of users that access a particular subject and assesses if similarities of the community with and without a certain user are sufficiently different. We present a theoretical basis and perform an extensive empirical evaluation with the access logs of two distinct environments: those of a large electronic health record system (6,015 users, 130,457 patients and 1,327,500 accesses) and the editing logs of Wikipedia (2,388,955 revisors, 55,200 articles and 6,482,780 revisions). We compare SNAD with several competing methods and demonstrate it is significantly more effective: on average it achieves 20-30% greater area under an ROC curve. You Chen 0001, Steve Nyemba, Bradley A. Malin |
ISI | 4 |
| 2011 | A Constraint Satisfaction Cryptanalysis of Bloom Filters in Private Record Linkage
Mehmet Kuzu, Murat Kantarcioglu, Elizabeth Durham, Bradley A. Malin |
PETS | 4 |
| 2011 | Towards understanding the usage pattern of web-based electronic medical record systemsabstractThe benefits and importance of electronic medical record (EMR) systems have been well recognized in the healthcare industry. Yet, their wide adoption still face significant barriers in providing on-demand secure medical information access while preserving patients' privacy. Understanding the usage pattern of an EMR system is the first essential step towards building such environment. This paper conducts an in-depth trace analysis of a large-scale EMR system that has been in operation for more than a decade at the Vanderbilt Medical Center. Our study demonstrates several important characteristics of EMR system usage from the perspective of user-initiated sessions. First, the workload of the EMR system is highly stable and consistent with a weekly pattern. Second, EMR behavior varies between users, but each user's behavior tends to be consistent with a slow rate of migration across sessions. Finally, the degree of access between users and medical records is sparse, echoing the limits of patient-caregiver relationships that manifest in real healthcare operations. We believe these observations can assist in the development of system security measures, such as EMR-specific anomaly detection systems, and facilitate system performance optimization. Xiaowei Li 0003, Yuan Xue 0001, Bradley A. Malin |
WOWMOM | 3 |
| 2011 | An entropy approach to disclosure risk assessment: Lessons from real applications and simulated domains
Edoardo M. Airoldi, Xue Bai 0003, Bradley A. Malin |
Decis. Support Syst. | 3 |
| 2011 | A secure protocol for protecting the identity of providers when disclosing data for disease surveillanceabstractBACKGROUND: Providers have been reluctant to disclose patient data for public-health purposes. Even if patient privacy is ensured, the desire to protect provider confidentiality has been an important driver of this reluctance. METHODS: Six requirements for a surveillance protocol were defined that satisfy the confidentiality needs of providers and ensure utility to public health. The authors developed a secure multi-party computation protocol using the Paillier cryptosystem to allow the disclosure of stratified case counts and denominators to meet these requirements. The authors evaluated the protocol in a simulated environment on its computation performance and ability to detect disease outbreak clusters. RESULTS: Theoretical and empirical assessments demonstrate that all requirements are met by the protocol. A system implementing the protocol scales linearly in terms of computation time as the number of providers is increased. The absolute time to perform the computations was 12.5 s for data from 3000 practices. This is acceptable performance, given that the reporting would normally be done at 24 h intervals. The accuracy of detection disease outbreak cluster was unchanged compared with a non-secure distributed surveillance protocol, with an F-score higher than 0.92 for outbreaks involving 500 or more cases. CONCLUSION: The protocol and associated software provide a practical method for providers to disclose patient data for sentinel, syndromic or other indicator-based surveillance while protecting patient privacy and the identity of individual providers. Khaled El Emam, Jay Mercer, Liam Peyton, Murat Kantarcioglu, Bradley A. Malin, David L. Buckeridge, Saeed Samet, Craig Earle |
J. Am. Medical Informatics Assoc. | 6 |
| 2011 | Never too old for anonymity: a statistical standard for demographic data sharing via the HIPAA Privacy RuleabstractOBJECTIVE: Healthcare organizations must de-identify patient records before sharing data. Many organizations rely on the Safe Harbor Standard of the HIPAA Privacy Rule, which enumerates 18 identifiers that must be suppressed (eg, ages over 89). An alternative model in the Privacy Rule, known as the Statistical Standard, can facilitate the sharing of more detailed data, but is rarely applied because of a lack of published methodologies. The authors propose an intuitive approach to de-identifying patient demographics in accordance with the Statistical Standard. DESIGN: The authors conduct an analysis of the demographics of patient cohorts in five medical centers developed for the NIH-sponsored Electronic Medical Records and Genomics network, with respect to the US census. They report the re-identification risk of patient demographics disclosed according to the Safe Harbor policy and the relative risk rate for sharing such information via alternative policies. MEASUREMENTS: The re-identification risk of Safe Harbor demographics ranged from 0.01% to 0.19%. The findings show alternative de-identification models can be created with risks no greater than Safe Harbor. The authors illustrate that the disclosure of patient ages over the age of 89 is possible when other features are reduced in granularity. LIMITATIONS: The de-identification approach described in this paper was evaluated with demographic data only and should be evaluated with other potential identifiers. CONCLUSION: Alternative de-identification policies to the Safe Harbor model can be derived for patient demographics to enable the disclosure of values that were previously suppressed. The method is generalizable to any environment in which population statistics are available. Bradley A. Malin, Kathleen Benitez, Daniel R. Masys |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Learning relational policies from electronic health record access logs
Bradley A. Malin, Steve Nyemba, John Paulett |
J. Biomed. Informatics | 1 |
| 2011 | COAT: COnstraint-based anonymization of transactions
Grigorios Loukides, Aris Gkoulalas-Divanis, Bradley A. Malin |
Knowl. Inf. Syst. | 3 |
| 2010 | Secure construction of k-unlinkable patient records from distributed providers
Bradley A. Malin |
Artif. Intell. Medicine | 1 |
| 2010 | Evaluating re-identification risks with respect to the HIPAA privacy ruleabstractOBJECTIVE: Many healthcare organizations follow data protection policies that specify which patient identifiers must be suppressed to share "de-identified" records. Such policies, however, are often applied without knowledge of the risk of "re-identification". The goals of this work are: (1) to estimate re-identification risk for data sharing policies of the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule; and (2) to evaluate the risk of a specific re-identification attack using voter registration lists. MEASUREMENTS: We define several risk metrics: (1) expected number of re-identifications; (2) estimated proportion of a population in a group of size g or less, and (3) monetary cost per re-identification. For each US state, we estimate the risk posed to hypothetical datasets, protected by the HIPAA Safe Harbor and Limited Dataset policies by an attacker with full knowledge of patient identifiers and with limited knowledge in the form of voter registries. RESULTS: The percentage of a state's population estimated to be vulnerable to unique re-identification (ie, g=1) when protected via Safe Harbor and Limited Datasets ranges from 0.01% to 0.25% and 10% to 60%, respectively. In the voter attack, this number drops for many states, and for some states is 0%, due to the variable availability of voter registries in the real world. We also find that re-identification cost ranges from $0 to $17,000, further confirming risk variability. CONCLUSIONS: This work illustrates that blanket protection policies, such as Safe Harbor, leave different organizations vulnerable to re-identification at different rates. It provides justification for locally performed re-identification risk estimates prior to sharing data. Kathleen Benitez, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 2 |
| 2010 | The disclosure of diagnosis codes can breach research participants' privacyabstractOBJECTIVE: De-identified clinical data in standardized form (eg, diagnosis codes), derived from electronic medical records, are increasingly combined with research data (eg, DNA sequences) and disseminated to enable scientific investigations. This study examines whether released data can be linked with identified clinical records that are accessible via various resources to jeopardize patients' anonymity, and the ability of popular privacy protection methodologies to prevent such an attack. DESIGN: The study experimentally evaluates the re-identification risk of a de-identified sample of Vanderbilt's patient records involved in a genome-wide association study. It also measures the level of protection from re-identification, and data utility, provided by suppression and generalization. MEASUREMENT: Privacy protection is quantified using the probability of re-identifying a patient in a larger population through diagnosis codes. Data utility is measured at a dataset level, using the percentage of retained information, as well as its description, and at a patient level, using two metrics based on the difference between the distribution of Internal Classification of Disease (ICD) version 9 codes before and after applying privacy protection. RESULTS: More than 96% of 2800 patients' records are shown to be uniquely identified by their diagnosis codes with respect to a population of 1.2 million patients. Generalization is shown to reduce further the percentage of de-identified records by less than 2%, and over 99% of the three-digit ICD-9 codes need to be suppressed to prevent re-identification. CONCLUSIONS: Popular privacy protection methods are inadequate to deliver a sufficiently protected and useful result when sharing data derived from complex clinical systems. The development of alternative privacy protection models is thus required. Grigorios Loukides, Joshua C. Denny, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 3 |
| 2010 | Effects of personal identifier resynthesis on clinical text de-identificationabstractOBJECTIVE: De-identified medical records are critical to biomedical research. Text de-identification software exists, including "resynthesis" components that replace real identifiers with synthetic identifiers. The goal of this research is to evaluate the effectiveness and examine possible bias introduced by resynthesis on de-identification software. DESIGN: We evaluated the open-source MITRE Identification Scrubber Toolkit, which includes a resynthesis capability, with clinical text from Vanderbilt University Medical Center patient records. We investigated four record classes from over 500 patients' files, including laboratory reports, medication orders, discharge summaries and clinical notes. We trained and tested the de-identification tool on real and resynthesized records. MEASUREMENTS: We measured performance in terms of precision, recall, F-measure and accuracy for the detection of protected health identifiers as designated by the HIPAA Safe Harbor Rule. RESULTS: The de-identification tool was trained and tested on a collection of real and resynthesized Vanderbilt records. Results for training and testing on the real records were 0.990 accuracy and 0.960 F-measure. The results improved when trained and tested on resynthesized records with 0.998 accuracy and 0.980 F-measure but deteriorated moderately when trained on real records and tested on resynthesized records with 0.989 accuracy 0.862 F-measure. Moreover, the results declined significantly when trained on resynthesized records and tested on real records with 0.942 accuracy and 0.728 F-measure. CONCLUSION: The de-identification tool achieves high accuracy when training and test sets are homogeneous (ie, both real or resynthesized records). The resynthesis component regularizes the data to make them less "realistic," resulting in loss of performance particularly when training on resynthesized data and testing on real data. Reyyan Yeniterzi, John S. Aberdeen, Samuel Bayer, Ben Wellner, Lynette Hirschman, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 6 |
| 2009 | Formal anonymity models for efficient privacy-preserving joins
Murat Kantarcioglu, Ali Inan, Wei Jiang 0026, Bradley A. Malin |
Data Knowl. Eng. | 4 |
| 2008 | A Privacy-Preserving Framework for Integrating Person-Specific Databases
Murat Kantarcioglu, Wei Jiang 0026, Bradley A. Malin |
Privacy in Statistical Databases | 3 |
| 2008 | Recent advances in preserving privacy when mining data
Francesco Bonchi, Bradley A. Malin, Yücel Saygin |
Data Knowl. Eng. | 2 |
| 2008 | k-Unlinkability: A privacy protection model for distributed data
Bradley A. Malin |
Data Knowl. Eng. | 1 |
| 2008 | A Cryptographic Approach to Securely Share and Query Genomic SequencesabstractTo support large-scale biomedical research projects, organizations need to share person-specific genomic sequences without violating the privacy of their data subjects. In the past, organizations protected subjects' identities by removing identifiers, such as name and social security number; however, recent investigations illustrate that deidentified genomic data can be "reidentified" to named individuals using simple automated methods. In this paper, we present a novel cryptographic framework that enables organizations to support genomic data mining without disclosing the raw genomic sequences. Organizations contribute encrypted genomic sequence records into a centralized repository, where the administrator can perform queries, such as frequency counts, without decrypting the data. We evaluate the efficiency of our framework with existing databases of single nucleotide polymorphism (SNP) sequences and demonstrate that the time needed to complete count queries is feasible for real world applications. For example, our experiments indicate that a count query over 40 SNPs in a database of 5000 records can be completed in approximately 30 min with off-the-shelf technology. We further show that approximation strategies can be applied to significantly speed up query execution times with minimal loss in accuracy. The framework can be implemented on top of existing information and network technologies in biomedical environments. Murat Kantarcioglu, Wei Jiang 0026, Ying Liu 0007, Bradley A. Malin |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2007 | A Modeling Environment for Patient Portals
Sean Duncavage, Janos L. Mathe, Jan Werner, Bradley A. Malin, Ákos Lédeczi, Janos Sztipanovits |
AMIA | 4 |
| 2007 | A computational model to protect patient data from location-based re-identification
Bradley A. Malin |
Artif. Intell. Medicine | 1 |
| 2007 | Research Paper: A Longitudinal Social Network Analysis of the Editorial Boards of Medical Informatics and Bioinformatics JournalsabstractOBJECTIVE: The goal of this research is to learn how the editorial staffs of bioinformatics and medical informatics journals provide support for cross-community exposure. Models such as co-citation and co-author analysis measure the relationships between researchers; but they do not capture how environments that support knowledge transfer across communities are organized. METHODS: In this paper, we propose a social network analysis model to study how editorial boards integrate researchers from disparate communities. We evaluate our model by building relational networks based on the editorial boards of approximately 40 journals that serve as research outlets in medical informatics and bioinformatics. We track the evolution of editorial relationships through a longitudinal investigation over the years 2000 through 2005. RESULTS: Our findings suggest that there are research journals that support the collocation of editorial board members from the bioinformatics and medical informatics communities. Network centrality metrics indicate that editorial board members are located in the intersection of the communities and that the number of individuals in the intersection is growing with time. CONCLUSIONS: Social network analysis methods provide insight into the relationships between the medical informatics and bioinformatics communities. The number of editorial board members facilitating the publication intersection of the communities has grown, but the intersection remains dependent on a small group of individuals and fragile. Bradley A. Malin, Kathleen M. Carley |
J. Am. Medical Informatics Assoc. | 1 |
| 2006 | Re-identification of Familial Database Records
Bradley A. Malin |
AMIA | 1 |
| 2006 | Composition and Disclosure of Unlinkable Distributed DatabasesabstractAn individual’s location-visit pattern, or trail, can be leveraged to link sensitive data back to identity. We propose a secure multiparty computation protocol that enables locations to provably prevent such linkages. The protocol incorporates a controllable parameter specifying the minimum number of identities a sensitive piece of data must be linkable to via its trail. Bradley A. Malin, Latanya Sweeney |
ICDE | 1 |
| 2005 | A Secure Protocol to Distribute Unlinkable Health Data
Bradley A. Malin, Latanya Sweeney |
AMIA | 1 |
| 2005 | Configurable Security Protocols for Multi-party Data Analysis with Malicious ParticipantsabstractStandard multi-party computation models assume semi-honest behavior, where the majority of participants implement protocols according to specification, an assumption not always plausible. In this paper we introduce a multi-party protocol for collaborative data analysis when participants are malicious and fail to follow specification. The protocol incorporates a semi-trusted third party, which analyzes encrypted data and provides honest responses that only intended recipients can successfully decrypt. The protocol incorporates data confidentiality by enabling participants to receive encrypted responses tailored to their own encrypted data submissions without revealing plaintext to other participants, including the third party. As opposed to previous models, trust need only be placed on a single participant with no data at stake. Additionally, the proposed protocol is configurable in a way that security features are controlled by independent subprotocols. Various combinations of subprotocols allow for a flexible security system, appropriate for a number of distributed data applications, such as secure list comparison. Bradley A. Malin, Edoardo M. Airoldi, Samuel Edoho-Eket |
ICDE | 1 |
| 2005 | Technical Evaluation: An Evaluation of the Current State of Genomic Data Privacy Protection Technology and a Roadmap for the FutureabstractThe incorporation of genomic data into personal medical records poses many challenges to patient privacy. In response, various systems for preserving patient privacy in shared genomic data have been developed and deployed. Although these systems de-identify the data by removing explicit identifiers (e.g., name, address, or Social Security number) and incorporate sound security design principles, they suffer from a lack of formal modeling of inferences learnable from shared data. This report evaluates the extent to which current protection systems are capable of withstanding a range of re-identification methods, including genotype-phenotype inferences, location-visit patterns, family structures, and dictionary attacks. For a comparative re-identification analysis, the systems are mapped to a common formalism. Although there is variation in susceptibility, each system is deficient in its protection capacity. The author discovers patterns of protection failure and discusses several of the reasons why these systems are susceptible. The analyses and discussion within provide guideposts for the development of next-generation protection methods amenable to formal proofs. Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2005 | Preserving Privacy by De-Identifying Face ImagesabstractIn the context of sharing video surveillance data, a significant threat to privacy is face recognition software, which can automatically identify known people, such as from a database of drivers' license photos, and thereby track people regardless of suspicion. This paper introduces an algorithm to protect the privacy of individuals in video surveillance data by deidentifying faces such that many facial characteristics remain but the face cannot be reliably recognized. A trivial solution to deidentifying faces involves blacking out each face. This thwarts any possible face recognition, but because all facial details are obscured, the result is of limited use. Many ad hoc attempts, such as covering eyes, fail to thwart face recognition because of the robustness of face recognition methods. This work presents a new privacy-enabling algorithm, named k-Same, that guarantees face recognition software cannot reliably recognize deidentified faces, even though many facial details are preserved. The algorithm determines similarity between faces based on a distance metric and creates new faces by averaging image components, which may be the original image pixels (k-Same-Pixel) or eigenvectors (k-Same-Eigen). Results are presented on a standard collection of real face images with varying k. Elaine M. Newton, Latanya Sweeney, Bradley A. Malin |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2004 | How (not) to protect genomic data privacy in a distributed network: using trail re-identification to evaluate and design anonymity protection systems
Bradley A. Malin, Latanya Sweeney |
J. Biomed. Informatics | 1 |
| 2002 | Correlating web usage of health information with patient medical data
Bradley A. Malin |
AMIA | 1 |
| 2001 | Re-identification of DNA through an automated linkage process
Bradley A. Malin, Latanya Sweeney |
AMIA | 1 |
| 2000 | Determining the identifiability of DNA database entries
Bradley A. Malin, Latanya Sweeney |
AMIA | 1 |