EDBT 2026 Demo / reviewers in the wild / expert
Daniel Fabbri
dblp:17/7449
· DBLP profile ↗
34ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-0530-2510ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-authorSecurity and privacy · 2Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Secondary use of radiological imaging data: Vanderbilt's ImageVU approachabstractOBJECTIVE: To develop ImageVU, a scalable research imaging infrastructure that integrates clinical imaging data with metadata-driven cohort discovery, enabling secure, efficient, and regulatory-compliant access to imaging for secondary and opportunistic research use. This manuscript presents a detailed description of ImageVU's key components and lessons learned to assist other institutions in developing similar research imaging services and infrastructure. METHODS: ImageVU was designed to support the secondary use of radiological imaging data through a dedicated research imaging store. The system comprises four interconnected components: a Research PACS, an Ad Hoc Backfill Host, Cloud Storage System, and a De-Identification System. Imaging metadata are extracted and stored in the Research Derivative (RD), an identified clinical data repository, and the Synthetic Derivative (SD), a de-identified research data repository, with access facilitated through the RD Discover web portal. Researchers interact with the system via structured metadata queries and multiple data delivery options, including web-based viewing, bulk downloads, and dataset preparation for high-performance computing environments. RESULTS: The integration of metadata-driven search capabilities has streamlined cohort discovery and improved imaging data accessibility. As of December 2024, ImageVU has processed 12.9 million MRI and CT series from 1.36 million studies across 453,403 patients. The system has supported 75 project requests, delivering over 50 TB of imaging data to 55 investigators, leading to 66 published research papers. CONCLUSION: ImageVU demonstrates a scalable and efficient approach for integrating clinical imaging into research workflows. By combining institutional data infrastructure with cloud-based storage and metadata-driven cohort identification, the platform enables secure and compliant access to imaging for translational research. David S. Smith, Karthik Ramadass, Laura M. Jones, Jennifer Morse, Daniel Fabbri, Joseph R. Coco, Shunxing Bao, Melissa A. Basford, Peter J. Embí, Reed A. Omary, John C. Gore, Jill M. Pulley, Bennett A. Landman |
J. Biomed. Informatics | 5 |
| 2022 | Semi-Automated Data Curation from Biomedical Literature
Protiva Rahman, Daniel Fabbri |
AMIA | 2 |
| 2022 | Accelerated Data Curation of Colitis Cases
Protiva Rahman, Cheng Ye 0001, Kate Mittendorf, Michele L. LeNoue-Newton, Christine Micheel, Daniel Fabbri |
AMIA | 6 |
| 2022 | Classifying Infection Risk Following Pediatric Cardiac Surgery
Kaitlin C. Williamson, Daniel Fabbri |
AMIA | 2 |
| 2021 | Predicting Motor Responsiveness to Deep Brain Stimulation with Machine Learning
Kevin J. Krause, Fenna Phibbs, Thomas Davis, Daniel Fabbri |
AMIA | 4 |
| 2021 | Automatic Data Curation from Unstructured Text
Protiva Rahman, Daniel Fabbri |
AMIA | 2 |
| 2021 | Classifying Infection Risk Following Pediatric Cardiac Surgery
Kaitlin C. Williamson, Daniel Fabbri |
AMIA | 2 |
| 2021 | Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort
Sarah DeLozier, Sarah Bland, Melissa McPheeters, Quinn Stanton Wells, Eric Farber-Eger, Cosmin Adrian Bejan, Daniel Fabbri, S. Trent Rosenbloom, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei, Josh F. Peterson, Lisa Bastarache |
J. Biomed. Informatics | 7 |
| 2020 | To Warn or Not to Warn: Online Signaling in Audit GamesabstractRoutine operational use of sensitive data is often governed by law and regulation. For instance, in the medical domain, there are various statues at the state and federal level that dictate who is permitted to work with patients' records and under what conditions. To screen for potential privacy breaches, logging systems are usually deployed to trigger alerts whenever a suspicious access is detected. However, such mechanisms are often inefficient because 1) the vast majority of triggered alerts are false positives, 2) small budgets make it unlikely that a real attack will be detected, and 3) attackers can behave strategically, such that traditional auditing mechanisms cannot easily catch them. To improve efficiency, information systems may invoke signaling, so that whenever a suspicious access request occurs, the system can, in real time, warn the user that the access may be audited. Then, at the close of a finite period, a selected subset of suspicious accesses are audited. This gives rise to an online problem in which one needs to determine 1) whether a warning should be triggered and 2) the likelihood that the data request event will be audited. In this paper, we formalize this auditing problem as a Signaling Audit Game (SAG), in which we model the interactions between an auditor and an attacker in the context of signaling and the usability cost is represented as a factor of the auditor's payoff. We study the properties of its Stackelberg equilibria and develop a scalable approach to compute its solution. We show that a strategic presentation of warnings adds value in that SAGs realize significantly higher utility for the auditor than systems without signaling. We perform a series of experiments with 10 million real access events, containing over 26K alerts, from a large academic medical center to illustrate the value of the proposed auditing model and the consistency of its advantages over existing baseline methods. Chao Yan 0004, Yevgeniy Vorobeychik, Bo Li 0026, Daniel Fabbri, Bradley A. Malin |
ICDE | 5 |
| 2019 | Feasibility Assessment of a Pre-Hospital Automated Sensing Clinical Documentation System
Sean M. Bloos, Candace D. McNaughton, Joseph R. Coco, Laurie L. Novak, Julie A. Adams, Bobby Bodenheimer, Jesse M. Ehrenfeld, Jamison Heard, Richard A. Paris, Christopher L. Simpson, Deirdre Scully, Daniel Fabbri |
AMIA | 12 |
| 2019 | Database Audit Workload Prioritization via Game TheoryabstractThe quantity of personal data that is collected, stored, and subsequently processed continues to grow rapidly. Given its sensitivity, ensuring privacy protections has become a necessary component of database management. To enhance protection, a number of mechanisms have been developed, such as audit logging and alert triggers, which notify administrators about suspicious activities. However, this approach is limited. First, the volume of alerts is often substantially greater than the auditing capabilities of organizations. Second, strategic attackers can attempt to disguise their actions or carefully choose targets, thus hide illicit activities. In this article, we introduce an auditing approach that accounts for adversarial behavior by (1) prioritizing the order in which types of alerts are investigated and (2) providing an upper bound on how much resource to allocate for each type. Specifically, we model the interaction between a database auditor and attackers as a Stackelberg game. We show that even a highly constrained version of such problem is NP-Hard. Then, we introduce a method that combines linear programming, column generation, and heuristic searching to derive an auditing policy. On the synthetic data, we perform an extensive evaluation on the approximation degree of our solution with the optimal one. The two real datasets, (1) 1.5 months of audit logs from Vanderbilt University Medical Center and (2) a publicly available credit card application dataset, are used to test the policy-searching performance. The findings demonstrate the effectiveness of the proposed methods for searching the audit strategies, and our general approach significantly outperforms non-game-theoretic baselines. Chao Yan 0004, Bo Li 0026, Yevgeniy Vorobeychik, Aron Laszka, Daniel Fabbri, Bradley A. Malin |
ACM Trans. Priv. Secur. | 5 |
| 2018 | Crowdsourcing Clinical Chart Reviews
Joseph R. Coco, Cheng Ye 0001, Chen Hajaj, Yevgeniy Vorobeychik, Joshua C. Denny, Laurie L. Novak, Bradley A. Malin, Thomas A. Lasko, Daniel Fabbri |
AMIA | 9 |
| 2018 | What Do EHR Access Logs Tell Us About Workflow Patterns?
Ioana Danciu, Stuart Weinberg, Daniel Fabbri, Kim M. Unertl |
AMIA | 3 |
| 2018 | Get Your Workload in Order: Game Theoretic Prioritization of Database AuditingabstractA wide variety of mechanisms, such as alert triggers and auditing routines, have been developed to notify administrators about types of suspicious activities in the daily use of large databases of personal and sensitive information. However, such mechanisms are limited in that: 1) the volume of such alerts is often substantially greater than the auditing capabilities of budget-constrained organizations and 2) strategic attackers may disguise their actions or carefully choose which records they touch, thus evading auditing routines. To address these problems, we introduce a novel approach to database auditing that explicitly accounts for adversarial behavior by 1) prioritizing the order in which types of alerts are investigated and 2) providing an upper bound on how much budget to allocate for auditing each alert type. We model the interaction between a database auditor and potential attackers as a Stackelberg game in which the auditor chooses an auditing policy and attackers choose which records in a database to target. We further introduce an efficient approach that combines linear programming, column generation, and heuristic search to derive an auditing policy, in the form of a mixed strategy. We assess the performance of the policy selection method using a publicly available credit card application dataset, the results of which indicate that our method produces high-quality database audit policies, significantly outperforming baselines that are not based in a game theoretic framing. Chao Yan 0004, Bo Li 0026, Yevgeniy Vorobeychik, Aron Laszka, Daniel Fabbri, Bradley A. Malin |
ICDE | 5 |
| 2018 | The therapy is making me sick: how online portal communications between breast cancer patients and physicians indicate medication discontinuationabstractObjective: Online platforms have created a variety of opportunities for breast patients to discuss their hormonal therapy, a long-term adjuvant treatment to reduce the chance of breast cancer occurrence and mortality. The goal of this investigation is to ascertain the extent to which the messages breast cancer patients communicated through an online portal can indicate their potential for discontinuing hormonal therapy. Materials and Methods: We studied the de-identified electronic medical records of 1106 breast cancer patients who were prescribed hormonal therapy at Vanderbilt University Medical Center over a 12-year period. We designed a data-driven approach to investigate patients' patterns of messaging with healthcare providers, the topics they communicated, and the extent to which these messaging behaviors associate with the likelihood that a patient will discontinue a prescribed 5-year regimen of therapy. Results: The results indicates that messaging rate over time [hazard ratio (HR) = 1.373, P = 0.002], mentions of side effects (HR = 1.214, P = 0.006), and surgery-related topics (HR = 1.170, P = 0.034) were associated with increased risk of early medication discontinuation. In contrast, seeking professional suggestions (HR = 0.766, P = 0.002), expressing gratitude to healthcare providers (HR = 0.872, P = 0.044), and mentions of drugs used to treat side effects (HR = 0.807, P = 0.013) were associated with decreased risk of medication discontinuation. Discussion and Conclusion: This investigation suggests that patient-generated content can inform the study of health-related behaviors. Given that approximately 50% of breast cancer patients do not complete a course of hormonal therapy as described, the identification of factors associated with medication discontinuation can facilitate real-time interventions to prevent early discontinuation. Zhijun Yin, Morgan Harrell, Jeremy L. Warner, Qingxia Chen, Daniel Fabbri, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | Development of an automated phenotyping algorithm for hepatorenal syndrome
Jejo Koola, Sharon E. Davis, Omar Al-Nimri, Sharidan K. Parr, Daniel Fabbri, Bradley A. Malin, Samuel B. Ho, Michael E. Matheny |
J. Biomed. Informatics | 5 |
| 2018 | Extracting similar terms from multiple EMR-based semantic embeddings to support chart reviews
Cheng Ye 0001, Daniel Fabbri |
J. Biomed. Informatics | 2 |
| 2017 | Mixed Methods Approach for Understanding Clinical Workflow
Ioana Danciu, Kim M. Unertl, Stuart Weinberg, Daniel Fabbri |
AMIA | 4 |
| 2017 | Evaluating the Effectiveness of Auditing Rules for Electronic Health Record Systems
Monica Hedda, Bradley A. Malin, Chao Yan 0004, Daniel Fabbri |
AMIA | 4 |
| 2017 | A comparative analysis of state-of-the-art SQL-on-Hadoop systems for interactive analyticsabstractHadoop is emerging as the primary data hub in enterprises, and SQL represents the de facto language for data analysis. This combination has led to the development of a variety of SQL-on-Hadoop systems that are in use today. While the various SQL-on-Hadoop systems target the same class of analytical workloads, their different architectures, design decisions and implementations impact query performance. In this work, we perform a comparative analysis of four state-of-the-art SQL-on-Hadoop systems (Impala, Drill, Spark SQL and Phoenix) using the Web Data Analytics micro benchmark and the TPC-H benchmark on the Amazon EC2 cloud platform. The TPC-H experiment results show that, although Impala outperforms other systems (4.41x-6.65x) in the text format, trade-offs exists in the parquet format, with each system performing best on subsets of queries. A comprehensive analysis of execution profiles expands upon the performance results to provide insights into performance variations, performance bottlenecks and query execution characteristics. Ashish Tapdiya, Daniel Fabbri |
IEEE BigData | 2 |
| 2017 | A Comparative Analysis of Materialized Views Selection and Concurrency Control Mechanisms in NoSQL DatabasesabstractRelational databases are well suited for vertical scaling; however, specialized hardware can be expensive. Conversely, NewSQL and NoSQL data stores are designed to scale horizontally. NewSQL databases provide ACID transaction support; however, joins are limited to the partition keys, resulting in restricted query expressiveness. On the other hand, NoSQL databases are designed to scale out on commodity hardware; however, they are limited by slow join performance. Hence, we consider if the NoSQL join performance can be improved while ensuring ACID semantics and without drastically sacrificing write performance, disk utilization and query expressiveness.This paper presents the Synergy system that leverages schema and workload driven mechanisms to identify materialized views, and a specialized concurrency control system on top of a NoSQL database to enable scalable data management with familiar relational conventions. Synergy trades slight write performance degradation and increased disk utilization for faster join performance (compared to standard NoSQL databases) and improved query expressiveness (compared to NewSQL databases). Ashish Tapdiya, Daniel Fabbri |
CLUSTER | 3 |
| 2017 | Classifying patient portal messages using Convolutional Neural Networks
Lina M. Sulieman, David Gilmore, Christi French, Robert M. Cronin, Gretchen Purcell Jackson, Matthew Russell, Daniel Fabbri |
J. Biomed. Informatics | 7 |
| 2016 | Predicting Negative Events: Using Post-discharge Data to Detect High-Risk Patients
Lina M. Sulieman, Daniel Fabbri, Fei Wang 0001, Jianying Hu, Bradley A. Malin |
AMIA | 2 |
| 2016 | Data-Driven System for Perioperative Acuity Prediction
Linda Zhang 0002, Daniel Fabbri, Jonathan P. Wanderer |
AMIA | 2 |
| 2016 | #PrayForDad: Learning the Semantics Behind Why Social Media Users Disclose Health Information
Zhijun Yin, You Chen 0001, Daniel Fabbri, Jimeng Sun 0001, Bradley A. Malin |
ICWSM | 3 |
| 2015 | Automated Classification of Consumer Health Information Needs in Patient Portal Messages
Robert M. Cronin, Daniel Fabbri, Joshua C. Denny, Gretchen Purcell Jackson |
AMIA | 2 |
| 2015 | Mining Twitter as a First Step toward Assessing the Adequacy of Gender Identification Terms on Intake Forms
Amanda Hicks, William R. Hogan, Michael W. Rutherford, Bradley A. Malin, Mengjun Xie, Christiane Fellbaum, Zhijun Yin, Daniel Fabbri, Josh Hanna, Jiang Bian 0001 |
AMIA | 8 |
| 2015 | Comparison of Patient Portal Usage between Employees and Non-Employees
Lina M. Sulieman, Dara Eckerle Mize, Daniel Fabbri, S. Trent Rosenbloom |
AMIA | 3 |
| 2014 | Decide Now or Decide Later?: Quantifying the Tradeoff between Prospective and Retrospective Access DecisionsabstractOne of the greatest challenges an organization faces is determining when an employee is permitted to utilize a certain resource in a system. This "insider threat" can be addressed through two strategies: i) prospective methods, such as access control, that make a decision at the time of a request, and ii) retrospective methods, such as post hoc auditing, that make a decision in the light of the knowledge gathered afterwards. While it is recognized that each strategy has a distinct set of benefits and drawbacks, there has been little investigation into how to provide system administrators with practical guidance on when one or the other should be applied. To address this problem, we introduce a framework to compare these strategies on a common quantitative scale. In doing so, we translate these strategies into classification problems using a context-based feature space that assesses the likelihood that an access request is legitimate. We then introduce a technique called bispective analysis to compare the performance of the classification models under the situation of non-equivalent costs for false positive and negative instances, a significant extension on traditional cost analysis techniques, such as analysis of the receiver operator characteristic (ROC) curve. Using domain-specific cost estimates and access logs of several months from a large Electronic Medical Record (EMR) system, we demonstrate how bispective analysis can support meaningful decisions about the relative merits of prospective and retrospective decision making for specific types of hospital personnel. You Chen 0001, Thaddeus Cybulski, Daniel Fabbri, Carl A. Gunter, Patrick N. Lawlor, David M. Liebovitz, Bradley A. Malin |
CCS | 4 |
| 2013 | SELECT triggers for data auditingabstractAuditing is a key part of the security infrastructure in a database system. While commercial database systems provide mechanisms such as triggers that can be used to track and log any changes made to “sensitive” data using UPDATE queries, they are not useful for tracking accesses to sensitive data using complex SQL queries, which is important for many applications given recent laws such as HIPAA. In this paper, we propose the notion of SELECT triggers that extends triggers to work for SELECT queries in order to facilitate data auditing. We discuss the challenges in integrating SELECT triggers in a database system including specification, semantics as well as efficient implementation techniques. We have prototyped our framework in a commercial database system and present an experimental evaluation of our framework using the TPC-H benchmark. Daniel Fabbri, Ravishankar Ramamurthy, Raghav Kaushik |
ICDE | 1 |
| 2013 | Explaining accesses to electronic medical records using diagnosis informationabstractOBJECTIVE: Ensuring the security and appropriate use of patient health information contained within electronic medical records systems is challenging. Observing these difficulties, we present an addition to the explanation-based auditing system (EBAS) that attempts to determine the clinical or operational reason why accesses occur to medical records based on patient diagnosis information. Accesses that can be explained with a reason are filtered so that the compliance officer has fewer suspicious accesses to review manually. METHODS: Our hypothesis is that specific hospital employees are responsible for treating a given diagnosis. For example, Dr Carl accessed Alice's medical record because Hem/Onc employees are responsible for chemotherapy patients. We present metrics to determine which employees are responsible for a diagnosis and quantify their confidence. The auditing system attempts to use this responsibility information to determine the reason why an access occurred. We evaluate the auditing system's classification quality using data from the University of Michigan Health System. RESULTS: The EBAS correctly determines which departments are responsible for a given diagnosis. Adding this responsibility information to the EBAS increases the number of first accesses explained by a factor of two over previous work and explains over 94% of all accesses with high precision. CONCLUSIONS: The EBAS serves as a complementary security tool for personal health information. It filters a majority of accesses such that it is more feasible for a compliance officer to review the remaining suspicious accesses manually. Daniel Fabbri, Kristen LeFevre |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Explanation-Based AuditingabstractTo comply with emerging privacy laws and regulations, it has become common for applications like electronic health records systems (EHRs) to collect access logs , which record each time a user (e.g., a hospital employee) accesses a piece of sensitive data (e.g., a patient record). Using the access log, it is easy to answer simple queries (e.g., Who accessed Alice's medical record?), but this often does not provide enough information. In addition to learning who accessed their medical records, patients will likely want to understand why each access occurred. In this paper, we introduce the problem of generating explanations for individual records in an access log. The problem is motivated by user-centric auditing applications, and it also provides a novel approach to misuse detection. We develop a framework for modeling explanations which is based on a fundamental observation: For certain classes of databases, including EHRs, the reason for most data accesses can be inferred from data stored elsewhere in the database. For example, if Alice has an appointment with Dr. Dave, this information is stored in the database, and it explains why Dr. Dave looked at Alice's record. Large numbers of data accesses can be explained using general forms called explanation templates . Rather than requiring an administrator to manually specify explanation templates, we propose a set of algorithms for automatically discovering frequent templates from the database (i.e., those that explain a large number of accesses). We also propose techniques for inferring collaborative user groups, which can be used to enhance the quality of the discovered explanations. Finally, we have evaluated our proposed techniques using an access log and data from the University of Michigan Health System. Our results demonstrate that in practice we can provide explanations for over 94% of data accesses in the log. Daniel Fabbri, Kristen LeFevre |
Proc. VLDB Endow. | 1 |
| 2010 | PolicyReplay: Misconfiguration-Response Queries for Data Breach ReportingabstractRecent legislation has increased the requirements of organizations to report data breaches, or unauthorized access to data. While access control policies are used to restrict access to a database, these policies are complex and difficult to configure. As a result, misconfigurations sometimes allow users access to unauthorized data. In this paper, we consider the problem of reporting data breaches after such a misconfiguration is detected. To locate past SQL queries that may have revealed unauthorized information, we introduce the novel idea of a misconfiguration response (MR) query . The MR-query cleanly addresses the challenges of information propagation within the database by replaying the log of operations and returning all logged queries for which the result has changed due to the misconfiguration. A strawman implementation of the MR-query would go back in time and replay all the operations that occurred in the interim, with the correct policy. However, re-executing all operations is inefficient. Instead, we develop techniques to improve reporting efficiency by reducing the number of operations that must be re-executed and reducing the cost of replaying the operations. An extensive evaluation shows that our method can reduce the total runtime by up to an order of magnitude. Daniel Fabbri, Kristen LeFevre, Qiang Zhu 0001 |
Proc. VLDB Endow. | 1 |
| 2009 | PrivatePond: Outsourced Management of Web Corpuses
Daniel Fabbri, Arnab Nandi 0001, Kristen LeFevre, H. V. Jagadish |
WebDB | 1 |