EDBT 2026 Demo / reviewers in the wild / expert
You Chen 0001
dblp:31/3137-1
· DBLP profile ↗
28ranked-venue papers
9as first author
9since 2021 · last 2025
0000-0001-8232-8840ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 9 since 2021Security and privacy · 5 · 3 first-authorDatabases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Large language models are less effective at clinical prediction tasks than locally trained machine learning modelsabstractOBJECTIVES: To determine the extent to which current large language models (LLMs) can serve as substitutes for traditional machine learning (ML) as clinical predictors using data from electronic health records (EHRs), we investigated various factors that can impact their adoption, including overall performance, calibration, fairness, and resilience to privacy protections that reduce data fidelity. MATERIALS AND METHODS: We evaluated GPT-3.5, GPT-4, and traditional ML (as gradient-boosting trees) on clinical prediction tasks in EHR data from Vanderbilt University Medical Center (VUMC) and MIMIC IV. We measured predictive performance with area under the receiver operating characteristic (AUROC) and model calibration using Brier Score. To evaluate the impact of data privacy protections, we assessed AUROC when demographic variables are generalized. We evaluated algorithmic fairness using equalized odds and statistical parity across race, sex, and age of patients. We also considered the impact of using in-context learning by incorporating labeled examples within the prompt. RESULTS: Traditional ML [AUROC: 0.847, 0.894 (VUMC, MIMIC)] substantially outperformed GPT-3.5 (AUROC: 0.537, 0.517) and GPT-4 (AUROC: 0.629, 0.602) (with and without in-context learning) in predictive performance and output probability calibration [Brier Score (ML vs GPT-3.5 vs GPT-4): 0.134 vs 0.384 vs 0.251, 0.042 vs 0.06 vs 0.219)]. DISCUSSION: Traditional ML is more robust than GPT-3.5 and GPT-4 in generalizing demographic information to protect privacy. GPT-4 is the fairest model according to our selected metrics but at the cost of poor model performance. CONCLUSION: These findings suggest that non-fine-tuned LLMs are less effective and robust than locally trained ML for clinical prediction tasks, but they are improving across releases. Katherine E. Brown, Chao Yan 0004, Xinmeng Zhang, Benjamin X. Collins, You Chen 0001, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Discovering clinical drug-drug interactions with known pharmacokinetics mechanisms using spontaneous reporting systems and electronic health recordsabstractOBJECTIVE: Although the mechanisms behind pharmacokinetic (PK) drug-drug interactions (DDIs) are well-documented, bridging the gap between this knowledge and clinical evidence of DDIs, especially for serious adverse drug reactions (SADRs), remains challenging. While leveraging the FDA Adverse Event Reporting System (FAERS) database along with disproportionality analysis tends to detect a vast number of DDI signals, this abundance complicates further investigation, such as validation through clinical trials. Our study proposed a framework to efficiently prioritize these signals and assessed their reliability using multi-source Electronic Health Records (EHR) to identify top candidates for further investigation. METHODS: We analyzed FAERS data spanning from January 2004 to March 2023, employing four established disproportionality methods: Proportional Reporting Ratio (PRR), Reporting Odds Ratio (ROR), Multi-item Gamma Poisson Shrinker (MGPS), and Bayesian Confidence Propagating Neural Network (BCPNN). Building upon these models, we developed four ranking models to prioritize DDI-SADR signals and cross-referenced signals with DrugBank. To validate the top-ranked signals, we employed longitudinal EHRs from Vanderbilt University Medical Center and the All of Us research program. The performance of each model was assessed by counting how many of the top-ranked signals were confirmed by EHRs and calculating the average ranking of these confirmed signals. RESULTS: Out of 189 DDI-SADR signals identified by all four disproportionality methods, only two were documented in the DrugBank database. By prioritizing the top 20 signals as determined by each of the four disproportionality methods and our four ranking models, 58 unique DDI-SADR signals were selected for EHR validations. Of these, five signals were confirmed. The ranking model, which integrated the MGPS and BCPNN, demonstrated superior performance by assigning the highest priority to those five EHR-confirmed signals. CONCLUSION: The fusion of disproportionality analysis with ranking models, validated through multi-source EHRs, presents a groundbreaking approach to pharmacovigilance. Our study's confirmation of five significant DDI-SADRs, previously unrecorded in the DrugBank database, highlights the essential role of advanced data analysis techniques in identifying ADRs. Eugene Jeong, Yu Su 0001, Lang Li 0001, You Chen 0001 |
J. Biomed. Informatics | 4 |
| 2022 | Inferring EHR Utilization Workflows through Audit Logs
Xinmeng Zhang, Yuying Zhao, Chao Yan 0004, Tyler Derr, You Chen 0001 |
AMIA | 5 |
| 2021 | Telehealth Uptake and Continuing Usage During the COVID-19 Pandemic
Bradley A. Malin, You Chen 0001 |
AMIA | 3 |
| 2021 | Discovering Drug-Drug Interactions in COVID-19 Patients
Eugene Jeong, Anna K. Person, Joanna L. Stollings, Lang Li 0001, You Chen 0001 |
AMIA | 5 |
| 2021 | Predicting Next-Day Discharge via Electronic Health Record Audit Logs
Xinmeng Zhang, Chao Yan 0004, Mayur B. Patel, Bradley A. Malin, You Chen 0001 |
AMIA | 5 |
| 2021 | Mining tasks and task characteristics from electronic health record audit logs with unsupervised machine learningabstractOBJECTIVE: The characteristics of clinician activities while interacting with electronic health record (EHR) systems can influence the time spent in EHRs and workload. This study aims to characterize EHR activities as tasks and define novel, data-driven metrics. MATERIALS AND METHODS: We leveraged unsupervised learning approaches to learn tasks from sequences of events in EHR audit logs. We developed metrics characterizing the prevalence of unique events and event repetition and applied them to categorize tasks into 4 complexity profiles. Between these profiles, Mann-Whitney U tests were applied to measure the differences in performance time, event type, and clinician prevalence, or the number of unique clinicians who were observed performing these tasks. In addition, we apply process mining frameworks paired with clinical annotations to support the validity of a sample of our identified tasks. We apply our approaches to learn tasks performed by nurses in the Vanderbilt University Medical Center neonatal intensive care unit. RESULTS: We examined EHR audit logs generated by 33 neonatal intensive care unit nurses resulting in 57 234 sessions and 81 tasks. Our results indicated significant differences in performance time for each observed task complexity profile. There were no significant differences in clinician prevalence or in the frequency of viewing and modifying event types between tasks of different complexities. We presented a sample of expert-reviewed, annotated task workflows supporting the interpretation of their clinical meaningfulness. CONCLUSIONS: The use of the audit log provides an opportunity to assist hospitals in further investigating clinician activities to optimize EHR workflows. Bob Chen 0001, Mhd Wael Alrifai, Barrett Jones, Laurie L. Novak, Nancy M. Lorenzi, Daniel J. France, Bradley A. Malin, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2021 | Predicting brain function status changes in critically ill patients via Machine learningabstractOBJECTIVE: In intensive care units (ICUs), a patient's brain function status can shift from a state of acute brain dysfunction (ABD) to one that is ABD-free and vice versa, which is challenging to forecast and, in turn, hampers the allocation of hospital resources. We aim to develop a machine learning model to predict next-day brain function status changes. MATERIALS AND METHODS: Using multicenter prospective adult cohorts involving medical and surgical ICU patients from 2 civilian and 3 Veteran Affairs hospitals, we trained and externally validated a light gradient boosting machine to predict brain function status changes. We compared the performances of the boosting model against state-of-the-art models-an ABD predictive model and its variants. We applied Shapley additive explanations to identify influential factors to develop a compact model. RESULTS: There were 1026 critically ill patients without evidence of prior major dementia, or structural brain diseases, from whom 12 295 daily transitions (ABD: 5847 days; ABD-free: 6448 days) were observed. The boosting model achieved an area under the receiver-operating characteristic curve (AUROC) of 0.824 (95% confidence interval [CI], 0.821-0.827), compared with the state-of-the-art models of 0.697 (95% CI, 0.693-0.701) with P < .001. Using 13 identified top influential factors, the compact model achieved 99.4% of the boosting model on AUROC. The boosting and the compact models demonstrated high generalizability in external validation by achieving an AUROC of 0.812 (95% CI, 0.812-0.813). CONCLUSION: The inputs of the compact model are based on several simple questions that clinicians can quickly answer in practice, which demonstrates the model has direct prospective deployment potential into clinical practice, aiding in critical hospital resource allocation. Chao Yan 0004, Ziqi Zhang 0005, Wencong Chen, Bradley A. Malin, Eugene Wesley Ely, Mayur B. Patel, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | Predicting next-day discharge via electronic health record access logsabstractOBJECTIVE: Hospital capacity management depends on accurate real-time estimates of hospital-wide discharges. Estimation by a clinician requires an excessively large amount of effort and, even when attempted, accuracy in forecasting next-day patient-level discharge is poor. This study aims to support next-day discharge predictions with machine learning by incorporating electronic health record (EHR) audit log data, a resource that captures EHR users' granular interactions with patients' records by communicating various semantics and has been neglected in outcome predictions. MATERIALS AND METHODS: This study focused on the EHR data for all adults admitted to Vanderbilt University Medical Center in 2019. We learned multiple advanced models to assess the value that EHR audit log data adds to the daily prediction of discharge likelihood within 24 h and to compare different representation strategies. We applied Shapley additive explanations to identify the most influential types of user-EHR interactions for discharge prediction. RESULTS: The data include 26 283 inpatient stays, 133 398 patient-day observations, and 819 types of user-EHR interactions. The model using the count of each type of interaction in the recent 24 h and other commonly used features, including demographics and admission diagnoses, achieved the highest area under the receiver operating characteristics (AUROC) curve of 0.921 (95% CI: 0.919-0.923). By contrast, the model lacking user-EHR interactions achieved a worse AUROC of 0.862 (0.860-0.865). In addition, 10 of the 20 (50%) most influential factors were user-EHR interaction features. CONCLUSION: EHR audit log data contain rich information such that it can improve hospital-wide discharge predictions. Xinmeng Zhang, Chao Yan 0004, Bradley A. Malin, Mayur B. Patel, You Chen 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Learning Tasks of Pediatric Providers from Electronic Health Record Audit Logs
Barrett Jones, Xinmeng Zhang, Bradley A. Malin, You Chen 0001 |
AMIA | 4 |
| 2020 | Metrics for assessing physician activity using electronic health record log dataabstractElectronic health record (EHR) log data have shown promise in measuring physician time spent on clinical activities, contributing to deeper understanding and further optimization of the clinical environment. In this article, we propose 7 core measures of EHR use that reflect multiple dimensions of practice efficiency: total EHR time, work outside of work, time on documentation, time on prescriptions, inbox time, teamwork for orders, and an aspirational measure for the amount of undivided attention patients receive from their physicians during an encounter, undivided attention. We also illustrate sample use cases for these measures for multiple stakeholders. Finally, standardization of EHR log data measure specifications, as outlined here, will foster cross-study synthesis and comparative research. Christine A. Sinsky, Adam Rule, Genna R. Cohen, Brian G. Arndt, Tait D. Shanafelt, Christopher D. Sharp, Sally L. Baxter, Ming Tai-Seale, Sherry H. F. Yan, You Chen 0001, Julia Adler-Milstein, Michelle R. Hribar |
J. Am. Medical Informatics Assoc. | 10 |
| 2019 | Advancing Common Approaches to Working with EHR Log Data
Julia Adler-Milstein, You Chen 0001, Michelle R. Hribar, Jennifer R. Popovic, J. Marc Overhage |
AMIA | 2 |
| 2019 | Corpus Size Influences Clinical Concept Embeddings
Chao Yan 0004, Bradley A. Malin, You Chen 0001 |
AMIA | 4 |
| 2019 | Deep learning predicts extreme preterm birth from electronic health records
Sarah Osmundson, Digna Velez Edwards, Gretchen Purcell Jackson, Bradley A. Malin, You Chen 0001 |
J. Biomed. Informatics | 6 |
| 2018 | Interaction patterns of trauma providers are associated with length of stayabstractBackground: Trauma-related hospitalizations drive a high percentage of health care expenditure and inpatient resource consumption, which is directly related to length of stay (LOS). Robust and reliable interactions among health care employees can reduce LOS. However, there is little known about whether certain patterns of interactions exist and how they relate to LOS and its variability. The objective of this study is to learn interaction patterns and quantify the relationship to LOS within a mature trauma system and long-standing electronic medical record (EMR). Methods: We adapted a spectral co-clustering methodology to infer the interaction patterns of health care employees based on the EMR of 5588 hospitalized adult trauma survivors. The relationship between interaction patterns and LOS was assessed via a negative binomial regression model. We further assessed the influence of potential confounders by age, number of health care encounters to date, number of access action types care providers committed to patient EMRs, month of admission, phenome-wide association study codes, procedure codes, and insurance status. Results: Three types of interaction patterns were discovered. The first pattern exhibited the most collaboration between employees and was associated with the shortest LOS. Compared to this pattern, LOS for the second and third patterns was 0.61 days (P = 0.014) and 0.43 days (P = 0.037) longer, respectively. Although the 3 interaction patterns dealt with different numbers of patients in each admission month, our results suggest that care was provided for similar patients. Discussion: The results of this study indicate there is an association between LOS and the extent to which health care employees interact in the care of an injured patient. The findings further suggest that there is merit in ascertaining the content of these interactions and the factors that induce these differences in interaction patterns within a trauma system. You Chen 0001, Mayur B. Patel, Candace D. McNaughton, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Learning bundled care opportunities from electronic medical records
You Chen 0001, Abel N. Kho, David M. Liebovitz, Catherine Ivory, Sarah Osmundson, Jiang Bian 0001, Bradley A. Malin |
J. Biomed. Informatics | 1 |
| 2017 | Identifying collaborative care teams through electronic medical record utilization patternsabstractOBJECTIVE: The goal of this investigation was to determine whether automated approaches can learn patient-oriented care teams via utilization of an electronic medical record (EMR) system. MATERIALS AND METHODS: To perform this investigation, we designed a data-mining framework that relies on a combination of latent topic modeling and network analysis to infer patterns of collaborative teams. We applied the framework to the EMR utilization records of over 10 000 employees and 17 000 inpatients at a large academic medical center during a 4-month window in 2010. Next, we conducted an extrinsic evaluation of the patterns to determine the plausibility of the inferred care teams via surveys with knowledgeable experts. Finally, we conducted an intrinsic evaluation to contextualize each team in terms of collaboration strength (via a cluster coefficient) and clinical credibility (via associations between teams and patient comorbidities). RESULTS: The framework discovered 34 collaborative care teams, 27 (79.4%) of which were confirmed as administratively plausible. Of those, 26 teams depicted strong collaborations, with a cluster coefficient > 0.5. There were 119 diagnostic conditions associated with 34 care teams. Additionally, to provide clarity on how the survey respondents arrived at their determinations, we worked with several oncologists to develop an illustrative example of how a certain team functions in cancer care. DISCUSSION: Inferred collaborative teams are plausible; translating such patterns into optimized collaborative care will require administrative review and integration with management practices. CONCLUSIONS: EMR utilization records can be mined for collaborative care patterns in large complex medical centers. You Chen 0001, Nancy M. Lorenzi, Warren S. Sandberg, Kelly Wolgast, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Learning Clinical Workflows to Identify Subgroups of Heart Failure Patients
Chao Yan 0004, You Chen 0001, Bo Li 0026, David M. Liebovitz, Bradley A. Malin |
AMIA | 2 |
| 2016 | #PrayForDad: Learning the Semantics Behind Why Social Media Users Disclose Health Information
Zhijun Yin, You Chen 0001, Daniel Fabbri, Jimeng Sun 0001, Bradley A. Malin |
ICWSM | 2 |
| 2015 | Inferring Clinical Workflow Efficiency via Electronic Medical Record Utilization
You Chen 0001, Wei Xie 0002, Carl A. Gunter, David M. Liebovitz, Sanjay Mehrotra, Bradley A. Malin |
AMIA | 1 |
| 2015 | Rubik: Knowledge Guided Tensor Factorization and Completion for Health Data AnalyticsabstractComputational phenotyping is the process of converting heterogeneous electronic health records (EHRs) into meaningful clinical concepts. Unsupervised phenotyping methods have the potential to leverage a vast amount of labeled EHR data for phenotype discovery. However, existing unsupervised phenotyping methods do not incorporate current medical knowledge and cannot directly handle missing, or noisy data. We propose Rubik, a constrained non-negative tensor factorization and completion method for phenotyping. Rubik incorporates 1) guidance constraints to align with existing medical knowledge, and 2) pairwise constraints for obtaining distinct, non-overlapping phenotypes. Rubik also has built-in tensor completion that can significantly alleviate the impact of noisy and missing data. We utilize the Alternating Direction Method of Multipliers (ADMM) framework to tensor factorization and completion, which can be easily scaled through parallel computing. We evaluate Rubik on two EHR datasets, one of which contains 647,118 records for 7,744 patients from an outpatient clinic, the other of which is a public dataset containing 1,018,614 CMS claims records for 472,645 patients. Our results show that Rubik can discover more meaningful and distinct phenotypes than the baselines. In particular, by using knowledge guidance constraints, Rubik can also discover sub-phenotypes for several major diseases. Rubik also runs around seven times faster than current state-of-the-art tensor methods. Finally, Rubik is scalable to large datasets containing millions of EHR records. Yichen Wang 0001, Robert Chen 0001, Joydeep Ghosh, Joshua C. Denny, Abel N. Kho, You Chen 0001, Bradley A. Malin, Jimeng Sun 0001 |
KDD | 6 |
| 2015 | Building bridges across electronic health record systems through inferred phenotypic topics
You Chen 0001, Joydeep Ghosh, Cosmin Adrian Bejan, Carl A. Gunter, Siddharth Gupta 0005, Abel N. Kho, David M. Liebovitz, Jimeng Sun 0001, Joshua C. Denny, Bradley A. Malin |
J. Biomed. Informatics | 1 |
| 2014 | Decide Now or Decide Later?: Quantifying the Tradeoff between Prospective and Retrospective Access DecisionsabstractOne of the greatest challenges an organization faces is determining when an employee is permitted to utilize a certain resource in a system. This "insider threat" can be addressed through two strategies: i) prospective methods, such as access control, that make a decision at the time of a request, and ii) retrospective methods, such as post hoc auditing, that make a decision in the light of the knowledge gathered afterwards. While it is recognized that each strategy has a distinct set of benefits and drawbacks, there has been little investigation into how to provide system administrators with practical guidance on when one or the other should be applied. To address this problem, we introduce a framework to compare these strategies on a common quantitative scale. In doing so, we translate these strategies into classification problems using a context-based feature space that assesses the likelihood that an access request is legitimate. We then introduce a technique called bispective analysis to compare the performance of the classification models under the situation of non-equivalent costs for false positive and negative instances, a significant extension on traditional cost analysis techniques, such as analysis of the receiver operator characteristic (ROC) curve. Using domain-specific cost estimates and access logs of several months from a large Electronic Medical Record (EMR) system, we demonstrate how bispective analysis can support meaningful decisions about the relative merits of prospective and retrospective decision making for specific types of hospital personnel. You Chen 0001, Thaddeus Cybulski, Daniel Fabbri, Carl A. Gunter, Patrick N. Lawlor, David M. Liebovitz, Bradley A. Malin |
CCS | 2 |
| 2013 | Evolving role definitions through permission invocation patternsabstractIn role-based access control (RBAC), roles are traditionally defined as sets of permissions. Roles specified by administrators may be inaccurate, however, such that data mining methods have been proposed to learn roles from actual permission utilization. These methods minimize variation from an information theoretic perspective, but they neglect the expert knowledge of administrators. In this paper, we propose a strategy to enable a controlled evolution of RBAC based on utilization. To accomplish this goal, we extend a subset enumeration framework to search candidate roles for an RBAC model that addresses an objective function which balances administrator beliefs and permission utilization. The rate of role evolution is controlled by an administrator-specified parameter. To assess effectiveness, we perform an empirical analysis using simulations, as well as a real world dataset from an electronic medical record system (EMR) in use at a large academic medical center (over 8000 users, 140 roles, and 140 permissions). We compare the results with several state-of-the-art role mining algorithms using 1) an outlier detection method on the new roles to evaluate the homogeneity of their behavior and 2)a set-based similarity measure between the original and new roles. The results illustrate our method is comparable to the state-of-the-art, but allows for a range of RBAC models which tradeoff user behavior and administrator expectations. For instance, in the EMR dataset, we find the resulting RBAC model contains 22% outliers and a distance of 0.02 to the original RBAC model when the system is biased toward administrator belief, and 13% outliers and a distance of 0.26 to the original RBAC model when biased toward permission utilization. You Chen 0001, Carl A. Gunter, David M. Liebovitz, Bradley A. Malin |
SACMAT | 2 |
| 2012 | Auditing Medical Records Accesses via Healthcare Interaction Networks
You Chen 0001, Steve Nyemba, Bradley A. Malin |
AMIA | 1 |
| 2012 | Detecting Anomalous Insiders in Collaborative Information SystemsabstractCollaborative information systems (CISs) are deployed within a diverse array of environments that manage sensitive information. Current security mechanisms detect insider threats, but they are ill-suited to monitor systems in which users function in dynamic teams. In this paper, we introduce the community anomaly detection system (CADS), an unsupervised learning framework to detect insider threats based on the access logs of collaborative environments. The framework is based on the observation that typical CIS users tend to form community structures based on the subjects accessed (e.g., patients' records viewed by healthcare providers). CADS consists of two components: 1) relational pattern extraction, which derives community structures and 2) anomaly prediction, which leverages a statistical model to determine when users have sufficiently deviated from communities. We further extend CADS into MetaCADS to account for the semantics of subjects (e.g., patients' diagnoses). To empirically evaluate the framework, we perform an assessment with three months of access logs from a real electronic health record (EHR) system in a large medical center. The results illustrate our models exhibit significant performance gains over state-of-the-art competitors. When the number of illicit users is low, MetaCADS is the best model, but as the number grows, commonly accessed semantics lead to hiding in a crowd, such that CADS is more prudent. You Chen 0001, Steve Nyemba, Bradley A. Malin |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2011 | Detection of anomalous insiders in collaborative environments via relational analysis of access logsabstractCollaborative information systems (CIS) are deployed within a diverse array of environments, ranging from the Internet to intelligence agencies to healthcare. It is increasingly the case that such systems are applied to manage sensitive information, making them targets for malicious insiders. While sophisticated security mechanisms have been developed to detect insider threats in various file systems, they are neither designed to model nor to monitor collaborative environments in which users function in dynamic teams with complex behavior. In this paper, we introduce a community-based anomaly detection system (CADS), an unsupervised learning framework to detect insider threats based on information recorded in the access logs of collaborative environments. CADS is based on the observation that typical users tend to form community structures, such that users with low affinity to such communities are indicative of anomalous and potentially illicit behavior. The model consists of two primary components: relational pattern extraction and anomaly detection. For relational pattern extraction, CADS infers community structures from CIS access logs, and subsequently derives communities, which serve as the CADS pattern core. CADS then uses a formal statistical model to measure the deviation of users from the inferred communities to predict which users are anomalies. To empirically evaluate the threat detection model, we perform an analysis with six months of access logs from a real electronic health record system in a large medical center, as well as a publicly available dataset for replication purposes. The results illustrate that CADS can distinguish simulated anomalous users in the context of real user behavior with a high degree of certainty and with significant performance gains in comparison to several competing anomaly detection models. You Chen 0001, Bradley A. Malin |
CODASPY | 1 |
| 2011 | Leveraging social networks to detect anomalous insider actions in collaborative environmentsabstractCollaborative information systems (CIS) enable users to coordinate efficiently over shared tasks. T hey are often deployed in complex dynamic systems that provide users with broad access privileges, but also leave the system vulnerable to various attacks. Techniques to detect threats originating from beyond the system are relatively mature, but methods to detect insider threats are still evolving. A promising class of insider threat detection models for CIS focus on the communities that manifest between users based on the usage of common subjects in the system. However, current methods detect only when a user's aggregate behavior is intruding, not when specific actions have deviated from expectation. In this paper, we introduce a method called specialized network anomaly detection (SNAD) to detect such events. SNAD assembles the community of users that access a particular subject and assesses if similarities of the community with and without a certain user are sufficiently different. We present a theoretical basis and perform an extensive empirical evaluation with the access logs of two distinct environments: those of a large electronic health record system (6,015 users, 130,457 patients and 1,327,500 accesses) and the editing logs of Wikipedia (2,388,955 revisors, 55,200 articles and 6,482,780 revisions). We compare SNAD with several competing methods and demonstrate it is significantly more effective: on average it achieves 20-30% greater area under an ROC curve. You Chen 0001, Steve Nyemba, Bradley A. Malin |
ISI | 1 |