EDBT 2026 Demo / reviewers in the wild / expert
György J. Simon
dblp:39/4973
· DBLP profile ↗
39ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0003-4715-5934ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Combining self-supervision and privileged information for representation learning from tabular dataabstractAbstract When building predictive models for real-world applications, many data are discarded because conventional learning algorithms cannot utilize it, although such data could be very informative. This paper focuses on representation learning using two types of additional data: privileged information (PI) and unlabeled data. PI refers to data available only during training but not at test time. Existing methods transfer the knowledge embedded in PI via supervised mechanisms, making them unable to use unlabeled data. In contrast, self-supervised learning methods can use unlabeled data but cannot learn from PI. While these techniques appear complementary, as we demonstrate, combining them is non-trivial. This paper introduces the privileged information regularized (PIReg) self-supervised learning framework, which utilizes both PI and unlabeled data to learn better representations. Michael S. Steinbach, Genevieve B. Melton, Vipin Kumar 0001, György J. Simon |
Knowl. Inf. Syst. | 5 |
| 2024 | Combining Self-Supervision and Privileged Information for Representation Learning from Tabular DataabstractWhen building predictive models for real-world applications, many data are discarded because conventional learning algorithms cannot utilize it, although such data could be very informative. This paper focuses on representation learning using two types of additional data: privileged information (PI) and unlabeled data. PI refers to data available only during training but not at test time. Existing methods transfer the knowledge embedded in PI via supervised mechanisms, making them unable to use unlabeled data. In contrast, self-supervised learning methods can use unlabeled data but cannot learn from PI. While these techniques appear complementary, as we demonstrate, combining them is non-trivial. This paper introduces the Privileged Information Regularized (PIReg) self-supervised learning framework, which utilizes both PI and unlabeled data to learn better representations. Michael S. Steinbach, Genevieve B. Melton, Vipin Kumar 0001, György J. Simon |
ICDM | 5 |
| 2024 | Consensus modeling: Safer transfer learning for small health systems
Roshan Tourani, Dennis Murphree, Adam Sheka, Genevieve B. Melton, Daryl J. Kor, György J. Simon |
Artif. Intell. Medicine | 6 |
| 2022 | Application of Causal Discovery Algorithms in Studying the Nephrotoxicity of Remdesivir Using Longitudinal Data from the EHR
Erich Kummerfeld, Gretchen M. Hultman, Paul E. Drawz, Terrence Adams, György J. Simon, Genevieve B. Melton |
AMIA | 6 |
| 2022 | Predicting Cancer Treatments Induced Cardiotoxicity of Breast Cancer Patients Using Electronic Health Record
Rui Zhang 0028, Xinpeng Shen, Chetan Shenoy, Anne H. Blaes, György J. Simon |
AMIA | 6 |
| 2022 | Robust Methods for Quantifying the Effect of a Continuous Exposure From Observational DataabstractA cornerstone of clinical medicine is intervening on a continuous exposure, such as titrating the dosage of a pharmaceutical or controlling a laboratory result. In clinical trials, continuous exposures are dichotomized into narrow ranges, excluding large portions of the realistic treatment scenarios. The existing computational methods for estimating the effect of continuous exposure rely on a set of strict assumptions. We introduce new methods that are more robust towards violations of these assumptions. Our methods are based on the key observation that changes of exposure in the clinical setting are often achieved gradually, so effect estimates must be "locally" robust in narrower exposure ranges. We compared our methods with several existing methods on three simulated studies with increasing complexity. We also applied the methods to data from 14 k sepsis patients at M Health Fairview to estimate the effect of antibiotic administration latency on prolonged hospital stay. The proposed methods achieve good performance in all simulation studies. When the assumptions were violated, the proposed methods had estimation errors of one half to one fifth of the state-of-the-art methods. Applying our methods to the sepsis cohort resulted in effect estimates consistent with clinical knowledge. Roshan Tourani, Sisi Ma, Michael Usher, György J. Simon |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Validation of Administrative Coding and Clinical Notes for Hospital-Acquired Acute Kidney Injury in Adults
Paul E. Drawz, Gretchen M. Hultman, György J. Simon, Genevieve B. Melton |
AMIA | 5 |
| 2021 | NLP Methods for Extraction of Symptoms from Unstructured Data for Use in Prognostic COVID-19 Analytic ModelsabstractStatistical modeling of outcomes based on a patient's presenting symptoms (symptomatology) can help deliver high quality care and allocate essential resources, which is especially important during the COVID-19 pandemic. Patient symptoms are typically found in unstructured notes, and thus not readily available for clinical decision making. In an attempt to fill this gap, this study compared two methods for symptom extraction from Emergency Department (ED) admission notes. Both methods utilized a lexicon derived by expanding The Center for Disease Control and Prevention's (CDC) Symptoms of Coronavirus list. The first method utilized a word2vec model to expand the lexicon using a dictionary mapping to the Uni ed Medical Language System (UMLS). The second method utilized the expanded lexicon as a rule-based gazetteer and the UMLS. These methods were evaluated against a manually annotated reference (f1-score of 0.87 for UMLS-based ensemble; and 0.85 for rule-based gazetteer with UMLS). Through analyses of associations of extracted symptoms used as features against various outcomes, salient risks among the population of COVID-19 patients, including increased risk of in-hospital mortality (OR 1.85, p-value < 0.001), were identified for patients presenting with dyspnea. Disparities between English and non-English speaking patients were also identified, the most salient being a concerning finding of opposing risk signals between fatigue and in-hospital mortality (non-English: OR 1.95, p-value = 0.02; English: OR 0.63, p-value = 0.01). While use of symptomatology for modeling of outcomes is not unique, unlike previous studies this study showed that models built using symptoms with the outcome of in-hospital mortality were not significantly different from models using data collected during an in-patient encounter (AUC of 0.9 with 95% CI of [0.88, 0.91] using only vital signs; AUC of 0.87 with 95% CI of [0.85, 0.88] using only symptoms). These findings indicate that prognostic models based on symptomatology could aid in extending COVID-19 patient care through telemedicine, replacing the need for in-person options. The methods presented in this study have potential for use in development of symptomatology-based models for other diseases, including for the study of Post-Acute Sequelae of COVID-19 (PASC). Greg M. Silverman, Himanshu S. Sahoo, Nicholas Ingraham, Monica Lupei, Michael A. Puskarich, Michael Usher, James Dries, Raymond L. Finzel, Eric Murray, John Sartori, György J. Simon, Rui Zhang 0028, Genevieve B. Melton, Christopher J. Tignanelli, Serguei V. S. Pakhomov |
J. Artif. Intell. Res. | 11 |
| 2021 | A likelihood-based convolution approach to estimate major health events in longitudinal health records data: an external validation studyabstractOBJECTIVE: In electronic health record data, the exact time stamp of major health events, defined by significant physiologic or treatment changes, is often missing. We developed and externally validated a method that can accurately estimate these time stamps based on accurate time stamps of related data elements. MATERIALS AND METHODS: A novel convolution-based change detection methodology was developed and tested using data from the national deidentified clinical claims OptumLabs data warehouse, then externally validated on a single center dataset derived from the M Health Fairview system. RESULTS: We applied the methodology to estimate time to liver transplantation for waitlisted candidates. The median error between estimated date within the period of the actual true date was zero days, and median error was 92% and 84% of the transplants, in development and validation samples, respectively. DISCUSSION: The proposed method can accurately estimate missing time stamps. Successful external validation suggests that the proposed method does not need to be refit to each health system; thus, it can be applied even when training data at the health system is insufficient or unavailable. The proposed method was applied to liver transplantation but can be more generally applied to any missing event that is accompanied by multiple related events that have accurate time stamps. CONCLUSION: Missing time stamps in electronic healthcare record data can be estimated using time stamps of related events. Since the model was developed on a nationally representative dataset, it could be successfully transferred to a local health system without substantial loss of accuracy. Lisiane Pruinelli, Jiaqi Zhou 0007, Bethany Stai, Jesse Schold, Timothy Pruett, Sisi Ma, György J. Simon |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Strategies for building robust prediction models using data unavailable at prediction timeabstractOBJECTIVE: Hospital-acquired infections (HAIs) are associated with significant morbidity, mortality, and prolonged hospital length of stay. Risk prediction models based on pre- and intraoperative data have been proposed to assess the risk of HAIs at the end of the surgery, but the performance of these models lag behind HAI detection models based on postoperative data. Postoperative data are more predictive than pre- or interoperative data since it is closer to the outcomes in time, but it is unavailable when the risk models are applied (end of surgery). The objective is to study whether such data, which is temporally unavailable at prediction time (TUP) (and thus cannot directly enter the model), can be used to improve the performance of the risk model. MATERIALS AND METHODS: An extensive array of 12 methods based on logistic/linear regression and deep learning were used to incorporate the TUP data using a variety of intermediate representations of the data. Due to the hierarchical structure of different HAI outcomes, a comparison of single and multi-task learning frameworks is also presented. RESULTS AND DISCUSSION: The use of TUP data was always advantageous as baseline methods, which cannot utilize TUP data, never achieved the top performance. The relative performances of the different models vary across the different outcomes. Regarding the intermediate representation, we found that its complexity was key and that incorporating label information was helpful. CONCLUSIONS: Using TUP data significantly helped predictive performance irrespective of the model complexity. Roshan Tourani, Vipin Kumar 0001, Genevieve B. Melton, Michael S. Steinbach, György J. Simon |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | A Computational Method for Learning Disease Trajectories From Partially Observable EHR DataabstractDiseases can show different courses of progression even when patients share the same risk factors. Recent studies have revealed that the use of trajectories, the order in which diseases manifest throughout life, can be predictive of the course of progression. In this study, we propose a novel computational method for learning disease trajectories from EHR data. The proposed method consists of three parts: first, we propose an algorithm for extracting trajectories from EHR data; second, three criteria for filtering trajectories; and third, a likelihood function for assessing the risk of developing a set of outcomes given a trajectory set. We applied our methods to extract a set of disease trajectories from Mayo Clinic EHR data and evaluated it internally based on log-likelihood, which can be interpreted as the trajectories' ability to explain the observed (partial) disease progressions. We then externally evaluated the trajectories on EHR data from an independent health system, M Health Fairview. The proposed algorithm extracted a comprehensive set of disease trajectories that can explain the observed outcomes substantially better than competing methods and the proposed filtering criteria selected a small subset of disease trajectories that are highly interpretable and suffered only a minimal (relative 5%) loss of the ability to explain disease progression in both the internal and external validation. Wonsuk Oh, Michael S. Steinbach, Regina Castro, Kevin A. Peterson, Vipin Kumar 0001, Pedro J. Caraballo, György J. Simon |
IEEE J. Biomed. Health Informatics | 7 |
| 2020 | Consensus Modeling: A Transfer Learning Approach for Small Health Systems
Roshan Tourani, Dennis Murphree, Adam Sheka, Genevieve B. Melton, Daryl J. Kor, György J. Simon |
AIME | 7 |
| 2020 | Innovative Method to Build Robust Prediction Models When Gold-Standard Outcomes Are Scarce
Roshan Tourani, Adam Sheka, Elizabeth C. Wick, Genevieve B. Melton, György J. Simon |
AIME | 6 |
| 2019 | A new representation of disease conditions and treatment pathways accurately predicts mortality and chronic diseases
Che Ngufor, Pedro J. Caraballo, Thomas J. O'Byrne, David Chen 0003, Nilay D. Shah, Michael S. Steinbach, György J. Simon |
AMIA | 7 |
| 2019 | Using Electronic Health Record Data to Identify Patients with Prediabetes
Thomas J. O'Byrne, Regina Castro, Che Ngufor, György J. Simon, Pedro J. Caraballo |
AMIA | 4 |
| 2019 | Frequent Causal Pattern Mining: A Computationally Efficient Framework For Estimating Bias-Corrected EffectsabstractOur aging population increasingly suffers from multiple chronic diseases simultaneously, necessitating the comprehensive treatment of these conditions. Finding the optimal set of drugs for a combinatorial set of diseases is a combinatorial pattern exploration problem. Association rule mining is a popular tool for such problems, but the requirement of health care for finding causal, rather than associative, patterns renders association rule mining unsuitable. To address this issue, we propose a novel framework based on the Rubin-Neyman causal model for extracting causal rules from observational data, correcting for a number of common biases. Specifically, given a set of interventions and a set of items that define subpopulations (e.g., diseases), we wish to find all subpopulations in which effective intervention combinations exist and in each such subpopulation, we wish to find all intervention combinations such that dropping any intervention from this combination will reduce the efficacy of the treatment. A key aspect of our framework is the concept of closed intervention sets which extend the concept of quantifying the effect of a single intervention to a set of concurrent interventions. Closed intervention sets also allow for a pruning strategy that is strictly more efficient than the traditional pruning strategy used by the Apriori algorithm. To implement our ideas, we introduce and compare five methods of estimating causal effect from observational data and rigorously evaluate them on synthetic data to mathematically prove (when possible) why they work. We also evaluated our causal rule mining framework on the Electronic Health Records (EHR) data of a large cohort of 152000 patients from Mayo Clinic and showed that the patterns we extracted are sufficiently rich to explain the controversial findings in the medical literature regarding the effect of a class of cholesterol drugs on Type-II Diabetes Mellitus (T2DM). Pranjul Yadav, Michael S. Steinbach, Regina Castro, Pedro J. Caraballo, Vipin Kumar 0001, György J. Simon |
IEEE BigData | 6 |
| 2017 | Causal Phenotyping for Susceptibility to Cardiotoxicity from Antineoplastic Breast Cancer Medications
Deyu Sun, György J. Simon, Steven J. Skube, Anne H. Blaes, Genevieve B. Melton, Rui Zhang 0028 |
AMIA | 2 |
| 2017 | Strategies for handling missing clinical data for automated surgical site infection detection from the electronic health record
Genevieve B. Melton, Elliot G. Arsoniadis, Yan Wang 0025, Mary R. Kwaan, György J. Simon |
J. Biomed. Informatics | 6 |
| 2016 | Clinical Decision Support to Detect High Risk Patients with QTc Prolongation
Christopher A. Aakre, J. Martijn Bos, György J. Simon, Robert F. Tarrell, Michael J. Ackerman, Pedro J. Caraballo |
AMIA | 3 |
| 2016 | Accelerating Chart Review Using Automated Methods on Electronic Health Record Data for Postoperative Complications
Genevieve B. Melton, Nathan D. Moeller, Elliot G. Arsoniadis, Yan Wang 0025, Mary R. Kwaan, Eric Jensen, György J. Simon |
AMIA | 8 |
| 2016 | Quantifying the Effect of Data Quality on the Correctness of an eMeasure
Steven G. Johnson, Stuart M. Speedie, György J. Simon, Vipin Kumar 0001, Bonnie L. Westra |
AMIA | 3 |
| 2015 | Predicting the Factors of Improvement of Health Status of Home Health Care Patients: A Holistic Data Mining Approach
Sanjoy Dey, Katherine Hauwiller, Pranjul Yadav, Michael S. Steinbach, György J. Simon, Vipin Kumar 0001, Connie White-Delaney, Bonnie L. Westra |
AMIA | 5 |
| 2015 | A Data Quality Ontology for the Secondary Use of EHR Data
Steven G. Johnson, Stuart M. Speedie, György J. Simon, Vipin Kumar 0001, Bonnie L. Westra |
AMIA | 3 |
| 2015 | Clustering Health Data to Discover EBP Interventions for Sepsis Prevention and Treatment for Health Disparities
Lisiane Pruinelli, Pranjul Yadav, Andrew Hangsleben, Kevin Schiroo, Sanjoy Dey, György J. Simon, Maribet C. McCarty, Vipin Kumar 0001, Connie White-Delaney, Michael S. Steinbach, Bonnie L. Westra |
AMIA | 6 |
| 2015 | Forensic Style Analysis with Survival TrajectoriesabstractElectronic Health Records (EHRs) consists of patient information such as demographics, medications, laboratory test results, diagnosis codes and procedures. Mining EHRs could lead to improvement in patient healthcare management as EHRs contain detailed information related to disease prognosis for large patient populations. We hypothesize that a patient's condition does not deteriorate at random, the trajectories, sequences in which diseases appear in a patient, are determined by a finite number of underlying disease mechanisms. In this work, we exploit this idea by predicting a patient's risk of mortality in the context of the metabolic syndrome by assessing which of many available trajectories a patient is following and progression along this trajectory. Implementing this idea required innovative enhancements both for the study design and also for the fitting algorithm. We propose a forensic-style study design, which aligns patients on last follow-up and measures time backwards. We modify the time-dependent covariate Cox proportional hazards model to better capture coefficients of covariate that follow a particular temporal sequence, such as trajectories. Knowledge extracted from such analysis can lead to personalized treatments, thereby forming the basis for future trajectory-centered guidelines. Pranjul Yadav, Michael S. Steinbach, Lisiane Pruinelli, Bonnie L. Westra, Connie White-Delaney, Vipin Kumar 0001, György J. Simon |
ICDM | 7 |
| 2015 | Extending Association Rule Summarization Techniques to Assess Risk of Diabetes MellitusabstractEarly detection of patients with elevated risk of developing diabetes mellitus is critical to the improved prevention and overall clinical management of these patients. We aim to apply association rule mining to electronic medical records (EMR) to discover sets of risk factors and their corresponding subpopulations that represent patients at particularly high risk of developing diabetes. Given the high dimensionality of EMRs, association rule mining generates a very large set of rules which we need to summarize for easy clinical use. We reviewed four association rule set summarization techniques and conducted a comparative evaluation to provide guidance regarding their applicability, strengths and weaknesses. We proposed extensions to incorporate risk of diabetes into the process of finding an optimal summary. We evaluated these modified techniques on a real-world prediabetic patient cohort. We found that all four methods produced summaries that described subpopulations at high risk of diabetes with each method having its clear strength. For our purpose, our extension to the Buttom-Up Summarization (BUS) algorithm produced the most suitable summary. The subpopulations identified by this summary covered most high-risk patients, had low overlap and were at very high risk of diabetes. György J. Simon, Pedro J. Caraballo, Terry M. Therneau, Steven S. Cha, Regina Castro, Peter W. Li |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Divisive Hierarchical Clustering towards Identifying Clinically Significant Pre-Diabetes Subpopulations
Era Kim, Wonsuk Oh, David S. Pieczkiewicz, Regina Castro, Pedro J. Caraballo, György J. Simon |
AMIA | 6 |
| 2014 | Employment of Penalized Models to Predict Cognitive Domain Performance Using Proteomics
Shauna M. Overgaard, S. Charles Schulz, György J. Simon |
AMIA | 3 |
| 2014 | Data Mining Methodologies to Discover Best practices for Diabetic Patients with Health Disparities
Lisiane Pruinelli, Sanjoy Dey, György J. Simon, Pranjul Yadav, Andrew Hangsleben, Katherine Hauwiller, Vipin Kumar 0001, Connie White-Delaney, Michael S. Steinbach, Bonnie L. Westra |
AMIA | 3 |
| 2014 | Mining Interpretable and Predictive Diagnosis Codes from Multi-source Electronic Health RecordsabstractMining patterns from electronic health-care records (EHR) can potentially lead to better and more cost-effective treatments. We aim to find the groups of ICD-9 diagnosis codes from EHRs that can predict the improvement of urinary incontinence of home health care (HHC) patients and also are interpretable to domain experts. In this paper, we propose two approaches for increasing the interpretability of the obtained groups of ICD-9 codes. First, we incorporate prior information available from clinical domain knowledge using the clinical classification system (CCS). Second, we incorporate additional types of clinical information for the same patients, such as demographic, behavioral, physiological, and psycho-social variables available from survey questions during the hospital visits. Finally, we develop a hybrid framework that can combine both prior information and the data-driven clinical information in the predictive model framework. Our results obtained from a large-scale EHR data set show that the hybrid framework enhances clinical interpretability as compared to the baseline model obtained from ICD-9 codes only, while achieving almost the same predictive capability. Sanjoy Dey, György J. Simon, Bonnie L. Westra, Michael S. Steinbach, Vipin Kumar 0001 |
SDM | 2 |
| 2014 | A Gentle Introduction to Support Vector Machines in Biomedicine, Alexander Statnikov, Constantin F. Aliferis, Douglas P. Hardin, Isabelle Guyon. World Scientific (2013). Vols. 1(200 p.) and 2 (212 p.)
György J. Simon, Genevieve B. Melton |
J. Biomed. Informatics | 1 |
| 2013 | Data Mining to Predict Mobility Outcomes for Older Adults Receiving Home Health Care
Sanjoy Dey, Jeremy Weed, Joanna Fakhoury, Jacob Cooner, György J. Simon, Michael S. Steinbach, Bonnie L. Westra, Vipin Kumar 0001 |
AMIA | 5 |
| 2013 | Quantifying the Effect of Statin Use in Pre-Diabetic Phenotypes Discovered Through Association Rule Mining
John Schrom, Pedro J. Caraballo, Regina Castro, György J. Simon |
AMIA | 4 |
| 2013 | Survival Association Rule Mining Towards Type 2 Diabetes Risk Assessment
György J. Simon, John Schrom, Regina Castro, Peter W. Li, Pedro J. Caraballo |
AMIA | 1 |
| 2011 | A simple statistical model and association rule filtering for classificationabstractAssociative classification is a predictive modeling technique that constructs a classifier based on class association rules (also known as predictive association rules; PARs). PARs are association rules where the consequence of the rule is a class label. Associative classification has gained substantial research attention because it successfully joins the benefits of association rule mining with classification. These benefits include the inherent ability of association rule mining to extract high-order interactions among the predictors--an ability that many modern classifiers lack--and also the natural interpretability of the individual PARs. György J. Simon, Vipin Kumar 0001, Peter W. Li |
KDD | 1 |
| 2011 | Understanding atrophy trajectories in alzheimer's disease using association rules on MRI imagesabstractAlzheimer's disease (AD) is associated with progressive cognitive decline leading to dementia. The atrophy/loss of brain structure as seen on Magnetic Resonance Imaging (MRI) is strongly correlated with the severity of the cognitive impairment in AD. In this paper, we set out to find associations between predefined regions of the brain (regions of interest; ROIs) and the severity of the disease. Specifically, we use these associations to address two important issues in AD: (i) typical versus atypical atrophy patterns and (ii) the origin and direction of progression of atrophy, which is currently under debate. György J. Simon, Peter W. Li, Clifford R. Jack Jr., Prashanthi Vemuri |
KDD | 1 |
| 2008 | Semi-supervised approach to rapid and reliable labeling of large data setsabstractIn this paper, we propose a method, where the labeling of the data set is carried out in a semi-supervised manner with user-specified guarantees about the quality of the labeling. In our scheme, we assume that for each class, we have some heuristics available, each of which can identify instances of one particular class. The heuristics are assumed to have reasonable performance but they do not need to cover all instances of the class nor do they need to be perfectly reliable. We further assume that we have an infallible expert, who is willing to manually label a few instances. The aim of the algorithm is to exploit the cluster structure of the problem, the predictions by the imperfect heuristics and the limited perfect labels provided by the expert to classify (label) the instances of the data set with guaranteed precision (specificed by the user) with regards to each class. The specified precision is not always attainable, so the algorithm is allowed to classify some instances as dontknow. The algorithm is evaluated by the number of instances labeled by the expert, the number of dontknow instances (global coverage) and the achieved quality of the labeling. On the KDD Cup Network Intrusion data set containing 500,000 instances, we managed to label 96.6% of the instances while guaranteeing a nominal precision of 90% (with 95% confidence) by having the expert label 630 instances; and by having the expert label 1200 instances, we managed to guarantee 95% nominal precision while labeling 96.4% of the data. We also provide a case study of applying our scheme to label the network traffic collected at a large campus network. György J. Simon, Vipin Kumar 0001, Zhi-Li Zhang |
KDD | 1 |
| 2007 | Estimating False Negatives for Classification Problems with Cluster StructureabstractEstimating the number of false negatives for a classifier when the true outcome of the classification is ascertained only for a limited number of instances is an important problem, with a wide range of applications from epidemiology to computer/network security. The frequently applied method is random sampling. However, when the target (positive) class of the classification is rare, which is often the case with network intrusions and diseases, this simple method results in excessive sampling. In this paper, we propose an approach that exploits the cluster structure of the data to significantly reduce the amount of sampling needed while guaranteeing an estimation accuracy specified by the user. The basic idea is to cluster the data and divide the clusters into a set of “strata”, such that the proportion of positive instances in the stratum is very low, very high or in between, respectively. By taking advantage of the different characteristics of the strata, more efficient estimation strategies can be applied, thereby significantly reducing the amount of required sampling. We also develop a computationally efficient clustering algorithm – referred to as class-focused partitioning – which uses the (imperfect) labels predicted by the classifier as additional guidance. We evaluated our method on the KDDCup network intrusion data set. Our method achieved better precision and accuracy with a 5% sample than the best trial of simple random sampling with 40% samples. György J. Simon, Vipin Kumar 0001, Zhi-Li Zhang |
SDM | 1 |
| 2006 | Scan Detection: A Data Mining ApproachabstractA precursor to many attacks on networks is often a reconnaissance operation, more commonly referred to as a scan. Despite the vast amount of attention focused on methods for scan detection, the state-ofthe-art methods suffer from high rate of false alarms and low rate of scan detection. In this paper, we formalize the problem of scan detection as a data mining problem. We show how the network traffic data sets can be converted into a data set that is appropriate for running off-the-shelf classifiers on. Our method successfully demonstrates that data mining models can encapsulate expert knowledge to create an adaptable algorithm that can substantially outperform state-ofthe-art methods for scan detection in both coverage and precision. György J. Simon, Hui Xiong 0001, Eric Eilertson, Vipin Kumar 0001 |
SDM | 1 |