VLDB 2026 Research / reviewers in the wild / expert
Kristin P. Bennett
dblp:24/4209
· DBLP profile ↗
73ranked-venue papers
14as first author
21since 2021 · last 2026
0000-0002-8782-105XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 11 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Self-supervised contrastive pre-training for multivariate event streams
Xiao Shou, Dharmashankar Subramanian, Debarun Bhattacharjya, Kristin P. Bennett |
Neurocomputing | 5 |
| 2025 | Encoding Matters: Impact of Categorical Variable Encoding on Performance and BiasabstractEncoding categorical variables impacts model performance and can introduce bias in supervised learning, particularly affecting fairness when some groups are under-represented.We analyze the effects of different encoding methods on synthetic and real datasets to mitigate unintended model reliance on specific variables.We propose CaVaR (Categorical Variable Reliance) to quantify model reliance on variables and an Availability Index to measure CaVaR's sensitivity to partial encoding changes.A high Availability Disparity, measured by the standard deviation of the Availability Index across encodings, highlights potential bias from mixed encodings.The results suggest encoding all categorical variables uniformly, regardless of their ordinal or nominal nature, may reduce bias, with the choice guided by computational and performance considerations.1 To facilitate the reading of this paper, we added a glossary in appendix provided here: https://github.com/danielkopp4/ESANN2025Appendix.However, this paper should be readable without the appendix. Daniel Kopp, Benjamin Maudet, Lisheng Sun-Hosoya, Kristin P. Bennett |
ESANN | 4 |
| 2024 | Examining Trustworthiness of LLM-as-a-Judge Systems in a Clinical Trial Design BenchmarkabstractManual evaluation of Large Language Model (LLM) applications at scale presents significant resource challenges, making LLM-as-Judge (LaaJ) an attractive alternative. This study examines the reliability of LaaJ evaluation within CT-Bench, a benchmark for assessing LLMs’ capabilities in recommending clinical trial baseline features. LaaJ-alpha, our GPT-4o based prototype, semantically matches LLM-recommended features against reference features from clinical trials, accounting for semantic equivalence (e.g., ‘BMI’ and ‘Body Mass Index’). The system generates matched pairs and unmatched features from both sources to calculate precision, recall, and F1 scores. Laaj-alpha evaluates baseline feature recommendations across CTBench CT-Pub (100 trials) and CT-Repo (1,690 trials) for comparing results for GPT-4o and Llama-3-70B-Instruct under zero-shot and three-shot settings. Coherence checking revealed hallucinations in LaaJ-alpha’s evaluation, necessitating a post-processing correction step that yielded lower but more accurate performance metrics. Three different types of hallucination were observed. The hallucination rate provides a quantifiable coherence metric that can be systematically used to improve LaaJ reliability. Our findings underscore the challenges in developing reliable LLM evaluation methods in healthcare applications and demonstrate a potential framework for improving LaaJ systems. Corey Curran, Nafis Neehal, Keerthiram Murugesan, Kristin P. Bennett |
IEEE Big Data | 4 |
| 2024 | LLM - Based Code Generation for Querying Temporal Tabular Financial DataabstractWe examine the question of “how well large language models (LLMs) can answer questions using temporal tabular financial data by generating code?”. Leveraging advanced language models, specifically GPT-4 and Llama 3, we aim to scrutinize and compare their abilities to generate coherent and effective code for Python, R, and SQL based on natural language prompts. We design an experiment to assess the performance of LLMs on natural language prompts on a large temporal financial dataset. We created a set of queries with hand-crafted R code answers. To investigate the strengths and weaknesses of LLMs, each query was created with different factors that characterize the financial meaning of the queries and their complexity. We demonstrate how to create specific zero-shot prompts to generate code to answer natural language queries about temporal financial tabular data. We develop specific system prompts for each language to ensure they correctly answer time-oriented questions. We execute this experiment on two LLMs (GPT-4 and Llama 3), assess if the outputs produced are executable and correct, and assess the efficiency of the produced code for Python, SQL, and R. We find that while LLMs have promising performance, their performance varies greatly across the languages, models, and experimental factors. GPT-4 performs best on Python (95.2% correctness) but has significantly weaker performance on SQL (87.6% correctness) and R (79.0% correctness). Llama 3 is less successful at generating code overall, but it achieves its best results in R (71.4% correctness). A multi-factor statistical analysis of the results with respect to the defined experimental factors provides further insights into the specific areas of challenge in code generation for each LLM. Our preliminary results on this modest benchmark demonstrate a framework for developing larger, comprehensive, unique benchmarks for both temporal financial tabular data and R code generation. While Python and SQL already have benchmarks, we are filling in the gaps for R. Powerful AI agents for text-to-code generation, as demonstrated in this work, provide a critical capability required for the next-generation AI-based natural language financial intelligence systems and chatbots, directly addressing the complex challenges presented by querying temporal tabular financial data. Mohamed Lashuel, Gulrukh Kurdistan, Aaron Green 0001, John S. Erickson, Oshani Seneviratne, Kristin P. Bennett |
CIFEr | 6 |
| 2024 | DeFi Survival Analysis: Insights Into the Emerging Decentralized Financial EcosystemabstractWe propose a survival analysis approach for discovering and characterizing user behavior and risks for lending protocols in decentralized finance (DeFi). We demonstrate how to gather and prepare DeFi transaction data for survival analysis. We illustrate our approach using transactions in Aave, one of the largest lending protocols. We develop a DeFi survival analysis pipeline that first prepares transaction data for survival analysis through the selection of different index events (or transactions) and associated outcome events. Then we apply survival analysis statistical and visualization methods modified for competing risks when appropriate, such as Kaplan–Meier survival curves, cumulative incidence functions, Cox hazard regression, and Fine-Gray models for sub-distribution hazards to gain insights into usage patterns and risks within the protocol. We show how, by varying the index and outcome events as well as covariates, we can use DeFi survival analysis to answer questions like “How does loan size affect the repayment schedule of the loan?”; “How does loan size affect the likelihood that an account gets liquidated?”; “How does user behavior vary between Aave markets?”; “How has user behavior in Aave varied from quarter to quarter?” The proposed DeFi survival analysis can easily be generalized to other DeFi lending protocols. By defining appropriate index and outcome events, DeFi survival analysis can be applied to any cryptocurrency protocol with transactions. Aaron Green 0001, Michael P. Giannattasio, John S. Erickson, Oshani Seneviratne, Kristin P. Bennett |
Distributed Ledger Technol. Res. Pract. | 5 |
| 2023 | Concurrent Multi-Label Prediction in Event StreamsabstractStreams of irregularly occurring events are commonly modeled as a marked temporal point process. Many real-world datasets such as e-commerce transactions and electronic health records often involve events where multiple event types co-occur, e.g. multiple items purchased or multiple diseases diagnosed simultaneously. In this paper, we tackle multi-label prediction in such a problem setting, and propose a novel Transformer-based Conditional Mixture of Bernoulli Network (TCMBN) that leverages neural density estimation to capture complex temporal dependence as well as probabilistic dependence between concurrent event types. We also propose potentially incorporating domain knowledge in the objective by regularizing the predicted probability. To represent probabilistic dependence of concurrent event types graphically, we design a two-step approach that first learns the mixture of Bernoulli network and then solves a least-squares semi-definite constrained program to numerically approximate the sparse precision matrix from a learned covariance matrix. This approach proves to be effective for event prediction while also providing an interpretable and possibly non-stationary structure for insights into event co-occurrence. We demonstrate the superior performance of our approach compared to existing baselines on multiple synthetic and real benchmarks. Xiao Shou, Dharmashankar Subramanian, Debarun Bhattacharjya, Kristin P. Bennett |
AAAI | 5 |
| 2023 | Stress-Testing Bias Mitigation Algorithms to Understand Fairness VulnerabilitiesabstractTo address the growing concern of unfairness in Artificial Intelligence (AI), several bias mitigation algorithms have been introduced in prior research. Their capabilities are often evaluated on certain overly-used datasets without rigorously stress-testing them under simultaneous train and test distribution shifts. To address this, we investigate the fairness vulnerabilities of these algorithms across several distribution shift scenarios using synthetic data, to highlight scenarios where these algorithms do and don’t work to encourage their trustworthy use. The paper makes three important contributions. Firstly, we propose a flexible pipeline called the Fairness Auditor to systematically stress-test bias mitigation algorithms using multiple synthetic datasets with shifts. Secondly, we introduce the Deviation Metric for measuring the fairness and utility performance of these algorithms under such shifts. Thirdly, we propose an interactive reporting tool for comparing algorithmic performance across various synthetic datasets, mitigation algorithms and metrics called the Fairness Report. Karan Bhanot, Ioana Baldini, Dennis Wei, Jiaming Zeng, Kristin P. Bennett |
AIES | 5 |
| 2023 | Enabling Cross-Language Data Integration and Scalable Analytics in Decentralized FinanceabstractWith the agile development process of most academic and corporate entities, designing a robust computational back-end system that can support their ever-changing data needs is a constantly evolving challenge. We propose the implementation of a data and language-agnostic system design that handles different data schemes and sources while subsequently providing researchers and developers a way to connect to it that is supported by a vast majority of programming languages. To validate the efficacy of a system with this proposed architecture, we integrate various data sources throughout the decentralized finance (DeFi) space, specifically from DeFi lending protocols, retrieving tens of millions of data points to perform analytics through this system. We then access and process the retrieved data through several different programming languages (R-Lang, Python, and Java). Finally, we analyze the performance of the proposed architecture in relation to other high-performance systems and explore how this system performs under a high computational load. Conor Flynn, Kristin P. Bennett, John S. Erickson, Aaron Green 0001, Oshani Seneviratne |
IEEE Big Data | 2 |
| 2023 | Adversarial Auditing of Machine Learning Models under Compound ShiftabstractMachine learning (ML) models often perform differently under distribution shifts, in terms of utility, fairness, and other dimensions.We propose the Adversarial Auditor for measuring the utility and fairness performance of ML models under compound shifts of outcome and protected attributes.We use Multi-Objective Bayesian Optimization (MOBO) to account for multiple metrics and identify shifts where model performance is extreme, both good and bad.Using two case studies, we show that MOBO performed better than random and grid-based approaches in identifying scenarios by adversarially optimizing objectives, highlighting the value of such an auditor for developing fair, accurate and shift-robust models. Karan Bhanot, Dennis Wei, Ioana Baldini, Kristin P. Bennett |
ESANN | 4 |
| 2023 | Probabilistic Attention-to-Influence Neural Models for Event SequencesabstractDiscovering knowledge about which types of events influence others, using datasets of event sequences without time stamps, has several practical applications. While neural sequence models are able to capture complex and potentially long-range historical dependencies, they often lack the interpretability of simpler models for event sequence dynamics. We provide a novel neural framework in such a setting - a probabilistic attention-to-influence neural model - which not only captures complex instance-wise interactions between events but also learns influencers for each event type of interest. Given event sequence data and a prior distribution on type-wise influence, we efficiently learn an approximate posterior for type-wise influence by an attention-to-influence transformation using variational inference. Our method subsequently models the conditional likelihood of sequences by sampling the above posterior to focus attention on influencing event types. We motivate our general framework and show improved performance in experiments compared to existing baselines on synthetic data as well as real-world benchmarks, for tasks involving prediction and influencing set identification. Xiao Shou, Debarun Bhattacharjya, Dharmashankar Subramanian, Oktie Hassanzadeh, Kristin P. Bennett |
ICML | 6 |
| 2023 | Pairwise Causality Guided Transformers for Event SequencesabstractAlthough pairwise causal relations have been extensively studied in observational longitudinal analyses across many disciplines, incorporating knowledge of causal pairs into deep learning models for temporal event sequences remains largely unexplored. In this paper, we propose a novel approach for enhancing the performance of transformer-based models in multivariate event sequences by injecting pairwise qualitative causal knowledge such as `event Z amplifies future occurrences of event Y'. We establish a new framework for causal inference in temporal event sequences using a transformer architecture, providing a theoretical justification for our approach, and show how to obtain unbiased estimates of the proposed measure. Experimental results demonstrate that our approach outperforms several state-of-the-art models in terms of prediction accuracy by effectively leveraging knowledge about causal pairs.
We also consider a unique application where we extract knowledge around sequences of societal events by generating them from a large language model, and demonstrate how a causal knowledge graph can help with event prediction in such sequences.
Overall, our framework offers a practical means of improving the performance of transformer-based models in multivariate event sequences by explicitly exploiting pairwise causal information. Xiao Shou, Debarun Bhattacharjya, Dharmashankar Subramanian, Oktie Hassanzadeh, Kristin P. Bennett |
NeurIPS | 6 |
| 2023 | Planning and Monitoring Equitable Clinical Trial Enrollment Using Goal ProgrammingabstractRandomized clinical trial (RCT) studies are the gold standard for scientific evidence on treatment benefits to patients. RCT outcomes may not be generalizable to clinical practice if the trial population is not representative of the patients for which the treatment is intended. Specifically, enrollment plans may not adequately include groups of patients with protected attributes, such as gender, race, or ethnicity. Inequities in RCTs are a major concern for funding agencies such as the National Institutes of Health (NIH) and for policy makers. We address this challenge by proposing a goal-programming approach, explicitly integrating measurable enrollment goals, to design equitable enrollment plans for RCTs. We evaluate our model in both single and multisite settings using the enrollment criteria and study population from the Systolic Blood Pressure Intervention Trial (SPRINT) study. Our model can successfully generate equitable enrollment plans that satisfy multiple goals such as sample representativeness and minimum total financial cost. Our model can detect deviations from a target plan during the enrollment process and update the plan to reduce deviations in the remaining process. Finally, through appropriate site selection in the planning stage, the model can demonstrate the possibility of enrolling a nationally representative study population if geographic constraints exist in multisite recruitment (e.g., clinical centers in a particular region). Our model can be used to prospectively produce and retrospectively evaluate how equitable enrollment plans are based on subjects' protected attributes, and it allows researchers to provide justifications on validity of scientific analysis and evaluation of subgroup disparities. Miao Qi, Amar K. Das, Kristin P. Bennett |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | An Ontology for Fairness MetricsabstractRecent research has revealed that many machine-learning models and the datasets they are trained on suffer from various forms of bias, and a large number of different fairness metrics have been created to measure this bias. However, determining which metrics to use, as well as interpreting their results, is difficult for a non-expert due to a lack of clear guidance and issues of ambiguity or alternate naming schemes between different research papers. To address this knowledge gap, we present the Fairness Metrics Ontology (FMO), a comprehensive and extensible knowledge resource that defines each fairness metric, describes their use cases, and details the relationships between them. We include additional concepts related to fairness and machine learning models, enabling the representation of specific fairness information within a resource description framework (RDF) knowledge graph. We evaluate the ontology by examining the process of how reasoning-based queries to the ontology were used to guide the fairness metric-based evaluation of a synthetic data model. Jade S. Franklin, Karan Bhanot, Mohamed F. Ghalwash, Kristin P. Bennett, Jamie P. McCusker, Deborah L. McGuinness |
AIES | 4 |
| 2022 | Evaluating Fairness of Synthetic Healthcare Data Models
Karan Bhanot, Ioana Baldini, Dennis Wei, Jiaming Zeng, Kristin P. Bennett |
AMIA | 5 |
| 2022 | Subpopulation Analysis in Causal Inference: A Healthcare Case StudyabstractTreatment interventions are usually targeted to improve a specific outcome on as elected group of patients who are eligible to receive the treatment. The success of such treatments is determined by the post-intervention treatment effect on the population under consideration. There are cases when the treatment group contains multiple categories of eligible populations, with various effects, especially when the study’s criteria are loosely defined. I n such s tudies (non-targeted trials) non-eligible subjects may be treated, producing heterogeneous treatment effects within the treated group. Inferring the effectiveness of the treatment under this scenario is difficult since the average treatment effect on the treated is a combination of multiple effect levels. This can bias the resulting conclusion of the causal studies. We propose an end-to-end framework based on matching and unsupervised clustering for extracting population sub-groups based on their effect levels. We demonstrate our methods on a real-world healthcare application, highlighting the value of subpopulation analysis for recovering multiple effect groups. Georgios Mavroudeas, Nafis Neehal, Jason Kuruzovich, Kristin P. Bennett, Malik Magdon-Ismail |
BIBM | 4 |
| 2022 | Investigating synthetic medical time-series resemblance
Karan Bhanot, Joseph Pedersen, Isabelle Guyon, Kristin P. Bennett |
Neurocomputing | 4 |
| 2021 | Planning Equitable Clinical Trial Enrollment using Integer Programming
Miao Qi, Amar K. Das, Kristin P. Bennett |
AMIA | 3 |
| 2021 | Predictive Modeling for Complex Care ManagementabstractComplex care management (CCM) or hot spotting programs identify and manage high-need/high-cost patients, improving long-term health quality and medical costs. Typically, physicians refer patients to CCM. Despite strict guidelines to ensure that eligible patients are placed in appropriate programs, such a provider-based approach is limited by provider-capacity and the narrow view of a patient that a provider sees. We propose an ML workflow to augment the provider-based approach, that can flag patients who are suited to CCM. Our predictor uses a global view of a patient’s entire history across multiple providers and time to identify high-risk individuals from among all the individuals in a matter of seconds. On a monthly basis, we evaluate our predictions against physician referrals. In the test dataset, 41% of the top-500 highest risk individuals found by our model were referred to CCM by a physician at some point in the 6-month window following our prediction (top-500 is a parameter that can be set to match the CCM program’s capacity). Of those who were not referred in the 6-month window, 30% were referred at some time in their trajectory. The remaining false positives had a greater than 95% similarity when compared to true positive physician referrals in terms of cost profiles (both prior to referral and after referral) and patient profile. This remarkable similarity suggests that our machine learning predictor can identify new candidates for complex care management and/or predict referrals before a physician has an opportunity to do so. Georgios Mavroudeas, Nafis Neehal, Xiao Shou, Malik Magdon-Ismail, Jason Kuruzovich, Kristin P. Bennett |
BIBM | 6 |
| 2021 | Quantifying Resemblance of Synthetic Medical Time-SeriesabstractAccess to medical data is often restricted due to privacy laws e.g.HIPAA and GDPR.We address the viability of substituting real data with synthetic data to protect privacy while maintaining utility.Medical data records are fundamentally longitudinal, with one patient having multiple health events influenced by covariates like gender, age etc. Synthesis of medical data, hence, falls under time-series generative modeling.We demonstrate methods to measure synthetic medical time-series quality on datasets from previously published synthetic data research.We deploy four time-series metrics to quantify resemblance in synthetic and real covariate plots while comparing baseline data generation methods. Karan Bhanot, Saloni Dash, Joseph Pedersen, Isabelle Guyon, Kristin P. Bennett |
ESANN | 5 |
| 2021 | Causal Inference for Event Pairs in Multivariate Point ProcessesabstractCausal inference and discovery from observational data has been extensively studied across multiple fields. However, most prior work has focused on independent and identically distributed (i.i.d.) data. In this paper, we propose a formalization for causal inference between pairs of event variables in multivariate recurrent event streams by extending Rubin's framework for the average treatment effect (ATE) and propensity scores to multivariate point processes. Analogous to a joint probability distribution representing i.i.d. data, a multivariate point process represents data involving asynchronous and irregularly spaced occurrences of various types of events over a common timeline. We theoretically justify our point process causal framework and show how to obtain unbiased estimates of the proposed measure. We conduct an experimental investigation using synthetic and real-world event datasets, where our proposed causal inference framework is shown to exhibit superior performance against a set of baseline pairwise causal association scores. Dharmashankar Subramanian, Debarun Bhattacharjya, Xiao Shou, Nicholas Mattei, Kristin P. Bennett |
NeurIPS | 6 |
| 2021 | MOSAIC: a joint modeling methodology for combined circadian and non-circadian analysis of multi-omics dataabstractMOTIVATION: Circadian rhythms are approximately 24-h endogenous cycles that control many biological functions. To identify these rhythms, biological samples are taken over circadian time and analyzed using a single omics type, such as transcriptomics or proteomics. By comparing data from these single omics approaches, it has been shown that transcriptional rhythms are not necessarily conserved at the protein level, implying extensive circadian post-transcriptional regulation. However, as proteomics methods are known to be noisier than transcriptomic methods, this suggests that previously identified arrhythmic proteins with rhythmic transcripts could have been missed due to noise and may not be due to post-transcriptional regulation. RESULTS: To determine if one can use information from less-noisy transcriptomic data to inform rhythms in more-noisy proteomic data, and thus more accurately identify rhythms in the proteome, we have created the Multi-Omics Selection with Amplitude Independent Criteria (MOSAIC) application. MOSAIC combines model selection and joint modeling of multiple omics types to recover significant circadian and non-circadian trends. Using both synthetic data and proteomic data from Neurospora crassa, we showed that MOSAIC accurately recovers circadian rhythms at higher rates in not only the proteome but the transcriptome as well, outperforming existing methods for rhythm identification. In addition, by quantifying non-circadian trends in addition to circadian trends in data, our methodology allowed for the recognition of the diversity of circadian regulation as compared to non-circadian regulation. AVAILABILITY AND IMPLEMENTATION: MOSAIC's full interface is available at https://github.com/delosh653/MOSAIC. An R package for this functionality, mosaic.find, can be downloaded at https://CRAN.R-project.org/package=mosaic.find. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hannah De Los Santos, Kristin P. Bennett, Jennifer M. Hurley |
Bioinform. | 2 |
| 2020 | Medical Time-Series Data Generation Using Generative Adversarial Networks
Saloni Dash, Andrew Yale, Isabelle Guyon, Kristin P. Bennett |
AIME | 4 |
| 2020 | MortalityMinder: A Web Tool for Visualizing and Investigating Social Determinants of Premature Mortality in the United States
Kristin P. Bennett, Lilian Ngweta, Karan Bhanot, John S. Erickson |
AMIA | 1 |
| 2020 | Visualizing Inequities in Clinical Trials using ML Fairness Metrics
Miao Qi, Owen Cahan, Morgan Foreman, Dan Gruen, Amar K. Das, Kristin P. Bennett |
AMIA | 6 |
| 2020 | ECHO: an application for detection and analysis of oscillators identifies metabolic regulation on genome-wide circadian outputabstractMOTIVATION: Time courses utilizing genome scale data are a common approach to identifying the biological pathways that are controlled by the circadian clock, an important regulator of organismal fitness. However, the methods used to detect circadian oscillations in these datasets are not able to accommodate changes in the amplitude of the oscillations over time, leading to an underestimation of the impact of the clock on biological systems. RESULTS: We have created a program to efficaciously identify oscillations in large-scale datasets, called the Extended Circadian Harmonic Oscillator application, or ECHO. ECHO utilizes an extended solution of the fixed amplitude oscillator that incorporates the amplitude change coefficient. Employing synthetic datasets, we determined that ECHO outperforms existing methods in detecting rhythms with decreasing oscillation amplitudes and in recovering phase shift. Rhythms with changing amplitudes identified from published biological datasets revealed distinct functions from those oscillations that were harmonic, suggesting purposeful biologic regulation to create this subtype of circadian rhythms. AVAILABILITY AND IMPLEMENTATION: ECHO's full interface is available at https://github.com/delosh653/ECHO. An R package for this functionality, echo.find, can be downloaded at https://CRAN.R-project.org/package=echo.find. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hannah De Los Santos, Emily J. Collins, Catherine Mann, April Sagan, Meaghan S. Jankowski, Kristin P. Bennett, Jennifer M. Hurley |
Bioinform. | 6 |
| 2020 | Generation and evaluation of privacy preserving synthetic health data
Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavão, Kristin P. Bennett |
Neurocomputing | 6 |
| 2020 | A Precision Environment-Wide Association Study of Hypertension via Supervised Cadre ModelsabstractWe consider the problem in precision health of grouping people into subpopulations based on their degree of vulnerability to a risk factor. These subpopulations cannot be discovered with traditional clustering techniques because their quality is evaluated with a supervised metric: The ease of modeling a response variable for observations within them. Instead, we apply the more appropriate supervised cadre model (SCM). We extend the SCM formalism so that it may be applied to multivariate regression and binary classification problems and develop a way to use conditional entropy to assess the confidence in the process by which a subject is assigned their cadre. Using the SCM, we generalize the environment-wide association study (EWAS) to be able to model heterogeneity in population risk. In our EWAS, we consider more than 200 environmental exposure factors and find their association with diastolic blood pressure, systolic blood pressure, and hypertension. This requires adapting the SCM to be applicable to data generated by a complex survey design. After correcting for false positives, we found 25 exposure variables that had a significant association with at least one of our response variables. Eight of these were significant for a discovered subpopulation but not for the overall population. Some of these associations have been identified by previous researchers, whereas others appear to be novel. We examine discovered subpopulations in detail, finding that they are interpretable and suggestive of further research questions. Alexander New, Kristin P. Bennett |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Artificial Intelligence for Public HealthabstractIn this talk, we examine artificial intelligence approaches for extracting actionable insights from health care data in order to improve public health. Our goal is to simultaneously identify subpopulations with distinct health risks and health trajectories and find the distinct risk factors or determinants associated each subpopulation. These determinants can then be used treatments, programs, and policies in order to reduce mortality and comorbidity and provide more efficient healthcare. We examine novel cadre machine learning approaches that combine predictive neural network modeling with more traditional statistical epidemiology methods for risk and survival analyses. We embed the cadre methods into a Semantically Targeted Analytics (Semantalytics) System that combines semantics, inference, automatic machine learning, and explainable AI. The AI system translate the public health questions to an analysis plan, prepares data, conducts analysis and reports results with visualization and text. We demonstrated these approach on public health care surveillance datasets and electronic medical records. The award winning “MortalityMinder” app examines the social determinants of “Deaths Despair” (deaths from suicide and substance abuse) and other causes of mortality that are unexpectedly rising in the United States. Other applications include association of environment toxins associated with diseases, high needs patient management for a health management organization, and emergency department readmissions. We conclude with the discussion of the open challenges to create population health AI systems that can transform health care questions into data-driven actionable-insights on the fly. Kristin P. Bennett |
BIBM | 1 |
| 2019 | Supervised Mixture Models for Population HealthabstractWe examine a machine learning approach for deriving insights from observational healthcare data in order to improve public health. Our goal is to simultaneously identify patient subpopulations with differing health risks and find the distinct risk factors or determinants associated with each subpopulation. Here, we develop a supervised Gaussian Mixture Model (GMM) approach for subpopulation modeling that combines GMMs with L1-logistic regression. We demonstrate the approach on an analysis of high cost drivers of Medicaid expenditures for inpatient stays associated with Newborn, Pregnancy, and Circulatory Systems diagnostic categories. These conditions were chosen because they had the highest total inpatient expenditures in New York State (NYS) in 2016. When compared with state-of-the-art learning methods (random forests, boosting, deep learning), our approach provides comparable prediction performance but also extracts insightful explanations of the subpopulation structure and risk factors within each subpopulation. Sequentially applying unsupervised learning methods and then applying logistic regression fails to yield equally meaningful results: the unsupervised subpopulations are homogeneous and moderately predictable, while some of our subpopulations are highly predictable with easy-to-identify drivers of cost. Focusing on newborns, we unveil subpopulations indicative of the landscape of healthcare in NYS: about 90% of the discharges are healthy New York City babies and about 1% are costly complex cases. Subpopulations indicate regional disparities: for example newborns from Central, Southern and Western NY are of higher risk for high-cost stays associated with substance abuse. The results indicate the promise of the approach for future population health studies based on electronic health care records. Xiao Shou, Georgios Mavroudeas, Alexander New, Kofi Arhin, Jason Kuruzovich, Malik Magdon-Ismail, Kristin P. Bennett |
BIBM | 7 |
| 2019 | Privacy Preserving Synthetic Health Data
Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavão, Kristin P. Bennett |
ESANN | 6 |
| 2019 | Making Study Populations Visible Through Knowledge Graphs
Shruthi Chari, Miao Qi, Nkechinyere Agu, Oshani Seneviratne, Jamie P. McCusker, Kristin P. Bennett, Amar K. Das, Deborah L. McGuinness |
ISWC (2) | 6 |
| 2019 | Biases in feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Julie Josse, Mehreen Saeed, Isabelle Guyon |
Neurocomputing | 3 |
| 2018 | Analysis of imputation bias for feature selection with missing data
Borja Seijo-Pardo, Amparo Alonso-Betanzos, Kristin P. Bennett, Verónica Bolón-Canedo, Isabelle Guyon, Julie Josse, Mehreen Saeed |
ESANN | 3 |
| 2018 | Cadre Modeling: Simultaneously Discovering Subpopulations and Predictive ModelsabstractWe consider the problem in regression analysis of identifying subpopulations that exhibit different patterns of response, where each subpopulation requires a different underlying model. Unlike statistical cohorts, these subpopulations are not known a priori; thus, we refer to them as cadres. When the cadres and their associated models are interpretable, modeling leads to insights about the subpopulations and their associations with the regression target. We introduce a discriminative model that simultaneously learns cadre assignment and target-prediction rules. Sparsity-inducing priors are placed on the model parameters, under which independent feature selection is performed for both the cadre assignment and targetprediction processes. We learn models using adaptive step size stochastic gradient descent, and we assess cadre quality with bootstrapped sample analysis. We present simulated results showing that, when the true clustering rule does not depend on the entire set of features, our method significantly outperforms methods that learn subpopulation-discovery and targetprediction rules separately. In a materials-by-design case study, our model provides state-of-the-art prediction of polymer glass transition temperature. Importantly, the method identifies cadres of polymers that respond differently to structural perturbations, thus providing design insight for targeting or avoiding specific transition temperature ranges. It identifies chemically meaningful cadres, each with interpretable models. Further experimental results show that cadre methods have generalization that is competitive with linear and nonlinear regression models and can identify robust subpopulations. Alexander New, Curt M. Breneman, Kristin P. Bennett |
IJCNN | 3 |
| 2018 | Knowledge Integration for Disease Characterization: A Breast Cancer Example
Oshani Seneviratne, Sabbir M. Rashid, Shruthi Chari, Jamie P. McCusker, Kristin P. Bennett, James A. Hendler, Deborah L. McGuinness |
ISWC (2) | 5 |
| 2016 | Wind turbine fault prediction using soft label SVMabstractIn this paper, we address the problem of predicting wind turbine electrical subsystem fault using time series data obtained from multiple sensors on wind turbine. While considering this as a time series classification problem, we are facing with the challenge that there is no explicit label information regarding the temporal location and duration of symptoms of the fault. Besides, significant data variation caused by both external and internal factors make the identification of change point non-trivial. To address these challenges, we propose a soft label SVM method where the probability of fault instead of binary label is used to train classifier to handle the uncertainty in label information. The probability is determined using temporal information of fault instances. We consider this as a weakly supervised learning problem. To handle large variation within data, we perform customized normalization on different sensor data based on their physical meanings and relationships. Finally, we evaluate our method on 38 different forced outage instances. The experiment on real SCADA data obtained from wind turbines show promising results where we can predict the triggering of fault 18 hours beforehand with an average AUC value 0.91. Rui Zhao 0015, Md. Ridwan Al Iqbal, Kristin P. Bennett |
ICPR | 3 |
| 2015 | Design of the 2015 ChaLearn AutoML challengeabstractChaLearn is organizing the Automatic Machine Learning (AutoML) contest for IJCNN 2015, which challenges participants to solve classification and regression problems without any human intervention. Participants' code is automatically run on the contest servers to train and test learning machines. However, there is no obligation to submit code; half of the prizes can be won by submitting prediction results only. Datasets of progressively increasing difficulty are introduced throughout the six rounds of the challenge. (Participants can enter the competition in any round.) The rounds alternate phases in which learners are tested on datasets participants have not seen, and phases in which participants have limited time to tweak their algorithms on those datasets to improve performance. This challenge will push the state of the art in fully automatic machine learning on a wide range of real-world problems. The platform will remain available beyond the termination of the challenge. Isabelle Guyon, Kristin P. Bennett, Gavin C. Cawley, Hugo Jair Escalante, Sergio Escalera, Tin Kam Ho, Núria Macià, Bisakha Ray, Mehreen Saeed, Alexander R. Statnikov, Evelyne Viegas |
IJCNN | 2 |
| 2012 | Fast Bundle Algorithm for Multiple-Instance LearningabstractWe present a bundle algorithm for multiple-instance classification and ranking. These frameworks yield improved models on many problems possessing special structure. Multiple-instance loss functions are typically nonsmooth and nonconvex, and current algorithms convert these to smooth nonconvex optimization problems that are solved iteratively. Inspired by the latest linear-time subgradient-based methods for support vector machines, we optimize the objective directly using a nonconvex bundle method. Computational results show this method is linearly scalable, while not sacrificing generalization accuracy, permitting modeling on new and larger data sets in computational chemistry and other applications. This new implementation facilitates modeling with kernels. Charles Bergeron, Gregory M. Moore, Jed Zaretzki, Curt M. Breneman, Kristin P. Bennett |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2012 | Signal Timing Estimation Using Sample Intersection Travel TimesabstractSignal timing information is important in signal operations and signal/arterial performance measurement. Such information, however, may not be available for wide areas. This imposes difficulty, particularly for real-time signal/arterial performance measurement and traffic information provisions that have received much attention recently. We study, in this paper, the possibility of using intersection travel times, i.e., those collected between upstream and downstream locations of an intersection, to estimate signal timing parameters. The method contains three steps: 1) cycle breaking that determines whether a new cycle starts; 2) exact cycle boundary detection that determines when exactly a cycle starts or ends; and 3) effective red (or green) time estimation that estimates the actual duration of the red (or green) time. The proposed method is a combination of traffic flow theory and learning/estimation algorithms and can be used to estimate the cycle-by-cycle signal timing parameters for a specific movement of a signal. The method is tested using data from microscopic simulation, field experiments, and next-generation simulation with promising results. Peng Hao 0001, Xuegang Ban, Kristin P. Bennett, Zhanbo Sun |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2011 | Data-Driven Insights into Deletions of Mycobacterium tuberculosis Complex Chromosomal DR Region Using SpoligoforestsabstractBiomarkers of Mycobacterium tuberculosis complex (MTBC) mutate over time. Among the biomarkers of MTBC, spacer oligonucleotide type (spoligotype) and Mycobacterium Interspersed Repetitive Unit (MIRU) patterns are commonly used to genotype clinical MTBC strains. In this study, we present an evolution model of spoligotype rearrangements using MIRU patterns to disambiguate the ancestors of spoligotypes, in a large patient dataset from the United States Centers for Disease Control and Prevention (CDC). Based on the contiguous deletion assumption and rare observation of convergent evolution, we first generate the most parsimonious forest of spoligotypes, called a spoligoforest, using three genetic distance measures. An analysis of topological attributes of the spoligoforest and number of variations at the direct repeat (DR) locus of each strain reveals interesting properties of deletions in the DR region. First, we compare our mutation model to existing mutation models of spoligotypes and find that our mutation model produces as many within-lineage mutation events as other models, with slightly higher segregation accuracy. Second, based on our mutation model, the number of descendant spoligotypes follows a power law distribution. Third, contrary to prior studies, the power law distribution does not plausibly fit to the mutation length frequency. Finally, the total number of mutation events at consecutive DR loci follows a bimodal distribution, which results in accumulation of shorter deletions in the DR region. The two modes are spacers 13 and 40, which are hotspots for chromosomal rearrangements. The change point in the bimodal distribution is spacer 34, which is absent in most MTBC strains. This bimodal separation results in accumulation of shorter deletions, which explains why a power law distribution is not a plausible fit to the mutation length frequency. Cagri Ozcaglar, Amina Shabbeer, Natalia Kurepina, Bülent Yener, Kristin P. Bennett |
BIBM | 5 |
| 2011 | Model selection for primal SVM
Gregory M. Moore, Charles Bergeron, Kristin P. Bennett |
Mach. Learn. | 3 |
| 2010 | Examining the sublineage structure of Mycobacterium tuberculosis complex strains with multiple-biomarker tensorsabstractStrains of the Mycobacterium tuberculosis complex (MTBC) can be classified into coherent lineages of similar traits based on their genotype. We present a tensor clustering framework to group MTBC strains into sublineages of the known major lineages based on two biomarkers: spacer oligonucleotide type (spoligotype) and mycobacterial interspersed repetitive units (MIRU). We represent genotype information of MTBC strains in a high-dimensional array in order to include information about spoligotype, MIRU, and their coexistence using multiple-biomarker tensors. We use multiway models to transform this multidimensional data about the MTBC strains into two-dimensional arrays and use the resulting score vectors in a stable partitive clustering algorithm to classify MTBC strains into sublineages. We validate clusterings using cluster stability and accuracy measures, and find stabilities of each cluster. Based on validated clustering results, we present a sublineage structure of MTBC strains and compare it to the sublineage structures of SpolDB4 and MIRU-VNTRplus. Cagri Ozcaglar, Amina Shabbeer, Scott L. Vandenberg, Bülent Yener, Kristin P. Bennett |
BIBM | 5 |
| 2010 | Online Knowledge-Based Support Vector Machines
Gautam Kunapuli, Kristin P. Bennett, Amina Shabbeer, Richard Maclin, Jude W. Shavlik |
ECML/PKDD (2) | 2 |
| 2010 | A conformal Bayesian network for classification of Mycobacterium tuberculosis complex lineagesabstractBACKGROUND: We present a novel conformal Bayesian network (CBN) to classify strains of Mycobacterium tuberculosis Complex (MTBC) into six major genetic lineages based on two high-throuput biomarkers: mycobacterial interspersed repetitive units (MIRU) and spacer oligonucleotide typing (spoligotyping). MTBC is the causative agent of tuberculosis (TB), which remains one of the leading causes of disease and morbidity world-wide. DNA fingerprinting methods such as MIRU and spoligotyping are key components in the control and tracking of modern TB. RESULTS: CBN is designed to exploit background knowledge about MTBC biomarkers. It can be trained on large historical TB databases of various subsets of MTBC biomarkers. During TB control efforts not all biomarkers may be available. So, CBN is designed to predict the major lineage of isolates genotyped by any combination of the PCR-based typing methods: spoligotyping and MIRU typing. CBN achieves high accuracy on three large MTBC collections consisting of over 34,737 isolates genotyped by different combinations of spoligotypes, 12 loci of MIRU, and 24 loci of MIRU. CBN captures distinct MIRU and spoligotype signatures associated with each lineage, explaining its excellent performance. Visualization of MIRU and spoligotype signatures yields insight into both how the model works and the genetic diversity of MTBC. CONCLUSIONS: CBN conforms to the available PCR-based biological markers and achieves high performance in identifying major lineages of MTBC. The method can be readily extended as new biomarkers are introduced for TB tracking and control. An online tool (http://www.cs.rpi.edu/~bennek/tbinsight/tblineage) makes the CBN model available for TB control and research efforts. Minoo Aminian, Amina Shabbeer, Kristin P. Bennett |
BMC Bioinform. | 3 |
| 2009 | Determination of Major Lineages of Mycobacterium tuberculosis Complex Using Mycobacterial Interspersed Repetitive UnitsabstractWe present a novel Bayesian network (BN) to classify strains of Mycobacterium tuberculosis Complex (MTBC) into six major genetic lineages using mycobacterial interspersed repetitive units (MIRUs), a high-throughput biomarker. MTBC is the causative agent of tuberculosis (TB), which remains one of the leading causes of disease and morbidity world-wide. DNA fingerprinting methods such as MIRU are key components of modern TB control and tracking. The BN achieves high accuracy on four large MTBC genotype collections consisting of over 4700 distinct 12-loci MIRU genotypes. The BN captures distinct MIRU signatures associated with each lineage, explaining the excellent performance of the BN. The errors in the BN support the need for additional biomarkers such as the expanded 24-loci MIRU used in CDC genotyping labs since May 2009. The conditional independence assumption of each locus given the lineage makes the BN easily extensible to additional MIRU loci and other biomarkers. Minoo Aminian, Amina Shabbeer, Kristin P. Bennett |
BIBM | 3 |
| 2008 | Multiple instance rankingabstractThis paper introduces a novel machine learning model called multiple instance ranking (MIRank) that enables ranking to be performed in a multiple instance learning setting. The motivation for MIRank stems from the hydrogen abstraction problem in computational chemistry, that of predicting the group of hydrogen atoms from which a hydrogen is abstracted (removed) during metabolism. The model predicts the preferred hydrogen group within a molecule by ranking the groups, with the ambiguity of not knowing which hydrogen atom within the preferred group is actually abstracted. This paper formulates MIRank in its general context and proposes an algorithm for solving MIRank problems using successive linear programming. The method outperforms multiple instance classification models on several real and synthetic datasets. Charles Bergeron, Jed Zaretzki, Curt M. Breneman, Kristin P. Bennett |
ICML | 4 |
| 2008 | Guest editor's introduction: special issue on inductive transfer learning
Daniel L. Silver, Kristin P. Bennett |
Mach. Learn. | 2 |
| 2007 | Typing Staphylococcus aureus Using the spa Gene and Novel Distance MeasuresabstractWe developed an approach for identifying groups or families of Staphylococcus aureus bacteria based on genotype data. With the emergence of drug resistant strains, S. aureus represents a significant human health threat. Identifying the family types efficiently and quickly is crucial in community settings. Here, we develop a hybrid sequence algorithm approach to type this bacterium using only its spa gene. Two of the sequence algorithms we used are well established, while the third, the Best Common Gap-Weighted Sequence (BCGS), is novel. We combined the sequence algorithms with a weighted match/mismatch algorithm for the spa sequence ends. Normalized similarity scores and distances between the sequences were derived and used within unsupervised clustering methods. The resulting spa groupings correlated strongly with the groups defined by the well-established Multi locus sequence typing (MLST) method. Spa typing is preferable to MLST typing which types seven genes instead of just one. Furthermore, our spa clustering methods can be fine-tuned to be more discriminative than MLST, identifying new strains that the MLST method may not. Finally, we performed a multidimensional scaling of our distance matrices to visualize the relationship between isolates. The proposed methodology provides a promising new approach to molecular epidemiology. Phaedra Agius, Barry Kreiswirth, Steve Naidich, Kristin P. Bennett |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2007 | Introduction to special issue ACM SIGKDD 2006abstractNo abstract available. Roberto J. Bayardo, Kristin P. Bennett, Gautam Das 0001, Dimitrios Gunopulos, Johannes Gunopulos |
ACM Trans. Knowl. Discov. Data | 2 |
| 2006 | Model Selection via Bilevel OptimizationabstractA key step in many statistical learning methods used in machine learning involves solving a convex optimization problem containing one or more hyper-parameters that must be selected by the users. While cross validation is a commonly employed and widely accepted method for selecting these parameters, its implementation by a grid-search procedure in the parameter space effectively limits the desirable number of hyper-parameters in a model, due to the combinatorial explosion of grid points in high dimensions. This paper proposes a novel bilevel optimization approach to cross validation that provides a systematic search of the hyper-parameters. The bilevel approach enables the use of the state-of-the-art optimization methods and their well-supported softwares. After introducing the bilevel programming approach, we discuss computational methods for solving a bilevel cross-validation program, and present numerical results to substantiate the viability of this novel approach as a promising computational tool for model selection in machine learning. Kristin P. Bennett, Xiaoyun Ji, Gautam Kunapuli, Jong-Shi Pang |
IJCNN | 1 |
| 2006 | The Interplay of Optimization and Machine Learning ResearchabstractThe fields of machine learning and mathematical programming are increasingly intertwined. Optimization problems lie at the heart of most machine learning approaches. The Special Topic on Machine Learning and Large Scale Optimization examines this interplay. Machine learning researchers have embraced the advances in mathematical programming allowing new types of models to be pursued. The special topic includes models using quadratic, linear, second-order cone, semi-definite, and semi-infinite programs. We observe that the qualities of good optimization algorithms from the machine learning and optimization perspectives can be quite different. Mathematical programming puts a premium on accuracy, speed, and robustness. Since generalization is the bottom line in machine learning and training is normally done off-line, accuracy and small speed improvements are of little concern in machine learning. Machine learning prefers simpler algorithms that work in reasonable computational time for specific classes of problems. Reducing machine learning problems to well-explored mathematical programming classes with robust general purpose optimization codes allows machine learning researchers to rapidly develop new techniques. In turn, machine learning presents new challenges to mathematical programming. The special issue include papers from two primary themes: novel machine learning models and novel optimization approaches for existing models. Many papers blend both themes, making small changes in the underlying core mathematical program that enable the develop of effective new algorithms. Kristin P. Bennett, Emilio Parrado-Hernández |
J. Mach. Learn. Res. | 1 |
| 2005 | Support Vector Machines and Other Kernel Methods
Kristin P. Bennett |
FUZZ-IEEE | 1 |
| 2005 | Kernelized set-membership approach to nonlinear adaptive filteringabstractIn linear filtering, the set-membership normalized least mean squares (SM-NLMS) algorithm has been shown to exhibit desirable features of selective update and optimized variable step size. In this paper, a kernel approach to the SM-NLMS algorithm is presented that makes it feasible to address nonlinear problems. An online greedy approximation technique to achieve sparsity is discussed. Simulation results are presented for two practical problems: equalization of nonlinear inter-symbol interference (ISI) channels and predistortion of nonlinear high power amplifiers (HPA). Amaresh V. Malipatil, Yih-Fang Huang, Srinivas Andra, Kristin P. Bennett |
ICASSP (4) | 4 |
| 2004 | Decision-tree learning in dwell point policies in autonomous vehicle storage and retrieval systems (AVSRS)abstractAutonomous vehicle storage and retrieval system (A VSRS) is a new material handling technology for unit load storage and retrieval. The dwell point issue is one of important aspects to increase the throughput. We examine two dwell point policies of an IO point policy and a last transaction floor policy. We combine two policies from findings in a decision tree. The result indicates significant improvements. The findings from a decision tree are used as inputs in the control system so that the control system selects an appropriate policy based on a current system state. Miki Fukunari, Kristin P. Bennett, Charles J. Malmborg |
ICMLA | 2 |
| 2004 | Column-generation boosting methods for mixture of kernelsabstractWe devise a boosting approach to classification and regression based on column generation using a mixture of kernels. Traditional kernel methods construct models based on a single positive semi-definite kernel with the type of kernel predefined and kernel parameters chosen according to cross-validation performance. Our approach creates models that are mixtures of a library of kernel models, and our algorithm automatically determines kernels to be used in the final model. The 1-norm and 2-norm regularization methods are employed to restrict the ensemble of kernel models. The proposed method produces sparser solutions, and thus significantly reduces the testing time. By extending the column generation (CG) optimization which existed for linear programs with 1-norm regularization to quadratic programs with 2-norm regularization, we are able to solve many learning formulations by leveraging various algorithms for constructing single kernel models. By giving different priorities to columns to be generated, we are able to scale CG boosting to large datasets. Experimental results on benchmark data are included to demonstrate its effectiveness. Jinbo Bi, Tong Zhang 0001, Kristin P. Bennett |
KDD | 3 |
| 2003 | Regression Error Characteristic Curves
Jinbo Bi, Kristin P. Bennett |
ICML | 2 |
| 2003 | A geometric approach to support vector regression
Jinbo Bi, Kristin P. Bennett |
Neurocomputing | 2 |
| 2003 | Dimensionality Reduction via Sparse Support Vector Machines
Jinbo Bi, Kristin P. Bennett, Mark J. Embrechts, Curt M. Breneman, Minghu Song |
J. Mach. Learn. Res. | 2 |
| 2002 | Exploiting unlabeled data in ensemble methodsabstractAn adaptive semi-supervised ensemble method, ASSEMBLE, is proposed that constructs classification ensembles based on both labeled and unlabeled data. ASSEMBLE alternates between assigning "pseudo-classes" to the unlabeled data using the existing ensemble and constructing the next base classifier using both the labeled and pseudolabeled data. Mathematically, this intuitive algorithm corresponds to maximizing the classification margin in hypothesis space as measured on both the labeled and unlabeled of data. Unlike alternative approaches, ASSEMBLE does not require a semi-supervised learning method for the base classifier. ASSEMBLE can be used in conjunction with any cost-sensitive classification algorithm for both two-class and multi-class problems. ASSEMBLE using decision trees won the NIPS 2001 Unlabeled Data Competition. In addition, strong results on several benchmark datasets using both decision trees and neural networks support the proposed method. Kristin P. Bennett, Ayhan Demiriz, Richard Maclin |
KDD | 1 |
| 2002 | MARK: a boosting algorithm for heterogeneous kernel modelsabstractSupport Vector Machines and other kernel methods have proven to be very effective for nonlinear inference. Practical issues are how to select the type of kernel including any parameters and how to deal with the computational issues caused by the fact that the kernel matrix grows quadratically with the data. Inspired by ensemble and boosting methods like MART, we propose the Multiple Additive Regression Kernels (MARK) algorithm to address these issues. MARK considers a large (potentially infinite) library of kernel matrices formed by different kernel functions and parameters. Using gradient boosting/column generation, MARK constructs columns of the heterogeneous kernel matrix (the base hypotheses) on the fly and then adds them into the kernel ensemble. Regularization methods such as used in SVM, kernel ridge regression, and MART, are used to prevent overfitting. We investigate how MARK is applied to heterogeneous kernel ridge regression. The resulting algorithm is simple to implement and efficient. Kernel parameter selection is handled within MARK. Sampling and "weak" kernels are used to further enhance the computational efficiency of the resulting additive algorithm. The user can incorporate and potentially extract domain knowledge by restricting the kernel library to interpretable kernels. MARK compares very favorably with SVM and kernel ridge regression on several benchmark datasets. Kristin P. Bennett, Michinari Momma, Mark J. Embrechts |
KDD | 1 |
| 2002 | A Pattern Search Method for Model Selection of Support Vector RegressionabstractWe develop a fully-automated pattern search methodology for model selection of support vector machines (SVMs) for regression and classification. Pattern search (PS) is a derivative-free optimization method suitable for low-dimensional optimization problems for which it is difficult or impossible to calculate derivatives. This methodology was motivated by an application in drug design in which regression models are constructed based on a few high-dimensional exemplars. Automatic model selection in such underdetermined problems is essential to avoid overfitting and overestimates of generalization capability caused by selecting parameters based on testing results. We focus on SVM model selection for regression based on leave-one-out (LOO) and cross-validated estimates of mean squared error, but the search strategy is applicable to any model criterion. Because the resulting error surface produces an extremely noisy map of the model quality with many local minima, the resulting generalization capacity of any single local optimal model illustrates high variance. Thus several locally optimal SVM models are generated and then bagged or averaged to produce the final SVM. This strategy of pattern search combined with model averaging has proven to be very effective on benchmark tests and in high-variance drug design domains with high potential of overfitting. Michinari Momma, Kristin P. Bennett |
SDM | 2 |
| 2002 | Linear Programming Boosting via Column Generation
Ayhan Demiriz, Kristin P. Bennett, John Shawe-Taylor |
Mach. Learn. | 2 |
| 2002 | Sparse Regression Ensembles in Infinite and Finite Hypothesis Spaces
Gunnar Rätsch, Ayhan Demiriz, Kristin P. Bennett |
Mach. Learn. | 3 |
| 2001 | Duality, Geometry, and Support Vector RegressionabstractWe develop an intuitive geometric framework for support vector regression (SVR). By examining when (cid:15)-tubes exist, we show that SVR can be regarded as a classi(cid:12)cation problem in the dual space. Hard and soft (cid:15)-tubes are constructed by separating the convex or reduced convex hulls respectively of the training data with the response variable shifted up and down by (cid:15). A novel SVR model is proposed based on choosing the max-margin plane between the two shifted datasets. Maximizing the margin corresponds to shrinking the e(cid:11)ective (cid:15)-tube. In the proposed approach the e(cid:11)ects of the choices of all parameters become clear geometrically. Jinbo Bi, Kristin P. Bennett |
NIPS | 2 |
| 2000 | A Column Generation Algorithm For Boosting
Kristin P. Bennett, Ayhan Demiriz, John Shawe-Taylor |
ICML | 1 |
| 2000 | Duality and Geometry in SVM Classifiers
Kristin P. Bennett, Erin J. Bredensteiner |
ICML | 1 |
| 2000 | A Linear Programming Approach to Novelty DetectionabstractNovelty detection involves modeling the normal behaviour of a sys(cid:173) tem hence enabling detection of any divergence from normality. It has potential applications in many areas such as detection of ma(cid:173) chine damage or highlighting abnormal features in medical data. One approach is to build a hypothesis estimating the support of the normal data i.e. constructing a function which is positive in the region where the data is located and negative elsewhere. Recently kernel methods have been proposed for estimating the support of a distribution and they have performed well in practice - training involves solution of a quadratic programming problem. In this pa(cid:173) per we propose a simpler kernel method for estimating the support based on linear programming. The method is easy to implement and can learn large datasets rapidly. We demonstrate the method on medical and fault detection datasets. 1 Colin Campbell, Kristin P. Bennett |
NIPS | 2 |
| 2000 | Enlarging the Margins in Perceptron Decision Trees
Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor, Donghui Wu |
Mach. Learn. | 1 |
| 1999 | Large Margin Trees for Induction and Transduction
Donghui Wu, Kristin P. Bennett, Nello Cristianini, John Shawe-Taylor |
ICML | 2 |
| 1999 | On support vector decision trees for database marketingabstractWe introduce a support vector decision tree method for customer targeting in the framework of large databases (database marketing). The goal is to provide a tool to identify the best customers based on historical data. This tool is then used to forecast the best potential customers among a pool of prospects. We begin by regressively constructing a decision tree. Each decision consists of a linear combination of independent attributes. A linear program motivated by the support vector machine method from Vapnik's statistical learning theory is used to construct each decision. This linear program automatically selects the relevant subset of attributes for each decision. Each customer is scored based on the decision tree. A gain chart table is used to verify the goodness-of-fit of the targeting, to determine the likely prospects and the expected utility or profit. Successful results are given for three industrial problems. Kristin P. Bennett, Donghui Wu, Leonardo Auslender |
IJCNN | 1 |
| 1999 | Density-Based Indexing for Approximate Nearest-Neighbor QueriesabstractWe consider the problem of performing Nearest-neighbor queries efficiently over large high-dimensional databases.To avoid a full database scan, we target constructing a multidimensional index structure.It is well-accepted that traditional database indexing algorithms fail for high-dimensional data (say d > 10 or 20 depending on the scheme).Some arguments have advocated that nearest-neighbor queries do not even make sense for high-dimensional data.We show that these arguments are based on over-restrictive assumptions, and that in the general case it is meaningful and possible to build an index for such queries.Our approach, called DBIN, scales to high-dimensional databases by exploiting statistical properties of the data.The approach is based on statistically modeling the density of the content of the data table.DBIN uses the density model to derive a single index over the data table and requires physically rewriting data in a new table sorted by the newly created index (i.e.create a clustered-index).The indexing scheme produces a mapping between a query point (a data record) and an ordering on the clustered index values.Data is then scanned according to the index.We present theoretical and empirical justification for DBIN.The scheme supports a family of distance functions which includes the traditional Euclidean distance measure. Kristin P. Bennett, Usama M. Fayyad, Dan Geiger |
KDD | 1 |
| 1998 | Semi-Supervised Support Vector Machines
Kristin P. Bennett, Ayhan Demiriz |
NIPS | 1 |
| 1997 | A Parametric Optimization Method for Machine LearningabstractThe classification problem of constructing a plane to separate the members of two sets can be formulated as a parametric bilinear program. This approach was originally created to minimize the number of points misclassified. However, a novel interpretation of the algorithm is that the subproblems represent alternative error functions of the misclassified points. Each subproblem identifies a specified number of outliers and minimizes the magnitude of the errors on the remaining points. A tuning set is used to select the best result among the subproblems. A parametric Frank-Wolfe method was used to solve the bilinear subproblems. Computational results on a number of datasets indicate that the results compare very favorably with linear programming and heuristic search approaches. The algorithm can be used as part of a decision tree algorithm to create nonlinear classifiers. Kristin P. Bennett, Erin J. Bredensteiner |
INFORMS J. Comput. | 1 |