EDBT 2026 Demo / reviewers in the wild / expert
David S. Matteson
dblp:52/11128
· DBLP profile ↗
8ranked-venue papers in the field
0as first author
4since 2021 · last 2022
0000-0002-2674-0387ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 6Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | IB-GAN: A Unified Approach for Multivariate Time Series Classification under Class ImbalanceabstractClassification of large multivariate time series with strong class imbalance is an important task in real-world applications. Standard methods of class weights, over-sampling, or parametric data augmentation do not always yield significant improvements for predicting minority classes of interest. Non-parametric data augmentation with Generative Adversarial Networks (GANs) offers a promising solution. We propose Imputation Balanced GAN (IB-GAN), a novel method that joins data augmentation and classification in a one-step process via an imputation-balancing approach. IB-GAN uses imputation and resampling techniques to generate higher quality samples from randomly masked vectors than from white noise, and augments classification through a class-balanced set of real and synthetic samples. Imputation hyperparameter pmiss allows for regularization of classifier variability by tuning innovations introduced via generator imputation. IB-GAN is simple to train and model-agnostic, pairing any deep learning classifier with a generator-discriminator duo and resulting in higher accuracy for under-observed classes. Empirical experiments on open-source UCR data and a 90K product dataset show significant performance gains against state-of-the-art parametric and GAN baselines. Grace Deng, Cuize Han, Tommaso Dreossi, Clarence Lee, David S. Matteson |
SDM | 5 |
| 2022 | Extended missing data imputation via GANs for ranking applications
Grace Deng, Cuize Han, David S. Matteson |
Data Min. Knowl. Discov. | 3 |
| 2021 | Risk Identification & Quantification in Complex Human-Natural Systems via Convergent Data Intensive ResearchabstractHuman-natural systems involve complex interdependent processes, but domain specific processes are traditionally studied in non-overlapping research silos. The Predictive Risk Investigation SysteM (PRISM) for multi-layer dynamic interconnection analysis is a group of collaborators across multiple domains who work to discover data driven connections specifically among domain risks. We bring our inter-disciplinary approach to risk assessment to our KDD'21 workshop. Our workshop is a step toward a holistic approach to systemic risk analysis by welcoming speakers in applied and technical research at the forefront of risk and complex systems. Toryn L. J. Schafer, Ryan M. McGranaghan, Mila Getmansky Sherman, Mei-Ling E. Feng, Olukunle O. Owolabi, Sean E. Ryan, Marie-Christine Düker, Michael Jauch, David S. Matteson |
KDD | 9 |
| 2021 | AURORA: A Unified fRamework fOR Anomaly detection on multivariate time series
Wenyu Zhang 0003, Maxwell McNeil, Nachuan Chengwang, David S. Matteson, Petko Bogdanov |
Data Min. Knowl. Discov. | 5 |
| 2019 | ABACUS: Unsupervised Multivariate Change Detection via Bayesian Source SeparationabstractChange detection involves segmenting sequential data such that observations in the same segment share some desired properties. Multivariate change detection continues to be a challenging problem due to the variety of ways change points can be correlated across channels and the potentially poor signal-to-noise ratio on individual channels. In this paper, we are interested in locating additive outliers (AO) and level shifts (LS) in the unsupervised setting. We propose ABACUS, Automatic BAyesian Changepoints Under Sparsity, a Bayesian source separation technique to recover latent signals while also detecting changes in model parameters. Multi-level sparsity achieves both dimension reduction and modeling of signal changes. We show ABACUS has competitive or superior performance in simulation studies against state-of-the-art change detection methods and established latent variable models. We also illustrate ABACUS on two real application, modeling genomic profiles and analyzing household electricity consumption. Wenyu Zhang 0003, Daniel E. Gilbert, David S. Matteson |
SDM | 3 |
| 2016 | Leveraging cloud data to mitigate user experience from 'breaking bad'abstractLow latency and high availability of an app or a web service are key, amongst other factors, to the overall user experience (which in turn directly impacts the bottoniline). Exogenic and/or endogenic factors often give rise to breakouts in cloud data which makes maintaining high availability and delivering high performance very challenging. Existing breakout detection techniques are not suitable for cloud data owing to not being robust in the presence of anomalies. To this end, we developed a novel statistical technique to automatically detect breakouts in cloud data. This technique employs Energy Statistics to detect breakouts in both app and system metrics. Further, the technique uses robust statistical metrics, viz., medians, and estimates the statistical significance of a breakout through a permutation test. To the best of our knowledge, this is the first work which addresses breakout detection in the presence of anomalies. We demonstrate the efficacy of the proposed technique using production data and report precision, recall, and f-measure measure. The proposed technique is 3.5× faster than a state-of-the-art technique for breakout detection and is being currently used on a daily basis at Twitter Inc. Nicholas A. James, Arun Kejariwal, David S. Matteson |
IEEE BigData | 3 |
| 2016 | Mixed data and classification of transit stopsabstractAn analysis of the characteristics and behavior of individual bus stops can reveal clusters of similar stops, which can be of use in making routing and scheduling decisions, as well as determining what facilities to provide at each stop. This paper provides an exploratory analysis, including several possible clustering results, of a dataset provided by the Regional Transit Service of Rochester, NY. The dataset describes ridership on public buses, recording the time, location, and number of entering and exiting passengers each time a bus stops. A description of the overall behavior of bus ridership is followed by a stop-level analysis. We compare multiple measures of stop similarity, based on location, route information, and ridership volume over time. Laura L. Tupper, David S. Matteson, John C. Handley |
IEEE BigData | 2 |
| 2015 | Predicting Ambulance Demand: a Spatio-Temporal Kernel ApproachabstractPredicting ambulance demand accurately at fine time and location scales is critical for ambulance fleet management and dynamic deployment. Large-scale datasets in this setting typically exhibit complex spatio-temporal dynamics and sparsity at high resolutions. We propose a predictive method using spatio-temporal kernel density estimation (stKDE) to address these challenges, and provide spatial density predictions for ambulance demand in Toronto, Canada as it varies over hourly intervals. Specifically, we weight the spatial kernel of each historical observation by its informativeness to the current predictive task. We construct spatio-temporal weight functions to incorporate various temporal and spatial patterns in ambulance demand, including location-specific seasonalities and short-term serial dependence. This allows us to draw out the most helpful historical data, and exploit spatio-temporal patterns in the data for accurate and fast predictions. We further provide efficient estimation and customizable prediction procedures. stKDE is easy to use and interpret by non-specialized personnel from the emergency medical service industry. It also has significantly higher statistical accuracy than the current industry practice, with a comparable amount of computational expense. Zhengyi Zhou, David S. Matteson |
KDD | 2 |