David S. Matteson

dblp:52/11128 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-2674-0387ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Dynamic Atomic Column Detection in Transmission Electron Microscopy Videos via Ridge Estimation
abstract
Ridge detection is a classical tool to extract curvilinear features in image processing. As such, it has great promise in applications to material science problems; specifically, for trend filtering relatively stable atom-shaped objects in image sequences, such as bright-field Transmission Electron Microscopy (TEM) videos. Standard analysis of TEM videos is limited to frame-by-frame object recognition. We instead harness temporal correlation across frames through simultaneous analysis of long image sequences, specified as a spatio-temporal image tensor. We define new ridge detection algorithms to non-parametrically estimate explicit trajectories of atomic-level object locations as a continuous function of time. Our approach is specially tailored to handle temporal analysis of objects that seemingly stochastically disappear and subsequently reappear throughout a sequence. We demonstrate that the proposed method is highly effective in simulation scenarios, and delivers notable performance improvements in TEM experiments compared to other material science benchmarks.
Yuchen Xu 0007, Andrew M. Thomas, Peter A. Crozier, David S. Matteson
IEEE Trans. Image Process.4
2022 IB-GAN: A Unified Approach for Multivariate Time Series Classification under Class Imbalance
abstract
Classification of large multivariate time series with strong class imbalance is an important task in real-world applications. Standard methods of class weights, over-sampling, or parametric data augmentation do not always yield significant improvements for predicting minority classes of interest. Non-parametric data augmentation with Generative Adversarial Networks (GANs) offers a promising solution. We propose Imputation Balanced GAN (IB-GAN), a novel method that joins data augmentation and classification in a one-step process via an imputation-balancing approach. IB-GAN uses imputation and resampling techniques to generate higher quality samples from randomly masked vectors than from white noise, and augments classification through a class-balanced set of real and synthetic samples. Imputation hyperparameter pmiss allows for regularization of classifier variability by tuning innovations introduced via generator imputation. IB-GAN is simple to train and model-agnostic, pairing any deep learning classifier with a generator-discriminator duo and resulting in higher accuracy for under-observed classes. Empirical experiments on open-source UCR data and a 90K product dataset show significant performance gains against state-of-the-art parametric and GAN baselines.
Grace Deng, Cuize Han, Tommaso Dreossi, Clarence Lee, David S. Matteson
SDM5
2022 Bayesian spillover graphs for dynamic networks
abstract
We present Bayesian Spillover Graphs (BSG), a novel method for learning temporal relationships, identifying critical nodes, and quantifying uncertainty for multi-horizon spillover effects in a dynamic system. BSG leverages both an interpretable framework via forecast error variance decompositions (FEVD) and comprehensive uncertainty quantification via Bayesian time series models to contextualize temporal relationships in terms of systemic risk and prediction variability. Forecast horizon hyperparameter h allows for learning both short-term and equilibrium state network behaviors. Experiments for identifying source and sink nodes under various graph and error specifications show significant performance gains against state-of-the-art Bayesian Networks and deep-learning baselines. Applications to real-world systems also showcase BSG as an exploratory analysis tool for uncovering indirect spillovers and quantifying systemic risk.
Grace Deng, David S. Matteson
UAI2
2022 Extended missing data imputation via GANs for ranking applications
Grace Deng, Cuize Han, David S. Matteson
Data Min. Knowl. Discov.3
2021 Graph-Based Continual Learning
Binh Tang, David S. Matteson
ICLR2
2021 Risk Identification & Quantification in Complex Human-Natural Systems via Convergent Data Intensive Research
abstract
Human-natural systems involve complex interdependent processes, but domain specific processes are traditionally studied in non-overlapping research silos. The Predictive Risk Investigation SysteM (PRISM) for multi-layer dynamic interconnection analysis is a group of collaborators across multiple domains who work to discover data driven connections specifically among domain risks. We bring our inter-disciplinary approach to risk assessment to our KDD'21 workshop. Our workshop is a step toward a holistic approach to systemic risk analysis by welcoming speakers in applied and technical research at the forefront of risk and complex systems.
Toryn L. J. Schafer, Ryan M. McGranaghan, Mila Getmansky Sherman, Mei-Ling E. Feng, Olukunle O. Owolabi, Sean E. Ryan, Marie-Christine Düker, Michael Jauch, David S. Matteson
KDD9
2021 Probabilistic Transformer For Time Series Analysis
abstract
Generative modeling of multivariate time series has remained challenging partly due to the complex, non-deterministic dynamics across long-distance timesteps. In this paper, we propose deep probabilistic methods that combine state-space models (SSMs) with transformer architectures. In contrast to previously proposed SSMs, our approaches use attention mechanism to model non-Markovian dynamics in the latent space and avoid recurrent neural networks entirely. We also extend our models to include several layers of stochastic variables organized in a hierarchy for further expressiveness. Compared to transformer models, ours are probabilistic, non-autoregressive, and capable of generating diverse long-term forecasts with uncertainty estimates. Extensive experiments show that our models consistently outperform competitive baselines on various tasks and datasets, including time series forecasting and human motion prediction.
Binh Tang, David S. Matteson
NeurIPS2
2021 AURORA: A Unified fRamework fOR Anomaly detection on multivariate time series
Wenyu Zhang 0003, Maxwell McNeil, Nachuan Chengwang, David S. Matteson, Petko Bogdanov
Data Min. Knowl. Discov.5
2020 High Dimensional Forecasting via Interpretable Vector Autoregression
abstract
Vector autoregression (VAR) is a fundamental tool for modeling multivariate time series. However, as the number of component series is increased, the VAR model becomes overparameterized. Several authors have addressed this issue by incorporating regularized approaches, such as the lasso in VAR estimation. Traditional approaches address overparameterization by selecting a low lag order, based on the assumption of short range dependence, assuming that a universal lag order applies to all components. Such an approach constrains the relationship between the components and impedes forecast performance. The lasso-based approaches perform much better in high-dimensional situations but do not incorporate the notion of lag order selection. We propose a new class of hierarchical lag structures (HLag) that embed the notion of lag selection into a convex regularizer. The key modeling tool is a group lasso with nested groups which guarantees that the sparsity pattern of lag coefficients honors the VAR's ordered structure. The proposed HLag framework offers three basic structures, which allow for varying levels of flexibility, with many possible generalizations. A simulation study demonstrates improved performance in forecasting and lag order selection over previous approaches, and macroeconomic, financial, and energy applications further highlight forecasting improvements as well as HLag's convenient, interpretable output.
William B. Nicholson, Ines Wilms, Jacob Bien, David S. Matteson
J. Mach. Learn. Res.4
2019 Independent Component Analysis Based on Mutual Dependence Measures
abstract
We apply both distance-based and kernel-based mutual dependence measures to independent component analysis (ICA), and generalize dCovICA to MDMICA, minimizing empirical dependence measures as an objective function in both deflation and parallel manners. Solving this minimization problem, we introduce Latin hypercube sampling (LHS), and a global optimization method, Bayesian optimization (BO) to improve the initialization of the Newton-type local optimization method. The performance of MDMICA is evaluated in various simulation studies and an image data example. When the ICA model is correct, MDMICA achieves competitive results compared to existing approaches. When the ICA model is misspecified, the estimated independent components are less mutually dependent than the observed components using MDMICA, while the estimated independent components are prone to be even more mutually dependent than the observed components using other approaches.
Ze Jin, David S. Matteson, Tianrong Zhang
ICMLA2
2019 ABACUS: Unsupervised Multivariate Change Detection via Bayesian Source Separation
abstract
Change detection involves segmenting sequential data such that observations in the same segment share some desired properties. Multivariate change detection continues to be a challenging problem due to the variety of ways change points can be correlated across channels and the potentially poor signal-to-noise ratio on individual channels. In this paper, we are interested in locating additive outliers (AO) and level shifts (LS) in the unsupervised setting. We propose ABACUS, Automatic BAyesian Changepoints Under Sparsity, a Bayesian source separation technique to recover latent signals while also detecting changes in model parameters. Multi-level sparsity achieves both dimension reduction and modeling of signal changes. We show ABACUS has competitive or superior performance in simulation studies against state-of-the-art change detection methods and established latent variable models. We also illustrate ABACUS on two real application, modeling genomic profiles and analyzing household electricity consumption.
Wenyu Zhang 0003, Daniel E. Gilbert, David S. Matteson
SDM3
2018 Testing for Conditional Mean Independence with Covariates through Martingale Difference Divergence
Ze Jin, Xiaohan Yan, David S. Matteson
UAI3
2016 Leveraging cloud data to mitigate user experience from 'breaking bad'
abstract
Low latency and high availability of an app or a web service are key, amongst other factors, to the overall user experience (which in turn directly impacts the bottoniline). Exogenic and/or endogenic factors often give rise to breakouts in cloud data which makes maintaining high availability and delivering high performance very challenging. Existing breakout detection techniques are not suitable for cloud data owing to not being robust in the presence of anomalies. To this end, we developed a novel statistical technique to automatically detect breakouts in cloud data. This technique employs Energy Statistics to detect breakouts in both app and system metrics. Further, the technique uses robust statistical metrics, viz., medians, and estimates the statistical significance of a breakout through a permutation test. To the best of our knowledge, this is the first work which addresses breakout detection in the presence of anomalies. We demonstrate the efficacy of the proposed technique using production data and report precision, recall, and f-measure measure. The proposed technique is 3.5× faster than a state-of-the-art technique for breakout detection and is being currently used on a daily basis at Twitter Inc.
Nicholas A. James, Arun Kejariwal, David S. Matteson
IEEE BigData3
2016 Mixed data and classification of transit stops
abstract
An analysis of the characteristics and behavior of individual bus stops can reveal clusters of similar stops, which can be of use in making routing and scheduling decisions, as well as determining what facilities to provide at each stop. This paper provides an exploratory analysis, including several possible clustering results, of a dataset provided by the Regional Transit Service of Rochester, NY. The dataset describes ridership on public buses, recording the time, location, and number of entering and exiting passengers each time a bus stops. A description of the overall behavior of bus ridership is followed by a stop-level analysis. We compare multiple measures of stop similarity, based on location, route information, and ridership volume over time.
Laura L. Tupper, David S. Matteson, John C. Handley
IEEE BigData2
2015 Predicting Ambulance Demand: a Spatio-Temporal Kernel Approach
abstract
Predicting ambulance demand accurately at fine time and location scales is critical for ambulance fleet management and dynamic deployment. Large-scale datasets in this setting typically exhibit complex spatio-temporal dynamics and sparsity at high resolutions. We propose a predictive method using spatio-temporal kernel density estimation (stKDE) to address these challenges, and provide spatial density predictions for ambulance demand in Toronto, Canada as it varies over hourly intervals. Specifically, we weight the spatial kernel of each historical observation by its informativeness to the current predictive task. We construct spatio-temporal weight functions to incorporate various temporal and spatial patterns in ambulance demand, including location-specific seasonalities and short-term serial dependence. This allows us to draw out the most helpful historical data, and exploit spatio-temporal patterns in the data for accurate and fast predictions. We further provide efficient estimation and customizable prediction procedures. stKDE is easy to use and interpret by non-specialized personnel from the emergency medical service industry. It also has significantly higher statistical accuracy than the current industry practice, with a comparable amount of computational expense.
Zhengyi Zhou, David S. Matteson
KDD2
2013 Locally stationary vector processes and adaptive multivariate modeling
abstract
The assumption of strict stationarity is too strong for observations in many financial time series applications; however, distributional properties may be at least locally stable in time. We define multivariate measures of homogeneity to quantify local stationarity and an empirical approach for robustly estimating time varying windows of stationarity. Finally, we consider bivariate series that are believed to be cointegrated locally, assess our estimates, and discuss applications in financial asset pairs trading.
David S. Matteson, Nicholas A. James, William B. Nicholson, Louis C. Segalini
ICASSP1